Traffic flow from outside → inside the cluster.
In an earlier post I put open-source Kong behind an F5 BIG-IP on OpenShift and showed the two are complementary: the ADC owns the edge (TLS, WAF, pod-direct load balancing), the gateway owns API management (routing, rate-limiting, auth). The whole point is that the pattern is modular — the edge layer and API gateway layer can be switched for alternatives.
I’ll show this again. This time the in-cluster gateway is MuleSoft’s Flex Gateway. Same BIG-IP, same CIS, same OpenShift cluster, same “client → BIG-IP → gateway → backends” architecture. Only the middle layer changed.
The end state is the same as my previous article:
client ──HTTPS──▶ F5 BIG-IP (TLS + LB via CIS) ──HTTP──▶ Flex Gateway (routing + rate-limit + JWT) ──▶ backend apps
Part 1 — The setup journey
local K8s configurations and --connected=false
How to install Mulesoft is outside the scope of this article, but I will outline some points so that the environment is clear.
When running Flex Gateway in K8s, you can run in a connected or disconnected mode. This is configured when generating a registration file, so it’s not something you can toggle once you’ve deployed in K8s.
Connected mode means the gateway receives it’s configuration from a hosted control plane, and not resources in the K8s cluster.
In disconnected mode, the gateway is configured with K8s custom resources that are local to the cluster. There is still a registration step at time of deployment, so your pods must have outbound connectivity. This is more like my previous example with Kong, so I’ve stuck with disconnected mode here.
Without diving into how to deploy Mulesoft on Openshift, I’ll note that when you create your registration.yaml file, ensure you’ve done so with --conected=false. Also, I’ll be careful with terms. I’m calling it “disconnected mode” but my searches of Mulesoft docs, Helm values, and logs appeared to also use the terms “local” and “offline” for this model of local configuration.
Part 2 — The architecture
Three layers, each independently testable, each ignorant of the others’ internals:
┌───────────────────────────────────────────────────────────────────┐
│ Layer 1 — Network / edge (F5 BIG-IP via CIS) │
│ • TLS termination (self-signed cert for this demo) │
│ • Pod-direct load balancing to the gateway (pool = pod IPs:8081) │
│ • WAF attach point (present via Policy CR; not exercised here) │
└───────────────────────────────────────────────────────────────────┘
│ HTTP :8081
▼
┌───────────────────────────────────────────────────────────────────┐
│ Layer 2 — API gateway (MuleSoft Omni Gateway v1.13.3, local mode) │
│ • Path routing: /a → app-a, /b → app-b │
│ • Rate limiting: 5 requests / 10 s, keyed on client IP │
│ • JWT validation: HS256 bearer token, offline (no API Manager) │
└───────────────────────────────────────────────────────────────────┘
│ HTTP :8080
▼
┌───────────────────────────────────────────────────────────────────┐
│ Layer 3 — Backends (namespace mulesoft-demo, 2 replicas each) │
│ • app-a: traefik/whoami (echoes request/pod metadata) │
│ • app-b: mendhak/http-https-echo (echoes headers + body) │
└───────────────────────────────────────────────────────────────────┘
In terms of the K8s cluster, CIS runs in namespace kube-system; The gateway lives in namespace gateway; the backends live in mulesoft-demo. BIG-IP routes on hostname to the Mulesoft gateway.
Here’s the pods and service in the gateway and mulesoft-demo namespaces.
$ oc get pods -n gateway -o wide
NAME READY STATUS RESTARTS AGE IP NODE
ingress-f88b6c598-txxlk 1/1 Running 1 4d21h 10.129.0.19 ip-10-0-2-184...
$ oc get pods -n mulesoft-demo -o wide
NAME READY STATUS RESTARTS AGE IP NODE
app-a-6fb74c6c6c-9qvj9 1/1 Running 1 4d19h 10.129.0.26 ip-10-0-2-184...
app-a-6fb74c6c6c-nxhvh 1/1 Running 1 4d19h 10.129.0.24 ip-10-0-2-184...
app-b-5ff456446f-txr29 1/1 Running 1 4d19h 10.129.0.27 ip-10-0-2-184...
app-b-5ff456446f-w87kj 1/1 Running 1 4d19h 10.129.0.28 ip-10-0-2-184...
$ oc get svc -n gateway
NAME TYPE CLUSTER-IP PORT(S)
ingress ClusterIP 172.30.194.53 80/TCP,443/TCP
ingress-8081 ClusterIP 172.30.137.137 8081/TCP # <- the demo listener
ingress-hl ClusterIP None <none>
The pod is named ingress because that’s the Helm release name for the Flex gateway chart — don’t let it fool you into thinking it’s the OpenShift ingress router.
Part 3 — MuleSoft Gateway configuration
Everything here is done the Kubernetes-native way: CRD’s. An ApiInstance for the listener + routing, PolicyBinding for policy. That’s why disconnected mode matters — in connected mode these CRDs are ignored.
Getting routing right: the ApiInstance
This is the resource that defines a listener on the gateway. I mentally think of this like a VirtualServer in CIS or NGINX, if you’re familiar with those.
apiVersion: gateway.mulesoft.com/v1alpha1
kind: ApiInstance
metadata:
name: demo-api
namespace: gateway
spec:
address: http://0.0.0.0:8081 # listen on :8081 — see the port note below
services:
backend-a:
address: http://app-a.mulesoft-demo.svc.cluster.local:8080
routes:
- rules:
- path: /a
backend-b:
address: http://app-b.mulesoft-demo.svc.cluster.local:8080
routes:
- rules:
- path: /b
Why :8081? The Helm chart already stands up listeners on :80 and :443 (you’ll see ApiInstance/ingress-http and ingress-https in the namespace). Rather than modify those, I gave the demo API its own listener on a high port.
Let’s prove routing works before we add security policy. Port-forward to the listener and hit each path:
oc port-forward -n gateway svc/ingress-8081 8081:8081 &
curl -s http://localhost:18081/a # -> app-a (traefik/whoami)
curl -s http://localhost:18081/b # -> app-b (mendhak/http-https-echo)
Requests to /a returns a whoami dump (hostname, pod IPs, request headers); /b returns the echo server’s JSON. Now let’s layer policy on top.
Rate limiting
After some back-and-forth to get this right, here’s my working config of rate-limiting applied to my API instance. The custom resource is a PolicyBinding.
apiVersion: gateway.mulesoft.com/v1alpha1
kind: PolicyBinding
metadata:
name: demo-rate-limit
namespace: gateway
spec:
policyRef:
name: rate-limiting-flex # NOT "rate-limit"
targetRef:
kind: ApiInstance # required — not optional
name: demo-api
config:
rateLimits: # keySelector/exposeHeaders/clusterizable
- maximumRequests: 5 # live INSIDE this array item
timePeriodInMilliseconds: 10000
exposeHeaders: true
clusterizable: false # false for this single-replica demo; true in prod
keySelector: "#[attributes.clientAddress.ip]" # rate-limit per client IP
Let’s test it — five in, then rejections:
$ for i in $(seq 1 10); do
curl -s -o /dev/null -w "req $i -> %{http_code}\n" localhost:18081/a
done
req 1 -> 200
req 2 -> 200
req 3 -> 200
req 4 -> 200
req 5 -> 200
req 6 -> 429
req 7 -> 429
req 8 -> 429
req 9 -> 429
req 10 -> 429
$ curl -s localhost:18081/a # body once the budget is spent
{"error":"Too Many Requests"}
JWT validation
This PolicyBinding took me even more back-and-forth to get right. But this is what worked for me to enforce JWT token auth. Notice my demo only key that is 32 characters long.
apiVersion: gateway.mulesoft.com/v1alpha1
kind: PolicyBinding
metadata:
name: demo-jwt-auth
namespace: gateway
spec:
policyRef:
name: jwt-validation-flex # NOT "jwt-validation"
targetRef:
kind: ApiInstance
name: demo-api
config:
signingMethod: hmac # not "jwtSigningMethod"
signingKeyLength: 256 # not "jwtSigningKeyLength"
jwtKeyOrigin: text
textKey: "12345678901234567890123456789012" # not "jwtKey"; plain text, NOT base64
jwtOrigin: httpBearerAuthenticationHeader
validateAudClaim: false
mandatoryExpClaim: true
skipClientIdValidation: true
Now let’s test auth. Firstly, generate a JWT using that demo only key (standard PyJWT):
import time, jwt
now = int(time.time())
payload = {"sub": "test", "iat": now, "nbf": now, "exp": now + 3600} # 1-hour expiry
print(jwt.encode(payload, "12345678901234567890123456789012", algorithm="HS256"))
Test the three cases:
$ curl -s -o /dev/null -w "%{http_code}\n" localhost:8081/a
400 # no token
$ curl -s localhost:18081/a
{"error":"JWT Token is required."}
$ curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer $TOKEN" localhost:8081/a
200 # valid token → whoami
$ curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer not.a.real.token" localhost:8081/a
401 # bad token
Note the status semantics: missing token is 400 (malformed request), present-but-invalid is 401 (unauthorized). Some clients treat those very differently.
Part 4 — F5 CIS integration
I’ve changed a few parameters with my CIS deployment since my previous article, just to show a different way to configure CIS. CIS will monitor either of these: Routes, CRD’s, or Ingresses. Since Ingresses are older, depracated and rarely used, the reality is more likely you will ask: “should I use Routes or CRD’s?”.
If you don’t use OpenShift, Routes are not an option, so use CRD’s. If you do use OpenShift, some organizations like to use Routes because they are comfortable and familiar with them. If you do this, you should specify the following arguments when deploying CIS:
- manage-routes=true
- route-vserver-addr=[IP address of VIP to create on BIG-IP]
OR, if you want more configuration options for your VIPs on BIG-IP, use “NextGen Routes”:
- extended-spec-configmap=“namespace/configmap-name”
- controller-mode=openshift
Because the concept of an additional configMap is one more configuration item, I typically find using F5 CRD’s simpler. Still, I’m using NextGen Routes here an an example.
CIS runs in kube-system, watching the relevant namespaces, in OpenShift controller mode with cluster (pod-direct) pool members:
$ oc get deploy f5cis-f5-bigip-ctlr -n kube-system -o jsonpath='{.spec.template.spec.containers[0].args}'
--bigip-partition=openshift
--bigip-url=10.0.4.11
--controller-mode=openshift
--extended-spec-configmap=["namespace/configmap-name"]
--pool-member-type=cluster # pool members = pod IPs, not NodePorts
--orchestration-cni=ovn-k8s
--static-routing-mode=true
--namespace=kube-system
--namespace=gateway
--namespace=mulesoft-demo
--log-level=debug
--insecure=true
pool-member-type=cluster with orchestration-cni=ovn-k8s and static-routing-mode=true is what gives us pod-direct load balancing on OVN-Kubernetes: BIG-IP balances straight across the gateway’s pod IPs, no NodePort or router hop, and pool membership tracks pods automatically.
The Route (edge TLS decryption, HTTP→HTTPS redirect, targeting the :8081 Service):
apiVersion: route.openshift.io/v1
kind: Route
metadata:
annotations:
virtual-server.f5.com/clientssl: "my-route-tls-secret"
name: mulesoft-route
namespace: gateway
labels:
f5type: systest # picked up by CIS
spec:
host: mulesoft.my-f5.com
path: /
port:
targetPort: 8081
tls:
termination: edge
insecureEdgeTerminationPolicy: Redirect
to:
kind: Service
name: ingress-8081
weight: 100
---
apiVersion: v1
kind: Secret
metadata:
name: my-route-tls-secret
namespace: default
type: kubernetes.io/tls
data:
tls.crt: <base64-encoded-certificate>
tls.key: <base64-encoded-private-key>
The extended-spec ConfigMap pins the VIP, partition, and a Policy CR (TCP profiles + SNAT; the WAF attach point lives here too):
...[see examples of extended spec config map for full format]
extendedRouteSpec:
- namespace: gateway
vserverAddr: 10.0.0.101 # the BIG-IP VIP
vserverName: mulesoft-demo-routes
bigIpPartition: mulesoft-demo
policyCR: kube-system/security-policy
Part 5 — End-to-end testing
The value of the three-layer design is that you can test each layer in isolation. If layer N misbehaves, the bug is in N — not upstream.
Layer 1 — gateway routing only (port-forward, no edge):
curl -s localhost:8081/a # app-a
curl -s localhost:8081/b # app-b
Layer 2 — rate limiting (port-forward):
$ for i in $(seq 1 10); do curl -s -o /dev/null -w "%{http_code} " localhost:18081/a; done
200 200 200 200 200 429 429 429 429 429
Layer 3 — JWT (port-forward):
$ curl -s -o /dev/null -w "%{http_code}\n" localhost:8081/a # 400 (no token)
$ curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer $TOKEN" localhost:8081/a # 200
$ curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer garbage" localhost:8081/a # 401
Layer 4 — the whole stack through BIG-IP:
# no token
$ curl -sk -o /dev/null -w "%{http_code}\n" https://mulesoft.my-f5.com/a
400
$ curl -sk https://mulesoft.my-f5.com/a
{"error":"JWT Token is required."}
# invalid token
$ curl -sk -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer bad" https://mulesoft.my-f5.com/a
401
$ curl -sk -H "Authorization: Bearer bad" https://mulesoft.my-f5.com/a
{"error":"Invalid token."}
# valid token → routed to app-a (whoami)
$ curl -sk -H "Authorization: Bearer $TOKEN" https://mulesoft.my-f5.com/a
Hostname: app-a-6fb74c6c6c-nxhvh
IP: 10.129.0.24
GET /a HTTP/1.1
Host: app-a.mulesoft-demo.svc.cluster.local
...
# valid token → routed to app-b (echo)
$ curl -sk -H "Authorization: Bearer $TOKEN" https://mulesoft.my-f5.com/b
{ "path": "/b", "headers": { "host": "app-b.mulesoft-demo.svc.cluster.local", ... } }
# HTTP is redirected to HTTPS at the edge
$ curl -sk -o /dev/null -w "%{http_code} -> %{redirect_url}\n" http://mulesoft.my-f5.com/a
302 -> https://mulesoft.my-f5.com:443/a
# rate limiting still enforced end-to-end (window mid-cycle here)
$ for i in $(seq 1 10); do
curl -sk -o /dev/null -w "%{http_code} " -H "Authorization: Bearer $TOKEN" https://mulesoft.my-f5.com/a
done
200 200 200 429 429 429 429 429 429 429
Same behavior through the edge as at the gateway — which is exactly what “the layers are decoupled” is supposed to mean.
Wrapping up
Two gateways, one edge pattern. Kong or MuleSoft Gateway, the BIG-IP side does not change: TLS and pod-direct load balancing at the edge via CIS, API routing/rate-limiting/auth inside the cluster. That’s the payoff of clean layering - a separation of concerns with easily supportable operations.
