Intent-Based Load Balancing for AI Traffic on BIG-IP

AI applications often start with a simple pattern: one client, one API endpoint, one model. That works for a demo, but it becomes limiting quickly.

In real environments, different requests may need different treatment. A technical support question may need a specialist model. A general question may go to a lower-cost model. A disallowed request should be blocked locally. A customer or application key may need its own routing policy. A provider API key may fail, hit a rate limit, or need controlled failover.

That is the idea behind Intent-Based Load Balancing.

Instead of routing only by network attributes, this project adds an AI-aware decision layer on BIG-IP. The gateway accepts OpenAI-compatible requests, classifies the user intent, applies routing policy, and then either responds locally or forwards the request to the selected model backend.

The current implementation runs natively on BIG-IP using:

  • iApps LX for the on-box management UI.
  • iRules and TMM for data-plane request handling.
  • iRules LX for lightweight runtime decisions.
  • BIG-IP LTM pools for backend and classifier egress.
  • OpenAI-compatible northbound traffic, primarily /v1/chat/completions.

The operator workflow is intentionally simple. Create a classifier, define backend targets, configure routing policies, create a listener, and deploy the configuration to BIG-IP.

A few example scenarios:

Route F5 questions to a specialist model

A user asks: “How do I configure an LTM pool monitor?”

The classifier tags the request as f5. The routing policy sends it to a backend target with an F5-focused system prompt and model configuration. A more general question can be routed to a different model.

Block unsafe or unwanted requests locally

If the classifier returns a tag such as blocked, BIG-IP can return a local response immediately. The prompt does not need to be forwarded to a model provider.

Use Virtual Keys for application-level control

Different internal applications can use different Virtual Keys. Policies can match by key pool, a specific key, or a key tag. This makes it possible to route or restrict traffic by application, tenant, or use case.

Fail over model provider API keys

Backend targets can use Model Credential pools. If a provider key returns failures such as 401, 403, or 429, the runtime can cool down that key and allow later requests to use the next available credential. This is intentionally conservative and avoids same-request retry behavior that could create duplicate billing risk.

Keep BIG-IP as the control point

The project does not require an external control-plane service. The UI, runtime, iRule, and installer are packaged for BIG-IP. Backend pool members, monitors, and load-balancing methods remain owned by BIG-IP Local Traffic, which keeps the design aligned with normal BIG-IP operations.

The current offline installer can build a self-contained package and install the gateway on a clean BIG-IP device. After installation, the operator configures the gateway from TMUI and uses Deploy Changes to apply the configuration.

This project is still evolving, but the direction is clear: AI traffic should not be treated as just another HTTP request. Intent, identity, safety, cost, and backend health all matter. BIG-IP already sits in the right place to make those decisions.

Demo Vedio

GitHub:

1 Like