AI and APIs
Over the past several years, the industry has seen an explosion of interest in large language models and AI driven applications. Much of the discussion has focused on the models themselves: their size, their capabilities, and their apparent ability to reason, summarize, and generate content. In the process, it is easy to overlook a more fundamental reality.
Modern AI systems are still API systems.
Despite new abstractions and new terminology, the underlying mechanics of AI applications remain familiar. Requests are sent, responses are returned. Identities are authenticated, authorization decisions are made, data is retrieved, and actions are executed. These interactions happen over APIs, and the reliability, security, and scalability of AI systems are constrained by the same architectural principles that have always governed distributed systems.
What is new is not the presence of APIs, but the nature of the consumer calling them. In traditional systems, API consumers are deterministic. They are code written by engineers who read the documentation and invoke endpoints in predictable ways. In AI systems, the consumer is increasingly a model, a probabilistic component that infers behavior from schemas, chains calls dynamically, and produces traffic patterns that were not explicitly programmed. That single shift is what makes every downstream concern in this series, including MCP design, token budgets, authorization, and operations, behave differently than in traditional API platforms.
Understanding this relationship is critical, not only for building AI systems, but for operating and securing them in production.
AI Applications as API Orchestration Platforms
At a high level, an AI application is best understood not as a single model invocation, but as an orchestration layer that coordinates multiple API interactions. A typical request may involve:
- A client calling an application API
- Authentication and authorization checks
- Retrieval of contextual data from internal or external services
- One or more calls to a model inference endpoint
- Follow-on tool or service calls triggered by the model’s output
- Aggregation and formatting of the final response
From an architectural perspective, this is not fundamentally different from any other multi-service application. Routing, observability, traffic management, and trust boundaries remain as relevant here as in any traditional platform. What has changed is that the decision logic, meaning when to call which service and with what parameters, is increasingly driven by model output rather than static application code.
That shift does not eliminate APIs. It increases their importance.
AI Application as an Orchestration Platform
Models as API Endpoints, Not Black Boxes
In production environments, models are consumed almost exclusively through APIs. Whether hosted by a third party or deployed internally, a model is exposed as an endpoint that accepts structured input and returns structured output.
Treating models as API endpoints clarifies several important points. A model does not “see” your system. It receives a request payload, processes it, and returns a response. Everything the model knows about your environment arrives through an API boundary.
What distinguishes model endpoints from conventional APIs is not their interface, but their operational profile. Responses are frequently streamed rather than returned as a single payload, which changes how load balancers, proxies, and timeouts behave. Payload sizes are highly variable, with both requests and responses ranging from a few hundred bytes to many megabytes depending on context and output length. Rate limits are often expressed in tokens per minute rather than requests per second, which complicates capacity planning and quota enforcement. Self-hosted models introduce additional concerns around GPU scheduling, cold start latency, and memory pressure that do not exist for traditional stateless services.
These characteristics do not change the fundamental nature of a model as an API endpoint. They do mean that the operational assumptions built into the existing API infrastructure may not hold without adjustment.
Tools, Retrieval, and Data Access Are Still APIs
As AI systems evolve beyond simple prompt-and-response interactions, they increasingly rely on tools: databases, search systems, ticketing platforms, code repositories, and internal business services. These tools are almost always accessed through APIs.
Retrieval-augmented generation, for example, is often described as a novel AI pattern. In practice, it is a sequence of API calls:
- An embedding service is called to encode a query
- A vector database is queried for relevant results
- A document store is accessed to retrieve source material
- The retrieved data is passed to the model as context
Each step carries the usual concerns: latency, authorization, data exposure, and error handling. The model may influence when these calls occur, but it does not change their fundamental nature.
Why API Design Matters More in AI Systems
If AI systems are built on APIs, why do they feel harder to manage?
The answer lies in amplification.
Model-driven systems tend to:
- Chain API calls dynamically
- Surface data in ways developers did not explicitly anticipate
- Expand the blast radius of a misconfigured authorization
- Increase sensitivity to payload size and response shape
A poorly designed API that returns excessive data may be tolerable in a traditional application. In an AI system, that same response can overflow context limits, leak sensitive information into prompts, or cascade into additional unintended tool calls.
This amplification rarely stays within a single domain. A schema decision that looks like an application concern becomes a traffic and routing concern when responses grow unpredictably, and an authorization concern when a model uses that response to drive the next call. Design choices that were once contained within one team’s scope now propagate across the stack.
In this sense, AI does not introduce entirely new architectural risks. It magnifies existing ones.
Introducing MCP as an API Coordination Layer
As models gain the ability to invoke tools directly, the need for consistent, structured access to APIs becomes more pressing. This is where Model Context Protocol (MCP) enters the picture.
At a conceptual level, MCP does not replace APIs. It standardizes how AI systems discover, describe, and invoke API-backed tools. MCP servers typically sit in front of existing services, exposing them in a model-friendly way while relying on the same underlying API infrastructure.
Seen through this lens, MCP is not a departure from established architecture patterns. It is an adaptation, one that acknowledges models as active participants in API-driven systems rather than passive consumers of text. But it is also the introduction of a new coordination layer, a tool plane, with its own operational, network, and security properties that do not map cleanly onto the API layer beneath it. The rest of this series examines what that means for the systems you build, run, and secure.
Looking Ahead
If AI systems are still API systems, then the familiar disciplines of API architecture, security, and operations remain essential. What changes is where decisions are made, how data flows, and how quickly small design flaws can propagate.
The next article looks more closely at MCP itself, examining how it standardizes tool access on top of APIs and why treating it as a tool plane helps clarify both its power and its risks. From there, the series turns to tokens as a first-class design constraint that shapes tool schemas, response shaping, and traffic behavior. The fourth article addresses authorization and the security implications of letting models invoke tools directly, including identity, delegation, and the expanded blast radius MCP introduces. The series closes with a look at operating MCP-enabled systems in production, where reliability, cost, and safety have to be enforced rather than assumed.
Resources:
Article Series:
- MCP, APIs, and Tokens: Building and Securing the Tool Plane of AI Systems (Intro)
- MCP, APIs, and Tokens (Part 1 - APIs First: Why AI Systems Are Still API Systems)
- MCP, APIs, and Tokens (Part 2 - MCP as the Tool Plane: Standardizing Access Across APIs)
- MCP, APIs, and Tokens (Part 3 - Tokens as a Design Constraint for MCP and APIs)
- MCP, APIs, and Tokens (Part 4 - Securing the Tool Plane: MCP, APIs, and Authorization)
- MCP, APIs, and Tokens (Part 5 - Designing for the Inference Track: Safe, Scalable MCP Systems)

