codefionn

The bad design of LLM APIs

I'm currently developing many different AI-enabled applications and specifically created llmleaf, an AI gateway proxy that unifies and enriches all kinds of LLM API providers.

During this I had some issues with these interfaces and sometimes with their documentation. This article covers the issues that the AI providers somehow ignore.

SSE

How would you implement a bidirectional communication protocol? Single requests? Custom binary protocol? Maybe WebSockets?

The answer chosen by the industry was of course SSE. It's a web standard for streaming events from the server to the client, which is good. Crucially, it doesn't stream the other way around, which is bad. Maybe this was a good decision in the early ChatGPT days, but the post-Opus 4.8 era needs something different.

This has two issues. When the client has to send messages back to the server, which in the age of agentic development includes a lot of tool call results, a completely new request has to be made to the server. This is made worse by prompt caching being a requirement for cost reasons and by thinking being hidden or encrypted.

When the new request is sent to answer a tool call, the API has to infer the last server response from the context, probably via hashing. Once it's identified, caching works properly and the user hopefully pays less for tokens that were already processed. The implementor has to ensure that the messages sent are detected as the same ones sent previously.

There are even special API features that try to work around this, e.g. Anthropic's cache_control or OpenAI's previous_response_id.

All this reconnecting has a latency cost, which would not be paid by proper implementations. If you want to experience this, just use Codex with the OpenAI subscription. It already uses WebSockets.

The agent sends the full history to the LLM API, gets an SSE stream until the next tool call, and sends the tool result as a new request with the full history again
Over SSE, every tool call means a new request with the whole history
The agent keeps one WebSocket open, messages flow both ways over it and only new tool results are sent, while the LLM API keeps the state in memory
Over a WebSocket, the connection stays open and only new input is sent

MCPs

The previous MCP specs were so bad that the 2026-07-28 spec had to fix them. FYI: the previous specs required stateful connections, which made stuff like letting your customers' agents read your documentation over MCP way harder and more resource intensive than it needed to be. The 2026-07-28 spec finally makes MCP stateless.

Another issue is that the previous specs made every agent session spawn its own MCP connections, which made 3 or more agents very costly with some larger MCP servers.

The list endpoint

AI providers have endpoints for listing supported models. The issue: they often don't provide necessary information. In the worst case, you will just get model ids, names and release dates.

curl request to the OpenAI models endpoint for gpt-6-astra, returning only id, object, created, owned_by and shutdown_date
OpenAI's model list: no context window, no pricing, no capabilities

But there's a lot of information required for implementing an AI-enabled application: context window length, max output tokens, pricing, supported parameters and capabilities like thinking, input and output modalities (e.g. can receive an image), automatic context window compression (summarization).

This is why providers like OpenRouter are an easy solution for applications that support "any" provider, because they provide essential information.

The future

I hope that more providers see how great WebSockets are in agentic workflows and support them more. MCP got fixed with the 2026-07-28 spec, but it will probably take years for the old MCP programs to disappear. List endpoints most likely won't get fixed due to lock-in effects, where you have to optimize your software for a specific provider.

Providers don't have many reasons to fix some of these issues, so gateway proxies like llmleaf or Portkey AI Gateway will have to fill the gaps.

Further resources