[Home](https://codefionn.eu/) · [Projects](https://codefionn.eu/projects/) · [About](https://codefionn.eu/about/) · [GitHub](https://github.com/codefionn)

---

# The bad design of LLM APIs

> What building an LLM gateway taught me about the state of AI

*Published on 2026-10-04 · [View as HTML](https://codefionn.eu/the-bad-design-of-llm-apis/) · [Auf Deutsch lesen](https://codefionn.eu/das-schlechte-design-von-llm-apis/)*

---


I'm currently developing many different AI-enabled applications and
specifically created [llmleaf](https://github.com/codefionn/llmleaf), an AI
gateway proxy that unifies and enriches all kinds of LLM API providers.

During this I had some issues with these interfaces and sometimes with their
documentation. This article covers the issues that the AI providers somehow
ignore.

## <abbr title="Server-sent events">SSE</abbr>

How would you implement a bidirectional communication protocol? Single
requests? Custom binary protocol? Maybe WebSockets?

The answer chosen by the industry was of course
<abbr title="Server-sent events">SSE</abbr>. It's a web standard for streaming
events from the server to the client, which is good. Crucially, it doesn't
stream the other way around, which is bad. Maybe this was a good decision in the
early ChatGPT days, but the post-Opus 4.8 era needs something different.

This has two issues. When the client has to send messages back to the server,
which in the age of agentic development includes a lot of tool call results, a
completely new request has to be made to the server. This is made worse by
prompt caching being a requirement for cost reasons and by thinking being hidden
or encrypted.

When the new request is sent to answer a tool call, the API has to infer the
last server response from the context, probably via hashing. Once it's
identified, caching works properly and the user hopefully pays less for tokens
that were already processed. The implementor has to ensure that the
messages sent are detected as the same ones sent previously.

There are even special API features that try to work around this, e.g.
Anthropic's `cache_control` or OpenAI's `previous_response_id`.

All this reconnecting has a latency cost, which would not be paid by proper
implementations. If you want to experience this, just use Codex with the
OpenAI subscription. It already uses WebSockets.

<div class="images images-column">
    <figure class="image">
        <img src="/static/img/llm-api-sse-turns-en.svg" alt="The agent sends the full history to the LLM API, gets an SSE stream until the next tool call, and sends the tool result as a new request with the full history again" width="674" height="220" />
        <figcaption>Over SSE, every tool call means a new request with the whole history</figcaption>
    </figure>
    <figure class="image">
        <img src="/static/img/llm-api-websocket-turns-en.svg" alt="The agent keeps one WebSocket open, messages flow both ways over it and only new tool results are sent, while the LLM API keeps the state in memory" width="600" height="225" />
        <figcaption>Over a WebSocket, the connection stays open and only new input is sent</figcaption>
    </figure>
</div>

## <abbr title="Model Context Protocol">MCP</abbr>s

The previous <abbr title="Model Context Protocol">MCP</abbr> specs were so bad
that the 2026-07-28 spec had to fix them. FYI: the previous specs required
stateful connections, which made stuff like letting your customers' agents read
your documentation over MCP way harder and more resource intensive than it
needed to be. The 2026-07-28 spec finally makes MCP stateless.

Another issue is that the previous specs made every agent session spawn its own
MCP connections, which made 3 or more agents very costly with some larger MCP
servers.

## The list endpoint

AI providers have endpoints for listing supported models. The issue: they often
don't provide necessary information. In the worst case, you will just get model
ids, names and release dates.

<div class="images">
    <figure class="image">
        <img src="/static/img/llm-api-model-list.svg" alt="curl request to the OpenAI models endpoint for gpt-6-astra, returning only id, object, created, owned_by and shutdown_date" width="520" height="306" />
        <figcaption>OpenAI's model list: no context window, no pricing, no capabilities</figcaption>
    </figure>
</div>

But there's a lot of information required for implementing an AI-enabled
application: context window length, max output tokens, pricing, supported
parameters and capabilities like thinking, input and output modalities (e.g.
can receive an image), automatic context window compression (summarization).

This is why providers like OpenRouter are an easy solution for applications
that support "any" provider, because they provide essential information.

## The future

I hope that more providers see how great WebSockets are in agentic workflows
and support them more. MCP got fixed with the 2026-07-28 spec, but it will
probably take years for the old MCP programs to disappear. List endpoints most
likely won't get fixed due to lock-in effects, where you have to optimize your
software for a specific provider.

Providers don't have many reasons to fix some of these issues, so gateway
proxies like [llmleaf](https://github.com/codefionn/llmleaf) or
[Portkey AI Gateway](https://github.com/Portkey-AI/gateway) will have to fill
the gaps.

## Further resources

- [Using server-sent events](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events)
- [Speeding up agentic workflows with WebSockets in the Responses API](https://openai.com/index/speeding-up-agentic-workflows-with-websockets/)
- [OpenAI WebSocket mode](https://developers.openai.com/api/docs/guides/websocket-mode)
- [The 2026-07-28 MCP specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/)

---

[Impressum](https://codefionn.eu/impressum/) · [Datenschutzerklärung](https://codefionn.eu/datenschutz/) · [Mastodon](https://c.im/@codefionn)

© Copyright 2022-2026 Fionn Langhans
