Skip to content
Piyush ๐Ÿ‘‹
all writing

Streaming LLMs through a 20-line edge proxy.

Apr 2026 ยท 5 min read

The entire backend of my Local-LLM experiment is one route file. POST comes in from the browser, gets forwarded to an OpenAI-compatible chat-completions endpoint, and the response stream flows back untouched. Twenty lines that earn their keep.

Three lines do all the work. export const runtime = "edge" puts the proxy geographically near the user instead of in one region. The upstream body streams back with text/event-stream, Cache-Control: no-cache, and X-Accel-Buffering: no โ€” the trio that tells every proxy in the chain do not hold my tokens. And non-200 upstreams map to clean JSON errors instead of leaking raw provider responses to the client.

Why proxy at all instead of calling the model from the browser? Two reasons, both non-negotiable: the API key stays server-side, and the browser never fights CORS with a third-party AI host. The proxy is a key-hider and a CORS-eraser that happens to also pick the closest region for you.

Honest naming complaint, filed against myself: the repo is called Local-LLM and nothing here runs locally โ€” it's a streaming pass-through to a hosted endpoint. The name was aspirational; the shelf it sits on is "experiments in talking to models." A truly local version would mean Ollama or vLLM behind this same route shape โ€” which, notably, this proxy already supports, since it speaks plain OpenAI-compatible SSE either way.