Selected work

LLM serving & traffic control

MaaS Gateway

Decide which replica serves a request, how long it may take, and what the client sees when a replica dies.

Python · asyncio · OpenAI-compatible HTTP and SSE · vLLM · pytest

Independent engineering project, September 2026. Validated end to end against real vLLM (Qwen2.5-0.5B-Instruct, vLLM 0.29.0) on one NVIDIA L4; every number below comes from a controlled experiment on that topology. The gateway has not served production customers and makes no capacity, uptime, or availability claim.

An LLM fleet is not a normal web backend

Requests to an inference fleet run for seconds rather than milliseconds, so a requests-per-second limit says little about actual load. A replica can answer its health probe while failing every completion, because the HTTP server is fine and the model never loaded. And once a streamed token is on the wire, the response status is committed and nothing can be retried transparently.

The gateway is where those problems are handled. It sits between OpenAI-compatible clients and a set of vLLM replicas and owns four decisions: which replica serves a request, how long the request may take, what happens when a replica dies mid-request, and what the client is told when it does.

Architecture, as built

The topology is literal: one gateway process (Python, asyncio), two vLLM processes, one GPU. No Kubernetes, Envoy, Redis, or service mesh was used, so none appear in the design. All control state is per-process, and the repository's milestone documents record what that costs.

Solid lines in the diagram are the request path; the control state is consulted by the router and the upstream call but never carries a request. Response bytes are relayed unmodified, so vLLM-specific fields the mock backend never emits reach the client intact.

The request path

  1. Rate limiter. Per-API-key token bucket: capacity is the burst, refill is the sustained rate. Over quota returns 429 with Retry-After.
  2. Admission control. Process-wide in-flight ceiling. When full it returns 503 overloaded immediately; it never queues.
  3. Validate and deadline. The client body is checked but forwarded untouched, and an end-to-end budget starts counting down.
  4. Route. Pool split (canary) first, then replica by smooth weighted round robin, filtered by health state and circuit-breaker state.
  5. Upstream call. Buffered, or SSE with the upstream headers awaited before the response status is committed, so a dead backend yields a clean 502 instead of an empty 200.
  6. Relay. Bytes pass through unmodified; time to first token is observed on the first relayed byte.

Four design decisions

01 — Rate limiting bounds arrival; admission control bounds work in flight

A rate limit cannot see concurrency when request durations vary by 100×. Ten requests per second is comfortable at 200 ms per request and fatal at 40 s per request, and only an in-flight ceiling can tell the difference. The two gates do different jobs, so both exist.

02 — Health checks gate eligibility; the circuit breaker gates attempts

Health probes are out-of-band and use hysteresis. The breaker is driven by in-band request outcomes and moves between closed, open, and half-open per backend. They are deliberately separate because each catches what the other cannot: the breaker is fast because it observes real traffic, and the probe catches a backend that is failing without being asked.

03 — One deadline across attempts, forwarded as remaining time

Each request carries one absolute budget. It is forwarded upstream as the remaining time, so a retry can never extend the client's wait. Per-hop timeouts would compose by addition; a single propagated deadline bounds real GPU work.

04 — Retry only before anything is committed to the client

A retry needs four conditions at once: a retryable error, nothing committed to the client, a deadline that still permits it, and credit in the retry budget. The budget caps retries as a fraction of traffic rather than a multiple of it. For streaming, the upstream headers are awaited before the status is committed; a failure after commit is signalled in band with an error event, partial: true, and no [DONE] marker, never replayed.

Controlled validation results

540 / 540Client requests completed when one of two vLLM engines was killed mid-traffic (experiment E3, non-streaming requests).
551Automated tests over real sockets, run without a GPU; the count reported by the repository at its validated tag.

Backend failure (E3). One engine was terminated while traffic was flowing. All 540 requests completed with zero client-visible errors. Four upstream failures were absorbed by four bounded retries, and the retry budget was never a constraint. The circuit breaker removed the dead engine from routing in about 2 s; the out-of-band health check agreed at about 10 s. The visible cost of masking the failure was p95 latency rising from about 345 ms to 709 ms, which an average would have hidden.

Burst, each mechanism isolated (E2). On an identical 3 → 40 → 3 requests-per-second profile, admission control alone admitted 470 of 800 requests (330 rejected with 503) and rate limiting alone admitted 219 of 800 (581 rejected with 429). Both held p50 latency for admitted traffic within about 5–6% and recovered as soon as the burst ended. Admission passed more useful work because it bounds concurrency and therefore adapts to service time.

Canary (E4). A configured 90/10 pool split was observed as 89.90% / 10.10% over 396 requests. When the canary engine was killed, traffic fell back to 100% stable with zero client-visible errors; the pool's share was reabsorbed rather than black-holed.

Steady load (E1). Completed throughput scaled from 3.04 to 79.58 requests per second between concurrency 1 and 32 with nothing rejected or dropped. The system was still scaling near-linearly at concurrency 32, so no saturation point was found and no capacity figure follows from this data.

Topology for every measurement: client → gateway → routing and traffic control → vLLM → GPU → generated tokens → client, with Qwen2.5-0.5B-Instruct in bfloat16 on vLLM 0.29.0, 64 output tokens, and both engines sharing a single NVIDIA L4. Raw per-request records and an evidence index that maps each number to its artifact are committed in the repository.

What went wrong in the harness

Three experiment defects were caught, each of which had produced green but invalid results. The invalid run is kept in the repository rather than deleted.

  1. The kill that killed nothing. vllm serve forks its API server; killing the launcher process orphaned it, the port stayed bound, and the "killed" backend kept serving. The first failover run reported 540/540 for the wrong reason.
  2. A backend that never started. Port 8001 was the host's nginx, which answered /health with 200, so a vLLM process that had died at startup looked permanently healthy.
  3. Orphaned engine processes held GPU memory after teardown and had to be cleaned up explicitly.
A green benchmark is not evidence unless the experiment itself is valid.

What this proves—and what it does not

It demonstrates a working OpenAI-compatible control layer for LLM serving, the reasoning behind each traffic-control mechanism, and a measurement discipline that separates measured facts from interpretation and from what was not tested.

Limits of the evidence: one NVIDIA L4 and one small 0.5B model, so nothing transfers to other GPUs or to 7B+ models; a fixed workload of short prompts and 64 output tokens; no saturation point and therefore no capacity limit; no multi-GPU failover, because both engines shared one card; no soak test beyond 90 s; no isolation of gateway overhead from backend time; no cost or multi-tenant fairness benchmark. It is an engineering project, not a production deployment.

Ownership

I designed the control layer, implemented it in four milestones (core gateway, traffic control, real vLLM integration, performance and failure measurement), built the deterministic mock backend with failure injection, and ran the validation and benchmark workflows. The repository contains the source, the tests, the committed artifacts, and the design review that argues each mechanism.