LLM serving & traffic control
MaaS Gateway
OpenAI-compatible inference gateway for vLLM replicas
An LLM serving fleet is not a normal web backend: requests run for seconds, a replica can pass its health probe while failing every completion, and a streamed token already on the wire cannot be retried. I built the control layer that decides which replica serves a request, how long it may take, what happens when a replica dies mid-request, and what the client is told.
Weighted and canary routing, one end-to-end deadline across attempts, per-client rate limits, concurrency admission control, circuit breakers with separate health checks, bounded retry, and byte-for-byte SSE passthrough. Validated end to end against real vLLM on one NVIDIA L4, with 551 tests and reproducible load and failure-injection workflows.
- Weighted / canary routing
- End-to-end deadline
- Admission control
- Breaker + health checks
- Bounded retry
- SSE passthrough
Engineering project · validated against real vLLM · Sep 2026
Client · OpenAI-compatible SDK or curl
- 01Rate limiterPer-API-key token bucket → 429
- 02Admission controlIn-flight ceiling, never queues → 503
- 03Validate + deadlineOne end-to-end budget per request
- 04Weighted / canary routerPool split, then replica (smooth WRR), filtered by health and breaker state
- 05Upstream callBuffered, or SSE opened headers-first; bytes relayed unmodified
Health probesCircuit breakerRetry budgetScheduler credit
E3 · ONE ENGINE KILLED MID-TRAFFIC
Breaker excluded it in ≈2 s; health probe agreed at ≈10 s. p95 345 → 709 ms: the measured cost of masking a failure.