Production AI backends · LLM serving · ML systems

Available full-time · May 2027

Haiyue Zhang
You can call me Heady.

AI Infrastructure Engineer

I build production AI backends, LLM serving layers, and ML systems—with a focus on performance, reliability, and end-to-end delivery.

M.S. ECE at USC · Available May 2027 · Los Angeles, CA

Measured results, with their scope3 systems
  1. Topify.ai · hybrid retrieval100k+ chunks · under 2 sRetrieval latency over a 100k+ chunk corpus (BGE-M3 dense search, BM25, reciprocal rank fusion, reranking). Retrieval step only; not end-to-end LLM workflow latency.
  2. PDD · MMoE serving22% lower inference timeInternal comparison against the previous feature-interaction implementation, after redesigning the interaction layer with CrossNet. Not a production-wide speedup.
  3. MaaS Gateway · failover540 / 540 requests completedControlled backend-process termination test: one of two vLLM engines killed mid-traffic, four bounded retries, dead engine excluded in about 2 s. An experiment on one GPU, not an SLA.
Each number stays attached to the measurement that produced it.
Now: CUDA C++ RMSNorm kernels · LLM post-training and inference workflows at USC FORTIS LabLos Angeles, CA

Built across the serving stack.

From an LLM inference gateway and a production AI backend to GPU training workflows and CUDA kernels.

02

Production backend

Topify.ai

Production AI backend for an end-to-end customer workflow

Founding Engineer · Nov 2025–Jan 2026

Founding Engineer for a customer-facing B2B AI platform: I owned the production Python/FastAPI backend and deployment, automating one workflow from AI-search visibility analysis and retrieval to LLM and tool execution, content generation, and WordPress/CMS publishing.

Parallel LLM queries in multi-step ReAct workflows across 30+ APIs, with timeout, rate-limit, and partial-failure controls. Hybrid retrieval over 100k+ chunks at under 2 s retrieval latency.

03

Security engineering

agent-audit + Argus

Static security analysis for LLM-agent code and MCP configurations

Open source · Argus Founder & Engineer · Feb–Aug 2026

Open-source taint and configuration analyzer for LLM-agent code and MCP configurations, shipped on PyPI and as a GitHub Action. Through Argus Security, I audited a production healthcare RAG application and agent-payment systems for external engineering teams.

Rules mapped to all ten OWASP Agentic Top 10 categories, rule-level regression tests, 1,500+ tests, 200+ GitHub stars. Rule counts are labeled by definition in the case study.

04

Research engineering

GPU systems & CUDA kernels

Training workflows, rollout inference, RMSNorm kernels & root-cause debugging

Research & kernel work · 2025–Present

Reproducible ToolRL/veRL post-training workflows with vLLM rollout inference at USC FORTIS Lab; multi-GPU MMoE training with Hugging Face Accelerate at PDD, where CPU and I/O data loading, not GPU compute, was the bottleneck.

Fixed an upstream prompt-builder evaluation defect; recovered GPU jobs from storage failures with checksum-verified model staging, memory limits, and resumable launches.

05

GPU kernels

CUDA RMSNorm kernels

Fused & two-stage FP32 RMSNorm in CUDA C++

In progress · Oct 2026–Present

Fused and two-stage FP32 RMSNorm kernels using warp-shuffle and shared-memory block reductions; the fused path does the reduction and normalization in one launch.

CPU-reference validation with FP64 accumulation, irregular-width and value-range fixtures, fixed tolerances, and a CUDA-event benchmark harness with warmup and alternating A/B timing. No speedup is claimed until measured.

From the lab to the customer.

2025–Present

USC FORTIS Lab & independent work

Research engineering · LLM post-training & inference systems
  • Built reproducible ToolRL/veRL (GRPO) workflows; fixed an upstream prompt-builder evaluation defect and released artifacts.
  • Recovered veRL/vLLM GPU jobs from storage failures with checksum-verified model staging, memory limits, and resumable launches.

Feb 2026–Aug 2026

Argus Security

Founder & Engineer · developer tooling & agent security
  • Built and shipped agent-audit, a taint/configuration analyzer for LLM-agent code and MCP configurations, on PyPI and as a GitHub Action; the open-source work continues.
  • Audited a production healthcare RAG application and agent-payment systems, translating security risks into actionable, auditable findings for external engineering teams.
  • Selected for DeepSeek Harness private beta testing; performed security testing and submitted security issues to the team.

Nov 2025–Jan 2026

Topify.ai

Founding Engineer · customer-facing B2B AI platform
  • Owned the production Python/FastAPI backend and deployment; automated an end-to-end customer workflow from AI-search visibility analysis and retrieval to LLM/tool execution, content generation, and WordPress/CMS publishing.
  • Parallelized LLM queries in multi-step ReAct workflows spanning 30+ APIs; added timeout, rate-limit, and partial-failure controls to coordinate long-running customer jobs.
  • Delivered hybrid retrieval over 100k+ chunks at under 2 s retrieval latency, combining BGE-M3 dense search, BM25, reciprocal rank fusion, and reranking.

May 2025–Aug 2025

Pinduoduo (PDD Holdings)

Algorithm Engineer Intern
  • Reduced MMoE model inference time by 22% against the prior feature-interaction implementation in internal comparisons; redesigned the interaction layer with CrossNet for review-incentive targeting, choosing CrossNet, residual connections, and Focal Loss through ablations and rejecting dynamic gating that added 30% training time for under 0.5% quality gain.
  • Reformulated EUEN's second-order FM interactions from quadratic to linear complexity in feature count, preserving algebraic equivalence, for causal-uplift models used in budget-constrained cashback allocation.
  • Implemented multi-GPU MMoE training with Hugging Face Accelerate; diagnosed CPU/I/O data loading as the primary pipeline bottleneck rather than GPU computation.
  • Also worked on LLM-based intent understanding and fine-tuned a BERT quality detector with DeepSpeed distributed training.

Make the work inspectable.

Public research on agent security, auditability, and execution governance, plus one research prototype.

Research prototype

CausalOps

Auditable AI decision and rollout workflows: typed claims, dependency checks, storage-boundary integrity, and Ed25519-signed approvals keep model proposals separate from execution authority. The case study records the mechanisms, their frozen test evidence, and an incomplete-evidence incident that stopped the provider evaluation.

Research-engineering prototype · 2026 · product value unvalidated

An engineer who follows the evidence.

I build the layer between models and the people who depend on them: serving gateways, backend workflows, training pipelines, and the kernels underneath.

Across a customer-facing backend, an open-source security analyzer, GPU research workflows, and an inference gateway, I care about how the running system behaves under load and failure, and about stating exactly what a number does and does not show.

University of Southern CaliforniaM.S. Electrical & Computer Engineering · Machine Learning & Data Science
Expected May 2027 · GPA 3.93

Shanghai Jiao Tong UniversityB.S. Electrical & Computer Engineering
2022–2025 · completed in three years

Languages
Python, C/C++, CUDA C++, SQL, Bash, JavaScript
ML / GPU
PyTorch, vLLM, veRL/GRPO, Hugging Face Accelerate, DeepSpeed; model optimization, RMSNorm, warp/block reductions
Backend / infrastructure
Linux, Docker, FastAPI, REST/SSE, Railway, Git, GitHub Actions, pytest; concurrency, load/fault testing, property-based and mutation testing

05 / Get in touch

Building serving or ML infrastructure?
Let's make it concrete.

Los Angeles, CA · Available full-time May 2027