Research engineering
GPU systems & CUDA kernels
Make the training, inference, and kernel path trustworthy enough to learn from its results.
PyTorch · veRL · vLLM · Hugging Face Accelerate · DeepSpeed · CUDA C++ · Linux
A reproduction is a systems problem
Reproducing an agent-training result means getting several layers to agree: the model, prompt construction, rollout inference, training configuration, and evaluation. A pipeline can finish without measuring the behavior the researcher intended.
At USC FORTIS Lab, I built reproducible ToolRL workflows with GRPO and veRL on the lab GPU cluster. The work included diagnosing and fixing an upstream prompt-builder evaluation defect and releasing the corrected artifacts to the lab.
Training & rollout inference
My research setup runs veRL with vLLM rollout inference on an NVIDIA A100 80GB, with explicit GPU-memory limits and model-loading configuration. During my PDD internship I implemented multi-GPU MMoE training with Hugging Face Accelerate and diagnosed CPU and I/O data loading, not GPU computation, as the primary pipeline bottleneck; earlier work there fine-tuned a BERT quality detector with DeepSpeed distributed training.
These are research and internship training and inference workflows. They support a concrete account of getting a model pipeline running and debugging it, without implying ownership of a production inference service or GPU fleet.
Two debugging lessons
Trace the prompt that evaluation actually uses
The ToolRL reproduction surfaced an upstream prompt-builder evaluation defect. I diagnosed and fixed the prompt-construction path. The practical lesson is to inspect the inputs reaching the model before attributing an evaluation result to the training method.
Separate model failures from storage failures
Network-storage failures also broke model staging mid-run. I recovered the veRL/vLLM jobs with checksum-verified local model staging, explicit memory limits, and resumable launches, so a restart could establish the identity of the files being loaded and continue rather than start over. Model-loading reliability is part of whether a research run can produce interpretable evidence.
CUDA kernel development — RMSNorm
CUDA C++ · October 2026–Present · in progress
Implementation
Fused and two-stage FP32 RMSNorm kernels in CUDA C++. Both use warp-shuffle reductions inside a warp and shared-memory reductions across the block. The fused path combines the sum-of-squares reduction and the normalization in one kernel launch; the two-stage path separates the reduction from the normalization pass so the two can be compared.
Correctness
Every kernel is validated against a CPU reference. The reference accumulates in FP64 so that the comparison is not limited by the reference's own rounding; the GPU kernels themselves compute in FP32. Fixtures cover irregular row widths and value ranges, and comparisons use fixed numerical tolerances rather than exact equality.
Measurement
A CUDA-event benchmark harness times the kernels with warmup iterations and alternating A/B runs, so that clock ramp-up and ordering effects do not favor whichever variant runs first.
No speedup figure is published here. The harness exists so that a performance claim can be measured before it is made, and the work is still in progress; the source is not yet public.
Evidence & scope
This account is supported by my current resume. It does not add an unverified throughput gain, benchmark score, kernel speedup, or production availability metric.
An exact public artifact URL for the lab runs was not established during the portfolio review, so no artifact-download link is presented. The case study stands on the specific workflows, hardware, debugging responsibilities, and kernel methodology described here.