Security engineering
agent-audit + Argus
Connect a code-level finding to the evidence a security reviewer can actually use.
Python · static analysis · taint tracking · OWASP Agentic Top 10 · PyPI · GitHub Actions
The gap between detection and a decision
LLM agent applications connect model outputs to tools, credentials, and deployment configuration. Reviewing only the model misses risks in the surrounding software. Developers need findings tied to code or configuration; security reviewers need a scoped assessment they can act on.
agent-audit and Argus address different parts of that workflow: the open-source tool supports repeatable detection, while the assessment work turns technical investigation into findings an external engineering team can fix and audit.
From source to reviewable findings
agent-audit analyzes agent code and deployment artifacts for risks such as untrusted data reaching dangerous operations, exposed credentials, excessive permissions, and poisoned MCP tool definitions. Findings link back to source locations or configuration paths and export into development workflows. It installs from PyPI and runs as a GitHub Action.
CI integration and rule-level regression tests make detection changes reviewable over time. A scanner result remains an input to an assessment; it does not by itself certify that an application is secure.
Keep the rule count honest
A rule count depends on its definition, and this project has published several that disagreed. On the current main branch (October 2, 2026), 52 unique rule IDs are defined in the built-in YAML files loaded by the rule engine; the advertised count of 66 also includes rules emitted directly by the MCP, skill, and framework scanners. An earlier published figure of 72 was corrected to 66 in the v0.20.0 release after six rules were withdrawn, and the test suite now fails if any user-facing surface disagrees with the advertised set.
A receiver-type false positive also exposed the need for discriminating tests: after correcting the misclassification, I checked the intended true-positive behavior with a control. Removing an unwanted finding is insufficient if the same change suppresses the vulnerability the rule is supposed to detect.
Argus: security assessment delivery
As Founder & Engineer of Argus Security from February to August 2026, I scoped and delivered AI-agent security assessments. The work included an OWASP-aligned audit of a production healthcare RAG application and audits of agent-payment systems, translating security risks into actionable, auditable findings for the external engineering teams that owned the code. Reports were written to answer buyer-side security questions during procurement reviews.
I was also selected for DeepSeek Harness private beta testing, where I performed security testing and submitted security issues to the team.
Client names, confidential findings, procurement documents, and beta-program details are intentionally excluded. This portfolio makes no claim that an assessment alone guaranteed a secure deployment.
Public engineering & research
The source repository and the first-author paper, Agent Audit: A Security Analysis System for LLM Agent Applications, provide public evidence of the analysis approach and developer workflow.
The paper describes a specific evaluation snapshot. Its benchmark results should not be conflated with later repository measurements or the outcome of a client assessment. The transferable work is the connection between detection engineering, measurement discipline, and security communication.