Selected work

Research prototype

CausalOps

Let models propose judgments; let deterministic code decide authority.

Python · typed contracts · state machines · Ed25519

Research-engineering prototype. Core workflow and authorization mechanisms have frozen test/review evidence; production scale and external product value remain unvalidated. A later provider run ended in an execution-integrity failure and was not evaluable.

The problem

A business action arrives with incomplete evidence, operating constraints, and assumptions that may need a human decision. A plausible model response does not tell an operator whether the proposed experiment is supported, whether its dependencies are valid, or whether anyone authorized execution.

I built CausalOps to turn that ambiguity into reviewable artifacts: typed claims, candidate designs, evidence links, and explicit rollout conditions. The goal is to make the route from suggestion to action inspectable.

Why this is technically hard

Models can mix a supported observation with an unsupported conclusion in the same response. Rejecting the entire draft loses useful material; accepting it wholesale can promote unverified text into an execution decision.

Authority also crosses several boundaries: parsing, dependency resolution, contract compilation, storage, and approval. A check in the command-line interface is insufficient if a caller can reach a lower-level write or execution path directly.

System architecture

The semantic proposal is a separate artifact from the deterministic decision. Claims and candidates retain provenance; dependency edges express what a candidate relies on. Lifecycle records preserve how an artifact moved through review, while execution approvals bind the authorized run to its configuration.

Three design decisions

01 — Keep semantic proposals separate from authority

Model-generated claims enter a typed review path. Deterministic checks decide which evidence and dependencies can support a candidate. An unresolved assumption remains a reason for human validation; it does not silently become an executable fact.

02 — Enforce integrity at the storage boundary

Content-addressed artifacts, append-only lifecycle events, and exclusive-create writes preserve prior records. Storage checks require artifact-kind and confidentiality classifications, so callers cannot bypass the boundary by avoiding the CLI.

03 — Bind approval to what will actually run

Ed25519 approvals bind execution identity, model configuration, budgets, and relevant artifact hashes. The execution path verifies the approval; it cannot sign its own authorization. Dependency-manifest checks protect against approving one configuration and running another.

Verification evidence

13 / 13Targeted mutations detected in one frozen offline gate-verification milestone.
152Focused tests passed for the standalone expert-case intake validator in a separate milestone.

The mutation campaign deliberately changed authority classification, retry behavior, manifest membership, and fixture or approval checks. Each changed implementation caused a behavioral test failure after a clean baseline. These results cover the selected gates, not every possible failure mode.

The separate intake-validator tests exercised envelope agreement, partitions, deterministic hashes, write-once sealing, and post-seal modification detection using fictional temporary data. This is a validator-specific count, not a cumulative project test total.

Both summaries come from frozen local records reviewed for this portfolio. Internal reports and private source material are not published here.

When execution evidence was incomplete

A later signed provider run recorded a semantic submission as attempted but persisted no provider response. That did not establish whether a request reached the provider. The run ended in an execution-integrity failure and was classified as not evaluable.

I kept the incident record and retired the affected confirmatory cases. The response was to stop further harness work and move toward independently acquired expert cases, rather than describe the offline gate results as a successful provider evaluation.

An authorization record is not proof of completed execution. Missing evidence must remain an explicit unknown.

The incident changed the project's validation plan and narrowed what its implementation evidence could support. It did not demonstrate external product value or expert-level decision quality.

What this project proves—and what is open

It demonstrates my approach to typed workflow design, deterministic authorization, storage integrity, and tests that challenge critical gates. It also demonstrates the engineering judgment to preserve an adverse result and stop when the evidence no longer supports progress.

Still unvalidated: production scale, provider operational readiness, expert-level decision quality, customer adoption, and external product value. The next scientific step is blind expert-case acquisition with a separately frozen evaluation protocol.

Ownership & AI-assisted engineering

I led the architecture, contract and boundary design, acceptance criteria, and review of the evidence. Implementation was AI-assisted. My ownership includes choosing the mechanisms, investigating review findings, evaluating proposed changes, and deciding which conclusions the artifacts support.

This case study is a public technical summary. It makes no claim that the system has been deployed to customers.