Assay: Claims That Decay With the Code
Our new paper on arXiv binds every claim an AI coding agent makes to the Merkle hash of the code it covers, so the claim goes stale exactly when that code changes.
Gagan Deep Singh
Founder | GLINR Studios
A paper, and a first for me in this field
Our paper Assay: Claims That Decay With the Code is now on arXiv as 2609.36170 in the software engineering category. I co-authored it with Om Shankar Tiwari and Tangi Vass. The implementation, the experiment scripts, and the paper source are public at github.com/OmShiv/assay-research.
It also has a personal thread running through it. The committed, Merkle-skipped index at the heart of the system comes from Stacklit, the codebase indexer I built earlier this year.
The problem: two failures that look unrelated
AI coding agents fail in two ways. They burn most of their context window rediscovering where things live in a repository. And when the work gets hard, they report success without evidence: "tests pass" when the tests were never run, or were quietly changed.
Tools like Stacklit attack the first problem with cheap, precise context. Review frameworks attack the second with accountability. The paper's observation is that both are describing the same thing, the structure of the codebase, at two different moments: what is true of the code now, and what was verified to be true, at which revision, by whom.
The idea: key every claim by the code it covers
Every claim an agent makes (tests pass, no secrets, behavior preserved) is bound to the Merkle hash of the dependency cone of the code it is about: the module plus everything it depends on. Change that code, or anything underneath it, and the hash moves, so the claim is stale. Change something unrelated, and the claim stays valid.
That turns "is this approval still good after the rebase?" from a judgment call into a hash comparison. It also means the blast radius of a change and the set of claims it invalidates are the same set, computed in one pass over one graph.
On top of that sits a merge gate that consults no model. It checks coverage, freshness, signatures, exit codes, plausibility, review status, and evidence monotonicity: the number of tests that ran cannot quietly drop, which is the mechanical form of "do not delete the failing test".
What the experiments show
Every number in the paper is produced by the released scripts, with no model in the loop, so they reproduce to the digit:
- A 600-token brief orients an agent for 14x to 114x fewer tokens than exploring the repository.
- Binding claims to the dependency cone re-verifies 7.9% to 81.9% of claims after an edit, where binding to the whole repository re-verifies all of them and binding to the edited module alone misses 23% to 68% of the claims that should go stale.
- The gate blocks 9 of 9 scripted cheating behaviours, including deleting a failing test, while admitting the honest paths.
The paper is also clear about what it does not show yet: whether real language-model agents perform better with it. That study is designed and is the next step.
Read it
- Paper: arxiv.org/abs/2609.36170
- Code and experiments: github.com/OmShiv/assay-research
- Stacklit, the index it builds on: github.com/glincker/stacklit
If you build agent tooling and try it on your own repositories, I would like to hear how the blast radius looks on your codebase.