Writing

Writing

Papers, field notes from building evaluation infrastructure, and technical write-ups.

Papers
The Potentiality Field→

What actually varies when an agentic system's evaluated behavior shifts, and why attributing that variance takes more than one axis.

Framing paper · self-hosted
Not identified: When holistic LLM evaluation cannot support an interaction claim

A preregistered, held-out test of whether defect-specific discrimination survives a second co-occurring defect. The instrument returned “not identified” and preserved why.

SoonUnder submission · Agent Evaluation Science, Fall 2026
Field Notes
Every gate agreed. And every gate was wrong→

Four automated checks came back green on a result whose capability ordering was backwards. The only thing that caught it was an operator who did not believe the number.

2026 · on coded gates and structural blindness
Seven confident, wrong substitutions from clean documentation

A capable model built structurally correct integration from the docs alone, and got the meaning wrong in seven specific ways. What the documentation layer could not transmit.

SoonDraft
The authority engine oversteps its own authority

An incident from building the authority-governance engine, caught in its own dogfood ledger.

SoonDraft · disclosure check pending
Open Source
Blacklight↗

A contract-driven black-box evaluation harness: test an implementation against a contract it did not supply, with an exact oracle, an independently implemented reference, and deliberately broken targets. Python, standard library only.

github.com/inspectorSlap/blacklight · early prototype
Study-Scout↗

Adversarial prior-art search: freeze the claim before searching, keep the full retrieval trail, route real matches to a human. There is deliberately no NOVEL verdict in its vocabulary.

github.com/inspectorSlap/Study-Scout
Technical
Whitepapers

Longer technical write-ups will live here, or link back to wherever they are published.

SoonComing