What actually varies when an agentic system's evaluated behavior shifts, and why attributing that variance takes more than one axis.
Writing
Papers, field notes from building evaluation infrastructure, and technical write-ups.
A preregistered, held-out test of whether defect-specific discrimination survives a second co-occurring defect. The instrument returned “not identified” and preserved why.
Four automated checks came back green on a result whose capability ordering was backwards. The only thing that caught it was an operator who did not believe the number.
A capable model built structurally correct integration from the docs alone, and got the meaning wrong in seven specific ways. What the documentation layer could not transmit.
An incident from building the authority-governance engine, caught in its own dogfood ledger.
A contract-driven black-box evaluation harness: test an implementation against a contract it did not supply, with an exact oracle, an independently implemented reference, and deliberately broken targets. Python, standard library only.
Adversarial prior-art search: freeze the claim before searching, keep the full retrieval trail, route real matches to a human. There is deliberately no NOVEL verdict in its vocabulary.
Longer technical write-ups will live here, or link back to wherever they are published.