Papers
Research on how AI systems get evaluated once they ship. Every claim in these is backed by code you can run.
Well Calibrated, Wrongly Ordered
Optimising against an LLM judge buys confident fabrication, but only where the judge's ranking disagrees with the truth.
Metric Blindness in Document Information Extraction
Why character-level accuracy cannot certify structured extraction from high-stakes documents.
A third paper is in progress.