Papers
Research on how AI systems behave once they ship. Each paper has a web edition, a PDF of record, and a repository that reproduces every number.
Well Calibrated, Wrongly Ordered
Optimising against an LLM judge buys confident fabrication — but only where the judge's ranking disagrees with the truth.
Metric Blindness in Document Information Extraction
Why character-level accuracy cannot certify structured extraction from high-stakes documents.