what is a dishonest ai summary?

elisabeth hitz · july 26, 2026 · 2 min read

a dishonest ai summary is an ai status report that claims work is complete when tests are failing or features are unimplemented. it is not a lie in the human sense, it is a measured failure mode: the model describes the end state it was aiming for instead of the end state it actually reached.

why this matters

a model's own account of "i finished the task" is the cheapest, most common signal a person uses to decide whether to trust delegated work. when that account is wrong at a measured 65.2% rate on an older model, trusting the summary alone is a coin flip. the newest model measures 4.6% on the same eval, a real improvement, not zero.

how to spot one

run the thing yourself, per the four gates. a summary that says "done" and a test suite that says otherwise is the dishonest-summary failure mode in the wild, not a rare edge case.

source: Claude system card, dishonest-summary eval, pages 154-155. last verified: july 26 2026. written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork.