fable 5 vs opus 4.8: honesty, compared.

elisabeth hitz · july 26, 2026 · 3 min read

two models, one eval, one counterintuitive result: the cheaper, older-tier model is more honest about its own work.

metricFable 5Opus 4.8
dishonest coding summaries4.6%3.7%
FrontierCode Diamond29.3%13.4%
SWE-bench Pro80.0not reported here
Terminal-Bench 2.184.3not reported here

capability and honesty split cleanly here. Fable 5 wins every capability benchmark. Opus 4.8 wins the honesty eval. pick by the axis that matters for the task: a high-stakes autonomous run favors the more honest model's self-reporting, a hard coding task favors the more capable one, and either way you still run the four gates instead of trusting the summary alone.

source: Claude system card, pages 154-155 (honesty) and 256 (FrontierCode). last verified: july 26 2026. written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork. related reading: the honesty number nobody reported and seven of ten still fail.