Benchmarks reward AI models for hallucinations. What if we called their bluff?
By Chip Holmes ·
Why does a wrong answer cost no more than a right answer earns?

If hallucinations make humans trust AI less, why aren't the benchmarks highlighting them instead of another coding exam?
What if we benchmaxxed trust?
What if we treated intentional and unintentional bluffing as a CI failure? If it bluffs, the build fails!
Would that at least point us in the right direction on "model alignment"?
Are frontier models being post-trained to bluff?
I propose the Hallucination-Weighted Score.
A right answer earns 2 points. "I don't know" earns 1 point. A wrong answer loses 3 points. Answering only beats "I don't know" when a model is right more than four times for every time it's wrong. This way, a "smarter" model can't just bluff its way to the top of the stack.
We created this secondary analysis from the results of Artificial Analysis's Omniscience test, which has 6,000 knowledge questions. Artificial Analysis (AA) scores it +1 for right, -1 for wrong, and 0 for "I don't know." And not hallucinating is only 5% of AA's Intelligence Index, a metric many people quote to claim a model is "smarter."
I'm what the AI insiders call a knowledge worker. I am a real estate appraiser, and I don't have a PhD in machine learning. In my work, a wrong answer costs me something. I am the one who loses when I rely on a hallucinated analysis that the AI stated confidently. I'm a regular AI user and small business owner, trying to make sense of all this model-scoring alchemy.




Claude Sonnet 5.5 dropped yesterday, so I scored it with the Hallucination-Weighted Score. At Max, it falls just short of Claude Opus 5.5 at Medium, and it costs nearly six times as much per Intelligence Index task.

So before I watch another YouTube video or read another Twitter post telling me how great Sonnet 5.5 is and how wonderful Anthropic is for releasing it, tell me how it helps me.
Knowing more is one thing. Knowing when to fold is another.
AA's hallucination rate measures wrong answers as a share of all the questions a model didn't get right. That includes wrong answers, partial answers, and questions it didn't attempt. Let's call that the bluff rate.
If more intelligence reliably meant less bluffing, the dots would run downhill. They barely do. That's why "smarter" isn't a good enough metric anymore. You can't just say "Pareto frontier" or "this one is smarter, so it's better."

Look at two models that are equally "smart" on paper. GPT-6 Luna scores about the same as MiniMax-M3 on AA's Intelligence Index. But when they don't know the answer, they play it completely differently. Luna bluffs 85% of the time. MiniMax bluffs 18% of the time. Which model do you want to use? I want the one that, most of the time, folds and says, "I don't know."
Why would I choose Luna for knowledge work? EVER!
And why on God's green Earth does AA use GPT-5.6 Luna at Medium to grade AA-Omniscience? Who's checking the grader?
Pre-training sets how much a model knows, and that sets how many hands it can win. In post-training, like RLHF, the lab decides what it does with the rest. And the benchmarks reward bluffing. Sooooooo?
So maybe frontier models are being post-trained to bluff. If we change the scoring and the labs change the training, do we get more AI "trust"?
When a right answer earns one point, a wrong answer loses one, and folding earns nothing, the pot odds say any guess better than a coin flip is worth making. Our score makes a wrong answer cost three and gives credit for "I don't know." Now a model has to be more than 80% sure before it bets. I want models to have a reason to fold.
I found a similar argument in Why Language Models Hallucinate, by Kalai, Nachum, Vempala and Zhang of OpenAI and Georgia Tech. Their confidence-target rule, set at 80%, gives the same answer-or-abstain threshold as my four-to-one rule. The paper also finds that most popular benchmarks give no credit for "I don't know," so guessing pays.





At Max, Claude Fable 5.1 gets more answers right than any model I scored, and only "dumber" models bluff more. Fable can go all in more because everyone knows he is "smarter," and humans and agents think he is probably right. (Fable has never once stopped to ask for directions.)
The AI industry calls these hallucinations strikes. They're not even close. These models damage knowledge workers like me. We need models that hallucinate and confabulate less, if at all. Our trust is built on correct numbers, not hallucinated numbers confidently stated as facts.
Until the labs train models for people who can't just shrug off a hallucination, I'll keep scoring every new model that drops.
See the methodology and sources below.
AI Use Statement: Of course I used AI to help me analyze this data and compose the secondary analysis. I'm not even sure how I would do it by hand at this point. Graph paper? Some of the named and unnamed agents that helped me were Lark, Plumb, Astra, Claude Opus 5.5 Max, Claude Sonnet 5.5, Claude Fable 5.1, Claude Opus 4.8 and GPT-6 Astra Low.
Sources and chart notes
Source: Artificial Analysis’s AA-Omniscience results, Intelligence Index methodology, and model pages. Data accessed September 28, 2026. These charts are my secondary analysis of AA’s published results.
The Hallucination-Weighted Score gives 2 points for a right answer, 1 for no answer, -3 for a wrong answer, and 0 for a partial answer. It is points per 100 questions, not a percentage or a confidence measure.
Right, partial, wrong and no answer are shares of all 6,000 questions. No answer means AA’s “not attempted,” including refusals and blanks. Rebuilding the answer shares assumes partial answers count as attempts.
Bluff rate is AA’s hallucination rate: wrong answers divided by everything not answered correctly, including wrong, partial and not-attempted responses.
The scatterplot shows the eight named models. Lines connect available reasoning settings in effort order. GPT and Claude lines run from Low through Max; Grok 4.7 connects High and Extra high. MiniMax-M3 has one point. Artificial Analysis’s Cost per Intelligence Index Task is its weighted average cost across the broader Intelligence Index workload at API list prices. It is not the cost of one AA-Omniscience question or subscription usage. The 5% non-hallucination weight comes from Intelligence Index v4.3.2.
AA uses GPT-5.6 Luna at Medium as its Omniscience grader. That model’s row in these charts shows its own test-answer performance, not the accuracy of its grading.
