Gemini 4 Argon's 15% hallucination rate isn't how often it's wrong

Gemini 4 Argon's 15% hallucination rate against 51% for GPT-6 Astra measures how often a model guesses when it doesn't know. Argon is also right less often.

Share
A pencil resting on a multiple choice answer sheet
Photo by Nguyen Dang Hoang Nhu on Unsplash.

On Wednesday, 30 September, Google released Gemini 4 Argon, first to a small group of cyber defenders and next, it says, to paid API customers and Google AI Ultra subscribers. The same day Artificial Analysis published its numbers. Argon ties OpenAI's GPT-6 Astra at 53 on the Intelligence Index, and it has "a 15% hallucination rate, the lowest of any model scoring 45+ on the Intelligence Index, compared with 51% for GPT-6 Astra (max)."

Read quickly, that says Astra makes things up half the time and Argon rarely does. I wanted to know what the 15% is a percentage of.

What does the hallucination rate actually measure?

Not the share of answers that are wrong. Artificial Analysis defines it as "the proportion of incorrect answers out of all non-correct responses, i.e. incorrect / (incorrect + partial answers + not attempted)."

Correct answers aren't in that sum at all. The question it answers is: when the model didn't get it right, how often did it give a wrong answer instead of saying it didn't know? A model that declines more often scores better on it, whatever its accuracy.

So is Argon more accurate than Astra?

No. On the same knowledge test, Artificial Analysis reports Argon's accuracy at 50% and Astra's at 63%. Their article says Argon "is much more likely to acknowledge when it does not know an answer rather than guess incorrectly."

My first reaction was that this makes the 15% figure misleading. Then I worked it through per hundred questions, using their numbers.

Argon gets 50 right. Of the other 50, 15% are wrong answers, so about 7 or 8 are wrong and the rest are blanks or partial answers. Astra gets 63 right. Of the other 37, 51% are wrong, so about 19 are wrong.

Their combined score, the Omniscience Index, is right answers minus wrong answers. That gives Argon about 42 and Astra about 44. Artificial Analysis reports 42 and 43, so the arithmetic holds up to rounding. (The tie at 53 is a different measure, the Intelligence Index.)

So the 15% isn't misleading. It's answering a narrower question than the headline suggests. Argon is wrong less than half as often as Astra and right noticeably less often. Both are true, and on the combined score they're level.

Per 100 questions: Gemini 4 Argon 50 right, 7.5 wrong, 42.5 blank, hallucination rate 15%. GPT-6 Astra 63 right, 18.9 wrong, 18.1 blank, rate 51%.
Per 100 questions on AA-Omniscience. Rates and accuracy from Artificial Analysis, 30 Sep 2026; the wrong and blank counts are my arithmetic.

Where have we seen this before?

On the SAT. Until March 2016 it took a quarter of a point off for each wrong answer, so a student who didn't know was better off leaving it blank. When the College Board redesigned the test it moved to "rights-only scoring" and told students "to select the best answer to every question." The Princeton Review's advice now is to "pick your favorite letter" for blind guesses. Same students, same knowledge, different behaviour, because the scoring changed what a guess cost.

OpenAI made the same point about models last year: "Leaving it blank guarantees a zero," so a test that only counts right answers teaches models to guess. Its proposed fix was to "penalize confident errors more than you penalize uncertainty." Artificial Analysis built its index that way. Argon behaves like a student sitting the negative-marking version of the test.

Which one should you use?

It depends on what a wrong answer costs you compared with a blank one.

Where an error is expensive and a human picks up the blanks, such as a compliance answer, a medical summary or a number going into a report, the model that says "I don't know" more often is worth the lower accuracy. Where a blank is the expensive part, such as a first draft or a brainstorm someone will check anyway, the model that answers more often gives you more to work with.

What I'm confident of, and what I'm not

The definitions and the 15%, 51%, 50% and 63% figures are established from Artificial Analysis. The per-question split is my own arithmetic from their numbers, and it reproduces their index for Argon. Which model suits which job is a judgement, and Argon isn't generally available yet, so I haven't tested it.

The claim, in one sentence: a hallucination rate tells you how a model behaves when it doesn't know, not how often it's right, so read it next to accuracy and choose by what a wrong answer costs you compared with a blank one.

It's the same kind of question as the collusion piece: before reacting to the number, find out what it's a percentage of.

Sources

Artificial Analysis, "Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved", 30 September 2026. https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs

Artificial Analysis, AA-Omniscience methodology. https://artificialanalysis.ai/evaluations/omniscience

AA-Omniscience paper, arXiv 2511.13029. https://arxiv.org/html/2511.13029v1

Koray Kavukcuoglu, Google, Gemini 4 Argon, 30 September 2026. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

Matthias Bastian, The Decoder, 1 October 2026. https://the-decoder.com/google-gemini-4-argon-closes-the-gap-with-openai-and-anthropic-but-doesnt-take-a-clear-lead/

OpenAI, "Why language models hallucinate", 5 September 2025. https://openai.com/index/why-language-models-hallucinate/

College Board, redesigned SAT announcement, 5 March 2014. https://newsroom.collegeboard.org/college-board-announces-bold-plans-expand-access-opportunity-redesign-sat

The Princeton Review, "Should you guess on the SAT and ACT?" https://www.princetonreview.com/college-advice/should-you-guess-on-the-sat-and-act