172
Points
112
Comments
aarondong
Author

Top Comments

andy99Jul 24
#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.
chmod775Jul 24
The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot.

At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.

didibusJul 25
What's interesting is this:

The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).

Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.

That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?

firasdJul 24
Very interesting that one of the components is "AA-Omniscience Index"

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.

This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)

I've thought for a while that Gemini 3.x has 'big model smell'

aarondongJul 24
Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...
nu11ptrJul 25
I don't have a horse in this race, but to me this makes GPT-5.6 Sol Max look better. It is about half the cost for nearly the exact same performance. It just goes to show how expensive Fable really is when Opus 5 is still this expensive relative to GPT 5.6.
kristopolousJul 25
I posted this before but I have a really simple shell tool to keep up with these charts over at

https://github.com/day50-dev/aa-eval-email

This also works

$ curl day50.dev/art-analysis.sh | bash

Artificial analysis knows about my tool and I'm working with them on getting their API improved.

zorminoJul 24
I'd be curious to see the results, especially with some models having 1.5m and 2m context sizes, if the first 75% of the context was filled with unrelated info.
Visit the Original Link

Read the full content on artificialanalysis.ai

Source
artificialanalysis.ai
Author
aarondong
Posted
July 24, 2026 at 07:45 PM


More Top Stories

anthropic.com Jul 24
Claude Opus 5
1365739 commentsby alvis
Details
dbos.dev Jul 24
Postgres LISTEN/NOTIFY actually scales
23340 commentsby KraftyOne
Details
news.st-andrews.ac.uk Jul 24
Sperm Whales blow bubbles to achieve restful, vertical sleep
453 commentsby hhs
Details
arstechnica.com Jul 20
India's first privately-developed rocket reaches orbit on debut launch
525149 commentsby sohkamyung
Details
hhh.hn Jul 24
My security camera shipped a GitHub admin token in its login page
532182 commentsby hhh
Details
globaloilnetwork.staffinganalytics.io Jul 23
Show HN: I simulated closing the Strait of Hormuz on real oil trade data
11761 commentsby eliotho
Details
👋 Need help with code?