One Night, Two Paths: Opus 5.5 Surges to First Place, GPT-6 Halves Costs at the Same Tier
Beating AI News Flash: Artificial Analysis has completed third-party testing of several new models released last night. Claude Opus 5.5 max scored 58 on Intelligence Index v4.3, currently ranking first. GPT-6 Astra and Fable 5.1 both scored 53, while Opus 5 scored 51. Opus 5.5 achieved the highest score in 6 out of 10 evaluations.
But this top score also used more tokens. Opus 5.5 max outputs an average of approximately 119,000 tokens per task, about 60% more than Opus 5's 73,000. Thanks to a 20% API price cut and a 60% reduction in cache read pricing, the final cost per task came to $5.98, roughly on par with Opus 5's $5.86. The more practical medium tier scored 51, already matching Opus 5 max, at a cost of only $1.34 per task.
OpenAI took a different path. GPT-6 Sol and Luna's Intelligence Index performance was close to the previous GPT-5.6 generation, but prices dropped significantly. Sol max's cost per task fell from $1.99 to $1.06, while Luna's dropped from $0.18 to $0.07. In coding agent evaluations, Sol rose from 55 to 57 points, with costs also dropping by about half; Luna fell from 43 to 41 points, but cost per task was about 60% lower.
The two companies' approaches have begun to diverge: Opus 5.5 pushes the capability ceiling to new heights with more reasoning, while GPT-6 Sol and Luna mainly make existing capabilities cheaper. On Artificial Analysis's performance/cost chart, both companies' new models occupy new positions on the Pareto frontier.