OpenAI has stated that the evaluation framework hampers the performance of GPT-5.6, but Opus 5 scored 2.3 times higher using the same framework.
According to Perceive Beating's monitoring, OpenAI stated that the general evaluation framework of ARC-AGI-3 did not retain GPT-5.6 Sol's previous reasoning and would also erase early records when the context was too long. As a result, the model often forgets the discovered game rules and has to reconsider them at each step.
Subsequently, OpenAI reworked the evaluation framework using its in-house Responses API, preserving historical reasoning and compressing excessively long contexts. As a result, GPT-5.6 Sol's score increased from 13.3% to 38.3%, with the output tokens reduced to about one-sixth of the original.
However, the issue lies in the fact that Opus 5 also utilizes this same general framework. It achieved 30.2% with a High reasoning tier, while GPT-5.6 Sol, even when using the higher Max tier, only reached 13.3%. The framework indeed weighed down GPT-5.6 but did not have the same detrimental effect on Opus 5.
Regardless of OpenAI's explanations, Opus 5 is significantly superior to GPT-5.6 in this globally most challenging AI evaluation.