ARC-AGI-3 Evaluation Body Considers Rule Change: Official Benchmark or Vendor Provided Memory Allowed
According to Dynamic AI Beating monitoring, the ARC-AGI-3 evaluation organization ARC Prize had previously requested all models to use the same simple invocation method to prevent vendors from inflating scores through customized frameworks. This makes it easier to see the actual capabilities of the models.
In response to the controversy over OpenAI's test, the official stance has started to relax. ARC Prize acknowledges that retaining historical reasoning, compressing lengthy contexts can indeed improve a model's performance on long tasks, and has begun researching how to incorporate such features into formal evaluations.
The boundary set by ARC-AGI founder François Chollet is: frameworks developed specifically for ARC are still not allowed; common configurations that all API users can access may be considered. He believes that the settings used by each party and how much money was spent should be made public.
Once vendors' built-in memory and context management are allowed, the comparison by ARC-AGI-3 will no longer be just about the models, but will also include the engineering capabilities behind each API.