Skip to content

AI as Analyst Not Yet Reliable Enough: Repeating Each Question 5 Times, Claude Opus 5 Only Gets 54% of Questions Correct

Aug 12, 15:13

According to TrendForce Beating Monitor, Artificial Analysis has launched a new AI analysis capability test called AA-AnalystAgent. This test allows models to process real tables and documents, perform data analytics, trend analysis, financial modeling, and other tasks. The test consists of 80 questions, with each question requiring the model to independently attempt 5 times, and only achieving 5 consecutive correct answers is considered a pass.

The top scorer was Claude Opus 5, but even they could only consecutively answer 54% of the questions correctly 5 times. GPT-5.5 ranked second with a pass rate of 50%, followed by Claude Fable 5 at 49%. Among open-weight models, Kimi K3 performed the best at 39%.

Looking at the average correctness rate for each attempt, GPT-5.5 actually had the highest score at 66%. However, when requiring 5 consecutive correct answers for the same question, its pass rate dropped to 50%. Despite having slightly lower single-attempt correctness rates, Claude Opus 5 demonstrated more stability, securing the first position in the final ranking.

The most common issue encountered by AI models is not miscalculations, but misinterpreting the questions from the start. Artificial Analysis analyzed 1567 instances of failed records, revealing that 57% of them involved the model incorrectly interpreting a question early on and then persisting in that wrong direction.

Source