Skip to content

25 models collectively lose to humans: AI still can't understand many common sense aspects of life

Oct 8, 19:32

Beating AI News Flash: Scale Labs, the research arm of Scale AI, has partnered with Elorian to release a new benchmark called Humanity's Sixth Sense (HSS), specifically designed to test whether AI can discern implicit information from images and videos. The 522 open-ended questions cover 288 images and 234 videos. For example, whether two cars can pass between each other, why a woman suddenly slows down while chasing a bus, and who holds more sway in a given situation.

The research team tested 25 multimodal models and had 20 human participants answer the questions. Human accuracy reached 93.1%, while GPT-6 Astra, which ranked first, scored only 53.6% even with maximum reasoning effort. GPT-6.1 Sol and Claude Opus 5.5 scored 46.6% and 44.6% respectively, and the median score across all models was just 30.9%.

The researchers analyzed 8,573 model failures and found that 94% were related to missing key clues, misidentifying subjects, or being unable to infer implicit relationships in the visuals. Only about 5% were classified as logical reasoning errors. Social understanding proved especially difficult, with 21 of the 25 models performing worst on this type of question. Video questions were also generally harder than image questions.

Increasing the amount of reasoning did not necessarily help. Models consumed an average of about 4,000 reasoning tokens per question, and for some questions, the longer they thought, the worse their scores became. The research team also tried having Agents zoom in, crop, and repeatedly inspect the visuals, which improved scores somewhat, but still left a clear gap compared with humans.

Source