Anthropic 发布 Claude Haiku 5.5:价格降 90%,多项评测超过 GPT-6 Luna 这一消息在 AI 圈引发了不小的震动。作为 Claude 系列中主打轻量、高速的 Haiku 产品线,5.5 版本的发布策略明显不同于以往——Anthropic 这次把重心放在了"性价比碾压"上。 **价格策略:降维打击** 90% 的降价幅度意味着 Haiku 5.5 的定价已经逼近开源模型的托管成本区间。对于大量依赖 API 调用的中小开发者和初创公司来说,这几乎改变了他们的技术选型逻辑。此前在成本和能力之间必须做的妥协,现在可能不再必要。 **评测表现:轻量模型的天花板被抬高** 据披露,Haiku 5.5 在多项基准测试中超过了 GPT-6 Luna。需要注意的是,Luna 是 OpenAI 面向轻量场景的对应产品线,两者定位直接对标。如果这一评测结果在第三方复现中站得住脚,说明 Anthropic 在模型蒸馏和推理效率优化上取得了实质性突破——用更小的参数量或更低的推理成本,达到了竞品旗舰轻量线的水平。 **背后的信号** 这次发布透露出几个趋势:
动察 Beating AI News Flash: Anthropic releases its new-generation lightweight model Claude Haiku 5.5, focusing on low cost, high speed, and Agent tasks. Compared with Haiku 4.5, the new model shows major improvements in coding, computer operation, and professional knowledge tasks, with some evaluations surpassing GPT-6 Luna. Anthropic says this is the company's fastest and most capable small model to date.
Official tests show that Haiku 5.5 reaches 72.4% on the OSWorld 2.1 computer operation benchmark, higher than GPT-6 Luna's 48.9%, while the previous generation Haiku 4.5 was only 15.7%. On the GDPval-AA professional work evaluation, the new model scores 1,620 points, also exceeding GPT-6 Luna's 1,437 points. Coding ability has improved significantly as well, rising from 0% in the previous generation to 39.2% on Terminal-Bench 4.0, surpassing GPT-6 Luna's 16.4%.
The price cut is even greater. For requests with no more than 100,000 input tokens, Haiku 5.5 charges $0.1 and $0.5 per million input and output tokens respectively, both 90% lower than the previous generation. Beyond that length, it charges $0.5 and $2.5 respectively. Considering changes in token consumption for actual tasks, Anthropic estimates that the average operating cost drops by about 75%.
Haiku 5.5 also adds adjustable reasoning intensity to the Haiku series for the first time, allowing developers to trade off between performance and cost. The model is now available through the Claude API, AWS, Google Cloud, and Azure, and is suitable for batch summarization, information retrieval, browser operations, and executing subtasks for large Coding Agents.