Only 5.1 Billion Parameters Activated, Ant's New Model Surpasses Trillion Flagship in 11 Tests
According to Dynamic Beating monitoring, Inclusion AI has released Ling-3.0-flash. The model has a total of 124 billion parameters, but only 5.1 billion are activated.
Based on the officially announced comparison, it outperforms Ling-2.6-1T in 11 out of 12 benchmarks, with activated parameters being only around one-twelfth of this trillion-parameter flagship model.
Architecturally, it combines Kimi Delta Attention with MLA attention layers in a 5:1 ratio to enhance long-text memory. The model natively supports a context of 256K, which can be scaled up to 1 million.
Ling-3.0-flash has been launched on the Inclusion AI API and OpenRouter, supporting both pensive and non-pensive modes. Currently, the model weights have not been made public, and the performance comparison is mainly based on Ant Group's official tests, awaiting third-party evaluation for validation.