Thousand Realms' $2.7B Hit to Opus 4.6 Door: Muse Glimmer Records 8 Consecutive Losses
According to Watchful AI monitoring, Qianwen has officially open-sourced Qwen3.8-27B, and the official benchmark scores have also been released. This local model with only 27B parameters has surpassed Meta's recently released 30B model, Muse Glimmer, in all 8 directly comparable tests listed by the official source.
The difference is particularly noticeable in Agent and Coding. In Terminal-Bench 2.1, Qwen3.8-27B scored 73.0, compared to Muse Glimmer's 51.7; in SWE-bench Pro, it was 61.7 versus 51.2; and in OSWorld-Verified, the scores were 84.3 versus 65.9. Qwen3.8-27B is also ahead in general inference, documentation, and visual tests.
Even more remarkable is the comparison with Claude Opus 4.6. Out of the 19 tests where both models achieved scores, Qwen3.8-27B won in 15. It has already surpassed SWE-bench Pro, LiveCodeBench, OSWorld, AndroidWorld, and multiple visual tasks; however, it still lags behind in Terminal-Bench, GPQA, HLE, and NL2Repo.
For a 27B model that can run locally on a personal computer, the official benchmark scores have reached the level of the previous generation's closed-source flagship.