Skip to content

Trillion parameters can't save dirty data either; domestic large models begin reworking pre-training.

Sep 23, 12:57

Beating AI News Flash, Leiphone reports that over the past six months or so, multiple domestic large model manufacturers have begun redoing pre-training, with issues all pointing to data.

One manufacturer once spent nearly a year training a trillion-parameter model, only for its actual performance to be surpassed by a small model with about 1% of the parameter count. The team ultimately discovered that the problem lay in the training data. Web junk, duplicate corpora, and low-quality annotations mixed into the training set, and even a larger model cannot handle it. This case comes from an anonymous source, and the specific company has not yet been disclosed.

Tencent Hunyuan has already publicly redone a round. Starting in February this year, the team rebuilt the pre-training and reinforcement learning infrastructure. Previously, media disclosed that the old Hunyuan had issues such as benchmark data mixed into the training set and chaotic annotation rules. The new team redefined data standards, cleaned the original corpora, and Hy3 also listed data quality and diversity as key improvement areas.

Alibaba and Baidu are also stepping up data governance. Qwen3 used the Qwen2.5 series models to clean documents and also massively supplemented mathematical and code synthetic data. Wenxin 4.5 added deduplication, low-quality filtering, data mapping, and manual review. The two companies did not publicly say they had reworked due to dirty data, but the new generation of models has put more effort into data processing.

Source