Skip to content

Banbury Road lets 32 models train together, each developing its own specialty.

Oct 6, 23:18

动察 Beating AI News Flash: AI lab Banbury Road releases Kardashev-0.7, composed of 32 different models. The team trained these models together using reinforcement learning, without predefining each one's responsibilities, allowing them to develop different specialties and complementary capabilities during training. This method is called RL for Population Scaling (RLPS).

Previous RLPS experiments scaled up to at most 16 models. In an 8-model experiment, the model population after joint training scored 81.70%, higher than the 71.04% of 8 independently trained models, and also higher than the 72.65% of repeatedly sampling the same model 8 times. In the 16-model experiment, there were 22 questions that only 1 of the models answered correctly, showing that different models did indeed learn different abilities.

Kardashev-0.7 expands the scale to 32 models. Banbury Road claims it can achieve frontier-model-level performance at 0.007 to 0.02 times the inference cost and 0.03 times the memory. Currently, the API Beta still requires joining a waitlist.

However, the population scores in the previous 8- and 16-model experiments were obtained by selecting the best result from multiple model outputs using the standard answers after the fact. In actual use, there are no standard answers, and Banbury Road has not yet publicly disclosed a complete plan for how the system automatically selects the correct response.

Source