Skip to content

Thinking Machines proposes a staged open-source approach: starting with the API, then fine-tuning, and finally releasing the weights

Aug 1, 11:49

According to Dynamic Beating monitoring, Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, has proposed a secure pathway for open-source models. The models are first provided to defense agencies through a restricted API, then gradually opened up for monitored usage and hosted fine-tuning. Only when the model risk is manageable and external defense systems are ready, will the team consider releasing the model weights. Releasing the weights is not a default endpoint, and the process can be paused at any stage.

The team used Inkling and Inkling-Small as examples. These two models underwent internal evaluations, testing by four independent organizations, and dedicated fine-tuning to remove the refusal-to-answer restriction. The results showed that they did not significantly surpass existing open-source models in biosecurity, cyberattack, and runaway risk. Therefore, the team concluded that releasing the model weights would not introduce substantial risk.

Thinking Machines is also researching methods to eliminate dangerous knowledge from the training data, aiming to preserve the model's general capabilities while reducing its biosecurity and cyberattack potential.

Source