Skip to content

Even a smart large-scale model needs practice: Former OpenAI Researcher Starts a Business Selling Data

Jul 30, 18:45

According to Dynamic Beating monitoring, Andrew Ho, former OpenAI researcher and co-author of GeneBench-Pro, has resigned to start his own venture. The new company will create high-quality reinforcement learning data for large models (training tasks that involve allowing the model to practice and then grading its performance). The initial products will focus on biology and statistical reasoning, and the company's name has not been disclosed yet.

Ho believes that large models are still heavily biased. Even in the vast field of programming, models tend to memorize tasks and modify code, requiring manual intervention for actual deployment.

Furthermore, there is a lack of appropriate training tasks for real-world applications. These tasks are domain-specific and it is challenging to automatically evaluate whether the model's output is correct.

His approach involves generating research-like data and questions with known answers. This allows the model to explore freely, while the training system can accurately assess correctness.

GeneBench-Pro has adopted this methodology. The evaluation includes 129 computational biology tasks, with GPT-5.6 Sol Pro achieving a pass rate of only 31.5%.

The new company will also create training data for experiment images such as petri dishes and Western blots, and later expand into chemistry, materials, healthcare, and office tasks.

Ho anticipates that the future AI frontier labs will invest over $100 billion in precise training data.

Source