Skip to content

RSIGym goes open-source: Letting AI research how to improve AI, aiming for recursive self-improvement

Oct 10, 16:16

Beating AI News Flash: Evolvent AI, which builds AI training and evaluation tools, has open-sourced RSIGym, aiming to study recursive self-improvement (RSI) in AI. Simply put, it lets AI find system weaknesses on its own, propose modifications, test the results, and then attempt to use the improved system for the next round of research.

RSIGym turns training, inference, evaluation, and sandboxing into ready-made services. Research agents can generate training data, adjust fine-tuning parameters, train models, and then revise their plans based on evaluation results. It can also modify the Harness, the program that controls how models call tools and execute tasks. Researchers can have agents modify only the training data, only the Harness, or optimize both simultaneously. The platform centrally manages experiment budgets, permissions, and execution logs.

The team also launched RSI-Index to test AI agents' ability to improve another AI system. Six research agents each started from the same Qwen3.5 base model and a simple Harness, optimizing across five task categories including coding and math. Opus 5 ranked first overall. Its modified system improved its score on a 100-question subset of SWE-bench Verified from 17.67% to 50.33%; on the AIME math test, it improved from 31.67% to 97.78%.

In another set of experiments, only the Harness was modified without touching model weights. Based on DeepSeek Harness execution logs, Opus 5 fixed issues such as stopping work after output truncation and prematurely declaring task completion. With Qwen3.6 model weights completely unchanged, the Terminal-Bench 2.0 score improved from 30.34% to 40.82%.

The team also tried having Qwen3.8-27B improve itself, with it serving as both the research agent and the model being improved. It found a fine-tuning plan that appeared effective on 6 test questions, but on the full 37 questions, the score actually dropped slightly from 56.66% to 56.26%. In addition, each group of experiments conducted only one research run, and the agent could also access evaluation questions during the development phase. There is currently no evidence that the improved AI can continue to enhance its own R&D capabilities and achieve multiple rounds of recursive self-improvement.

Source