FrogNano hits 61.5% on SWE-bench Verified with just 4 billion parameters
Researcher Minseon Kim posted the result on X. FrogNano, a 4-billion-parameter model, scores 61.5% on SWE-bench Verified without any distillation from a larger model. The base is Qwen3.5-4B, tuned with pure reinforcement learning on synthetic tasks generated by a pipeline called TaskPilot. The whole run took five iterations, 300 tasks each — a surprisingly small budget for performance that rivals models six to eight times larger.
TaskPilot is not a static dataset. It generates tasks "online" for every model checkpoint. At each iteration the system samples tasks inside that checkpoint's "learnability zone": problems the model can already make progress on but cannot yet solve completely. When improvement stalls, the pipeline repeats with the updated checkpoint. The approach replaces hundreds of thousands of human-labeled examples with a short, focused loop that produces exactly what the model needs right now.
The researchers found that at the 4B scale the execution harness matters critically. They built LEAF, a lightweight harness that lifted the base Qwen3.5-4B solve rate from 8.3% with R2E-Gym to 37.2% on SWE-bench Verified. That jump is not a model improvement — it is better environment access, less technical friction, more room to learn. Without LEAF the model would have started too low for reinforcement learning to pull it up.
Experiments show a consistent gap between pass@1 and pass@8 for FrogNano versus the base Qwen, evidence that reinforcement learning widens the