Sakana RL Conductor

RL Conductor is a 7B parameter reinforcement-learning-trained orchestrator built on Qwen2.5-7B that routes subtasks among models like GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro. It achieves state-of-the-art scores of 83.9% on LiveCodeBench and 87.5% on GPQA-Diamond at roughly 1.8K tokens per query, making it about six times cheaper than comparable multi-agent ensembles; the paper is dated April 27, 2026, with the Fugu beta expected in late April or early May 2026.