++

AI that builds
better AI.
Autonomously.

Each day it comes up with an idea, implements it, evaluates it in the environment, and learns from every experiment, all on its own. Or drop in your own idea, and it biases the search toward that.

NOW BUILDING · CUSTOMER-SUPPORT AGENTS
THE RESEARCH LOGBETWEEN RUNS
EXP 17Rewrite the agent with a dramatically shorter, priority-ordered system prompt, add programmatic step tracking via injected system messages after list-returning tool calls, fix the user-tool give→instruct→wait flow with f…21%
EXP 01Replace the minimal seed system prompt with a deeply policy-aware, workflow-structured system prompt that encodes the complete retail domain policy as precise operational procedures (with step-by-step tool-calling workfl…88%
EXP 01Replace the minimal seed agent with a retrieval-augmented agent that pre-processes the entire 698-document corpus into a persistent embedding index + discoverable-tool catalog on first use, then on each task proactively…21%
EXP 02Replace the minimal system prompt with a comprehensive procedure-driven operational playbook that embeds explicit workflow prerequisites, mandatory tool-use rules (especially for calculate and getproductdetails…78%
EXP 02Build a workflow-aware banking agent that pre-classifies customer requests into workflow archetypes (credit card application, cash back dispute, transaction dispute, transfer protocol, referral, account management), proa…24%
GRADED ON HELD-OUT TASKSSEE THE EXPLORATION ↓
01 · THE LOOP

Autonomous AI research loop.

Better AI system, built by an AI researcher that runs the whole loop: ideate, implement, experiment, evaluate, learn, repeat. Every attempt is preserved and every result feeds a growing archive of knowledge that steers the next idea, so the loop runs open-ended toward a measurably better system.

POWERED BY KAPSO ↗
EVOLVECONTEXT MANAGEREXPERIMENT GENERATOREXPERIMENT SELECTORREWARD SYNTHESIZERDEVELOPER AGENTAI SYSTEMthe starting pointIMPROVED AI SYSTEMWORLD KNOWLEDGEfeedback · code · reportsLEARNING ENGINE.learn()HISTORY OF EXPERIMENTS

SCROLL TO EXPLORE THE FULL LOOP →

02 · TRACK RECORD

Verified in the open.

EXHIBIT A · MLE-BENCH · OPENAINº1 OPEN SOURCE

Machine-Learning Engineering

Given a new problem, it builds the entire solution on its own, from raw data to a trained, tuned model, and outperforms every other open agent.

MEDAL RATE

Leeroo
50.7%
R&D-Agent
35.1%
AIRA-dojo
31.6%
ML-Master
29.3%
AIDE
17.1%

MEDAL RATE BY TASK COMPLEXITY

LEEROO BEST REPORTED
Low
68.2%
68.2%
Medium
44.7%
22.0%
High
40.0%
24.4%

BEST REPORTED PER TIER: R&D-AGENT (LOW) · AIRA-DOJO (MEDIUM) · ML-MASTER (HIGH)

HIGHEST MEDAL RATE AMONG OPEN, REPRODUCIBLE ML-ENGINEERING AGENTS
EXHIBIT B · ALE-BENCHELO 1909 · ▲ 30 VS BEST

Long-Horizon Algorithm Design

The problems have no perfect answer. It designs an algorithm, then rewrites it for hours to push the score higher, head to head with human experts.

LEEROO BEST REPORTED
Gravity Sorting
2528
2446
Coverage Design
2310
2880
Path Planning
2194
2153
Logistics
2040
1965
Signal Encoding
2022
1457
Urban Planning
1980
1980
Route Planning
1839
1740
Network Design
1607
1652
Load Balancing
1353
1331
Zoning
1221
1189
100020003000 ELO
RATED AGAINST HUMAN CONTESTANTS
EXHIBIT C · ONE SAMPLE, END TO END

Open-Ended Optimization Problem.

A sample of the loop on the hardest class of problems: open-ended optimization with no known best answer. This one is drawn from a live AtCoder contest. Encode up to a hundred messages as networks; the channel randomly rewires each one and erases its labels; the receiver must still decode it. There is no formula for the best design. The loop invented its own, rated on the contest's human ladder.

VIEW THE PROBLEM ON ATCODER ↗
2022ELO762 VS AVERAGE HUMAN

AVERAGE HUMAN 1260 · ALE-BENCH

011420021690031755041608052022061930
BEST IMPROVED EXPLORED
3 IMPROVED · 2 EXPLORED
03 · THE DOMAINS

A better support agent, for every domain.

SELECT A DOMAIN TO OPEN ITS WORKSPACE

Banking knowledge.

LIVE RUN · DOCUMENT RETRIEVAL
BEST SCORE · GRADED ON TASKS IT NEVER TRAINS ON
39%24 PTS VS TAU-BENCH

BASELINE 15% · TAU-BENCH · THE BASE MODEL ON THIS BENCHMARK

0121%0224%0324%0415%0524%0639%0727%0821%0924%1033%1115%1218%1324%1421%1527%1624%1721%
BEST IMPROVED EXPLORED
2 IMPROVED · 14 EXPLORED
04 · STEER THE RESEARCHER

Give the researcher an idea to try.

IT BIASES THE NEXT EXPERIMENT TOWARD YOUR IDEA

0 QUEUED · LEEROO SIGN-IN REQUIRED