Forge Daughter
Most good work begins as a question.
Then comes the slow part: finding sources, trying the idea, writing the code, seeing whether it holds. Each step tends to live in a different tool, and the thread is easy to lose between them. Many questions stop there.
Forge Daughter is a place to keep going. Think it through with a model, search the web, run code, and make images, all in one conversation you can return to.
One conversation
A question rarely stays the size it started. It grows a source worth reading, a number worth checking, a small program worth running. Here each of those happens in the same place, in order, where you can see it.
When the model needs to run something, a sandbox starts in the cloud. Until then, nothing is running. Longer work can go to a cloud agent while you keep thinking. The conversation keeps what was asked, what was tried, and what came back. Close it and it waits.
Why we measure
An agent is more than its model. The harness decides what the model sees, which tools it can reach, and how it recovers when a command fails. The sandbox decides whether the work actually runs. Change either one and the same model scores differently.
So we test the whole system on Terminal-Bench, a public set of real tasks in a real terminal. It is not the only measure that matters. It is one anyone can check.
Our runs are in progress. The Forge Daughter mark on these charts is a placeholder, and it will move when the runs finish.
- Codex CLI 82.2
- Factory Droid 77.3
- Junie CLI 71.0
- Gemini CLI 61.4
- Forge Daughter —
- Warp 61.2
- Claude Code 58.0
- Goose 54.3
- OpenHands 51.9
Cost
Accuracy is one axis. A trial that costs nineteen dollars is a different result from one that costs six. Terminal-Bench 4.0 is newer and harder, and it reports both.
The useful question is not which agent wins. It is which one does your work at a price you can live with.
Time
Minutes count too. A long trial is time spent waiting, and a sandbox kept running. Faster is not always better. Slower is not always more careful. Both are worth seeing plainly.
A moving line
Benchmarks age quickly. The best Terminal-Bench 2.0 score nearly doubled in eight months. Then the board was archived and a harder one took its place.
A score is a snapshot of one set of tasks at one moment. When our runs finish, the placeholder becomes a number.
Bring a question. See how far it goes.