Research Preview

Introducing Reason Agent 1.0

Reason Agent is the autonomous software engineer built by Reason Machines.

Today we’re introducing Reason Agent 1.0, our coding-agent harness built to help models solve software engineering tasks. Our first results on DeepSWE v1.1 are competitive with frontier harnesses, with up to 39% lower median solving time against our Codex benchmark baseline.1

DeepSWE score vs. median wall time

DeepSWE score
Higher score · less time ↗
Configs (60/71)
gpt-6-astra
gemini-3.8-flash
claude-opus-5
gpt-5.6-sol
claude-fable-5
gpt-5.6-terra
glm-5.3
kimi-k3
grok-4.6
gpt-5.6-luna
gpt-5.5
gemini-3.7-flash
glm-5.3-flash
deepseek-v4-pro
claude-opus-4.8
qwen3.8-max
muse-spark-1.2
claude-sonnet-5
grok-4.5
deepseek-v4-flash
muse-spark-1.1
gpt-5.4
gemini-3.6-flash
glm-5.2
gemini-3.5-flash
kimi-k2.7-code
claude-sonnet-4.6
gemini-3.1-pro-preview
Reason Agent
0%10%20%30%40%50%60%70%80%0 min10 min20 min30 min40 min50 min60 min70 min80 minMedian wall time per task (minutes)

Figure 1. Median recorded agent-execution wall time, excluding setup, queueing and verification, not end-to-end task latency. Published medians cover scored attempts, including failures; Reason pools timed attempts across both runs and retains all assignments in its score. Model lines connect configurations, not a fitted frontier. Different hosts, provider load and execution conditions make this descriptive, not a matched speed test. [1] [2]

Successful solutions, earlier

Reason is our coding-agent harness, built around a minimal set of general-purpose tools. The model decides how to inspect a repository, make changes and test its work. We give it room to adapt its approach to the problem, rather than prescribe a fixed sequence of task-specific procedures.

We reject the assumption that better agents necessarily need more orchestration. As models become more capable, our view is that the harness should impose less workflow and place more trust in the model’s problem-solving decisions. Its responsibility is to provide reliable tools and clear execution boundaries, without deciding every step on the model’s behalf.

The aim is to reduce harness-generated context and forced coordination calls, giving the model a more direct path from investigation to a tested change. That is a design hypothesis, not a mechanism established by these runs. A small toolset alone does not guarantee fewer model turns.

Alongside published reference points

Reason’s Astra result sits close to Datacurve’s leading published Astra configuration. The table combines our measurements with the public reference points.231

Reason Agent is competitive with leading harnesses

DeepSWE v1.1 · observed success

AgentEffortSuccess
GPT-6 Astraxhigh74%70–78%
Gemini 3.8 Flashhigh74%70–78%
Claude Opus 5max74%69–78%
GPT-5.6 Solmax73%68–77%
Claude Fable 5xhigh70%66–74%
Figure 2. Reason (Astra) is our measurement; all other rows use mini-swe-agent and published DeepSWE results. Ranges show 95% Wilson intervals, unadjusted for repeated-task dependence. Evaluation methods differ. [1][2][3][6]

A side-by-side comparison of harnesses

Same model and effort, different harnesses. Reason used 36.7% less mean solving time on shared successes, averaging its attempts per task. Its overall median solving time was about 40% lower (39% across the combined runs) than Codex’s.1

We compare Reason and Codex on the same 113 tasks, both using Astra at xhigh effort. The curve combines both Reason runs and the single Codex run. Below it, the task-by-task view compares run 1 from each harness, including the two Reason recoveries, so each dot shows how both performed on the same task.

The completion curve shows the measured result: successful attempts finish earlier, at a similar overall observed success rate. By 20 minutes of solving time, final passing attempts accounted for 59.3% of Reason assignments versus 20.4% for Codex. These are retrospective completion times, not results from an experiment with a 20-minute deadline.1

Reason and Codex: solving time

Astra xhigh · two Reason runs, one Codex run

Click a legend item to show or hide it.

Confirmed successes (%)

0
20
40
60
80
100
0102030405060

Solving time (minutes)

Figure 3. Final passing attempts by recorded agent time, not an early-stopping experiment. Reason and Codex retain all assignments in their denominators. Both use Astra at xhigh effort; separate provider routes and execution conditions limit causal claims about the harness. [1]

Reason and Codex: task by task

Astra xhigh · run 1 from each harness · one dot per task

Left half: Reason · Right half: Codex

Both passReason passesCodex passesNeither
Figure 4. Reason run 1 versus Codex, including the Kea setup-failure retry and unchanged-patch Pwntools regrade. Blue means a pass. Hover or focus a dot to inspect both outcomes. [1]

What comes next

These results are preliminary. We believe coding agents should be built around a minimal set of general-purpose tools, giving capable models more freedom to solve tasks. Next, we will test that principle on fresh tasks under matched conditions, varying tools and orchestration one at a time to see whether less context and fewer calls can deliver faster results without sacrificing correctness.

Evaluation notes

  • These measurements evaluate the extracted coding harness, not the hosted product. Separate provider routes and execution infrastructure limit causal and formal-parity claims.1
  • Solving time excludes setup, pre-trial queueing and verification. Figure 4 shows Reason run 1 only; aggregate figures retain both runs.1
  • Updated September 14: Reason has 165 passes, 61 failures and no ungraded outcomes across 226 assignments. Kea failed on a fresh attempt after its setup failure; Pwntools passed when its unchanged patch was regraded. No graded failures were replaced. The JSON preserves original outcomes and recovery details. Published scores retain Datacurve’s scoring rules.123
  • Run results, aggregate scores and solving times are in the results JSON.1

References

  1. Reason Machines. DeepSWE v1.1. Run results and aggregate data (JSON).
  2. Datacurve. DeepSWE v1.1. Reporting methodology · published leaderboard.
  3. Datacurve. DeepSWE v1.1 rollout data. Trial browser · raw trial index.
  4. Datacurve. DeepSWE: Measuring frontier coding agents. Evaluation methodology.
  5. Datacurve / Harbor. DeepSWE v1.1, 113 tasks. Dataset.
  6. NIST/SEMATECH. e-Handbook of Statistical Methods, §7.2.4.1. Wilson confidence intervals for a proportion.

Authors

Reason Machines