---
title: "Introducing Reason Agent 1.0"
description: "Our first results on DeepSWE v1.1. Competitive with frontier harnesses, with up to 39% lower median solving time against our benchmark baseline."
url: "https://reasonmachines.company/blog/reason-agent-deepswe"
section: "Blog"
---

# Introducing Reason Agent 1.0

_Our first results on DeepSWE v1.1. Competitive with frontier harnesses, with up to 39% lower median solving time against our benchmark baseline._

By Reason Machines

September 14, 2026

[Reason Agent](/) is the autonomous software engineer built by Reason Machines.

Today we’re introducing Reason Agent 1.0, our coding-agent harness built to help models solve software engineering tasks. Our first results on DeepSWE v1.1 are competitive with frontier harnesses, with up to 39% lower median solving time against our Codex benchmark baseline.[1](#ref-1)

### DeepSWE score vs. median wall time

DeepSWE score

Higher score · less time ↗Configs (60/71)

gpt-6-astraxhighhighmaxmediumlowgemini-3.8-flashhighmediumclaude-opus-5maxxhighhighmediumlowgpt-5.6-solmaxxhighhighmediumlowclaude-fable-5xhighmaxhighmediumlowgpt-5.6-terramaxxhighhighmediumlowglm-5.3maxkimi-k3maxgrok-4.6mediumxhighhighlowgpt-5.6-lunamaxxhighhighmediumlowgpt-5.5xhighhighmediumlowgemini-3.7-flashmediumhighlowglm-5.3-flashmaxdeepseek-v4-promaxclaude-opus-4.8maxxhighhighmediumlowqwen3.8-maxxhighmuse-spark-1.2xhighclaude-sonnet-5maxxhighhighmediumlowgrok-4.5highdeepseek-v4-flashmaxmuse-spark-1.1xhighgpt-5.4xhighgemini-3.6-flashhighglm-5.2maxhighgemini-3.5-flashhighkimi-k2.7-codedefaultclaude-sonnet-4.6highgemini-3.1-pro-previewhighReason AgentAstra xhigh

Reason AgentAstra · xhigh · our runs

73.0%DeepSWE score

13.80 minMedian wall time

Figure 1. Median recorded agent-execution wall time, excluding setup, queueing and verification, not end-to-end task latency. Published medians cover scored attempts, including failures; Reason pools timed attempts across both runs and retains all assignments in its score. Model lines connect configurations, not a fitted frontier. Different hosts, provider load and execution conditions make this descriptive, not a matched speed test. [[1]](#ref-1) [[2]](#ref-2)

## Successful solutions, earlier

Reason is our coding-agent harness, built around a minimal set of general-purpose tools. The model decides how to inspect a repository, make changes and test its work. We give it room to adapt its approach to the problem, rather than prescribe a fixed sequence of task-specific procedures.

We reject the assumption that better agents necessarily need more orchestration. As models become more capable, our view is that the harness should impose less workflow and place more trust in the model’s problem-solving decisions. Its responsibility is to provide reliable tools and clear execution boundaries, without deciding every step on the model’s behalf.

The aim is to reduce harness-generated context and forced coordination calls, giving the model a more direct path from investigation to a tested change. That is a design hypothesis, not a mechanism established by these runs. A small toolset alone does not guarantee fewer model turns.

## Alongside published reference points

Reason’s Astra result sits close to Datacurve’s leading published Astra configuration. The table combines our measurements with the public reference points.[2](#ref-2)[3](#ref-3)[1](#ref-1)

### Reason Agent is competitive with leading harnesses

DeepSWE v1.1 · observed success

Agent | Effort | Success |

Reason (Astra) | xhigh | 73%67–78% |

GPT-6 Astra | xhigh | 74%70–78% |

Gemini 3.8 Flash | high | 74%70–78% |

Claude Opus 5 | max | 74%69–78% |

GPT-5.6 Sol | max | 73%68–77% |

Claude Fable 5 | xhigh | 70%66–74% |

Figure 2. Reason (Astra) is our measurement; all other rows use mini-swe-agent and published DeepSWE results. Ranges show 95% Wilson intervals, unadjusted for repeated-task dependence. Evaluation methods differ. [[1]](#ref-1)[[2]](#ref-2)[[3]](#ref-3)[[6]](#ref-6)

## A side-by-side comparison of harnesses

Same model and effort, different harnesses. Reason used 36.7% less mean solving time on shared successes, averaging its attempts per task. Its overall median solving time was about 40% lower (39% across the combined runs) than Codex’s.[1](#ref-1)

We compare Reason and Codex on the same 113 tasks, both using Astra at xhigh effort. The curve combines both Reason runs and the single Codex run. Below it, the task-by-task view compares run 1 from each harness, including the two Reason recoveries, so each dot shows how both performed on the same task.

The completion curve shows the measured result: successful attempts finish earlier, at a similar overall observed success rate. By 20 minutes of solving time, final passing attempts accounted for 59.3% of Reason assignments versus 20.4% for Codex. These are retrospective completion times, not results from an experiment with a 20-minute deadline.[1](#ref-1)

### Reason and Codex: solving time

Astra xhigh · two Reason runs, one Codex run

Click a legend item to show or hide it.

Confirmed successes (%)

0

20

40

60

80

100

0102030405060

Solving time (minutes)

Figure 3. Final passing attempts by recorded agent time, not an early-stopping experiment. Reason and Codex retain all assignments in their denominators. Both use Astra at xhigh effort; separate provider routes and execution conditions limit causal claims about the harness. [[1]](#ref-1)

### Reason and Codex: task by task

Astra xhigh · run 1 from each harness · one dot per task

Left half: Reason · Right half: Codex

abs-module-cache-flags. Reason: pass; Codex: pass

abs-stepped-slices. Reason: pass; Codex: pass

actionlint-action-pinning-lint. Reason: pass; Codex: pass

adaptix-name-mapping-aliases. Reason: pass; Codex: pass

aiomonitor-task-snapshots-diff. Reason: pass; Codex: pass

anko-default-function-arguments. Reason: pass; Codex: pass

anko-typed-variable-bindings. Reason: pass; Codex: pass

arcane-drift-detection-baselines. Reason: pass; Codex: fail

arktype-json-schema-refs-dependencies. Reason: pass; Codex: pass

awilix-async-container-initialization. Reason: pass; Codex: fail

bandit-incremental-cache-control. Reason: pass; Codex: pass

bandit-interprocedural-taint-checks. Reason: pass; Codex: pass

bandit-structured-nosec-directives. Reason: fail; Codex: fail

boa-hierarchical-evaluation-cancellation. Reason: pass; Codex: pass

cattrs-partial-structuring-recovery. Reason: pass; Codex: pass

clack-async-autocomplete-options. Reason: pass; Codex: fail

claude-code-by-agents-recursive-delegation. Reason: pass; Codex: pass

cliffy-config-file-parsing. Reason: pass; Codex: pass

csstree-shorthand-expansion-compression. Reason: pass; Codex: pass

dasel-html-document-format. Reason: pass; Codex: pass

dateutil-rfc5545-timezone-interop. Reason: pass; Codex: pass

drizzle-orm-window-function-builders. Reason: pass; Codex: pass

dynamodb-toolbox-conditional-attribute-requirements. Reason: pass; Codex: pass

dynamodb-toolbox-lazy-recursive-schemas. Reason: pass; Codex: pass

effect-sse-httpapi-streaming. Reason: pass; Codex: fail

eicrud-keyset-pagination-cursor. Reason: pass; Codex: pass

etree-xml-diff-patch. Reason: fail; Codex: fail

expr-try-catch-errors. Reason: pass; Codex: pass

fastapi-deprecation-response-headers. Reason: pass; Codex: pass

fastapi-implicit-head-options. Reason: pass; Codex: pass

fd-deterministic-multi-key-sorting. Reason: pass; Codex: pass

geo-shapeindex-serialization. Reason: pass; Codex: pass

go-critic-doc-link-checker. Reason: fail; Codex: fail

go-genai-streamed-function-args. Reason: pass; Codex: pass

go-git-worktree-merge-conflicts. Reason: pass; Codex: pass

goreleaser-retry-publish-auditing. Reason: pass; Codex: pass

gql-incremental-graphql-delivery. Reason: fail; Codex: fail

happy-dom-abort-pending-body-reads. Reason: pass; Codex: pass

happy-dom-deterministic-intersectionobserver. Reason: fail; Codex: fail

helm-array-merge-strategies. Reason: fail; Codex: fail

helm-unified-manifest-stream. Reason: pass; Codex: pass

httpx-deterministic-cookie-store. Reason: pass; Codex: pass

httpx-multipart-response-parsing. Reason: pass; Codex: pass

httpx-streaming-json-iteration. Reason: pass; Codex: pass

igel-persist-feature-schema. Reason: fail; Codex: fail

ink-grid-box-layout. Reason: fail; Codex: fail

ipython-session-bundle-replay. Reason: pass; Codex: pass

katex-multicolumn-array-spans. Reason: pass; Codex: pass

kcp-go-multiplexed-kcp-streams. Reason: pass; Codex: fail

kea-atomic-signal-selectors. Reason: fail; Codex: fail

kgateway-consistent-hash-policy. Reason: pass; Codex: pass

kombu-single-active-consumer-priority. Reason: pass; Codex: pass

kombu-virtual-queue-dead-lettering. Reason: pass; Codex: pass

koota-composite-trait-aspects. Reason: pass; Codex: pass

koota-deferred-mutation-buffer. Reason: fail; Codex: pass

koota-entity-snapshot-rollback. Reason: pass; Codex: fail

koota-pair-relation-tracking. Reason: fail; Codex: fail

koota-query-predicates. Reason: pass; Codex: fail

kysely-window-grouping-helpers. Reason: pass; Codex: pass

langchain-request-coalescing. Reason: pass; Codex: pass

mashumaro-flattened-dataclass-fields. Reason: pass; Codex: pass

meriyah-explicit-resource-declarations. Reason: fail; Codex: fail

mnamer-daemon-watch-lifecycle. Reason: pass; Codex: pass

mobly-grouped-test-barriers. Reason: pass; Codex: pass

narwhals-rolling-window-suite. Reason: pass; Codex: pass

numba-stencil-boundary-modes. Reason: pass; Codex: pass

obsidian-linter-auto-table-of-contents. Reason: fail; Codex: fail

obsidian-linter-link-format-conversion. Reason: fail; Codex: fail

obsidian-linter-scoped-ignore-markers. Reason: pass; Codex: pass

ofetch-per-origin-circuit-breaker. Reason: pass; Codex: pass

onedump-dump-encryption-pipeline. Reason: fail; Codex: fail

opa-rego-rule-profiling. Reason: fail; Codex: fail

opa-template-string-reconstruction. Reason: pass; Codex: fail

optique-conditional-option-dependencies. Reason: fail; Codex: pass

oxvg-structural-selector-preservation. Reason: fail; Codex: pass

participle-grammar-conflict-analysis. Reason: fail; Codex: fail

pebble-durability-wait-apis. Reason: pass; Codex: pass

pest-character-class-coalescing. Reason: fail; Codex: fail

prometheus-transactional-reload-status. Reason: fail; Codex: pass

prometheus-typed-label-sorting. Reason: pass; Codex: pass

psd-tools-blend-range-api. Reason: pass; Codex: pass

pwntools-tube-multiplexing. Reason: pass; Codex: fail

python-statemachine-state-data-scoping. Reason: fail; Codex: fail

query-persist-restored-query-state. Reason: pass; Codex: pass

quill-shared-toolbar-focus. Reason: pass; Codex: pass

returns-validated-error-accumulation. Reason: pass; Codex: pass

scc-bounded-memory-spilling. Reason: pass; Codex: pass

scriggo-method-declarations. Reason: pass; Codex: pass

skrub-duration-encoding. Reason: pass; Codex: pass

sql-formatter-bigquery-pipe-formatting. Reason: pass; Codex: pass

sqlfmt-create-table-ddl-formatting. Reason: fail; Codex: fail

sqlite-utils-safe-import-checkpoints. Reason: pass; Codex: pass

superjson-error-stack-serialization. Reason: pass; Codex: pass

task-task-graph-export. Reason: pass; Codex: pass

tengo-callable-instance-isolation. Reason: pass; Codex: pass

tengo-destructuring-bindings. Reason: pass; Codex: pass

termenv-preserve-ansi-resets. Reason: fail; Codex: fail

testem-bail-on-test-failure. Reason: fail; Codex: fail

testem-per-launcher-reports. Reason: fail; Codex: pass

textual-kitty-key-phases. Reason: pass; Codex: pass

textual-richlog-follow-state. Reason: pass; Codex: pass

tomlkit-toml-table-converters. Reason: pass; Codex: pass

true-myth-iterable-collection-combinators. Reason: pass; Codex: pass

ts-pattern-match-each. Reason: fail; Codex: fail

updo-policy-alerting. Reason: pass; Codex: pass

valibot-recursive-schema-composition. Reason: pass; Codex: pass

vitest-duration-sharding. Reason: pass; Codex: pass

vulture-persistent-analysis-cache. Reason: fail; Codex: fail

wasmi-trap-coredumps. Reason: pass; Codex: pass

wazero-multi-module-snapshots. Reason: fail; Codex: fail

yaegi-go-embed-directives. Reason: pass; Codex: pass

yjs-map-conflict-detection. Reason: pass; Codex: pass

ytt-jsonpath-query-api. Reason: pass; Codex: pass

Both passReason passesCodex passesNeither

Figure 4. Reason run 1 versus Codex, including the Kea setup-failure retry and unchanged-patch Pwntools regrade. Blue means a pass. Hover or focus a dot to inspect both outcomes. [[1]](#ref-1)

## What comes next

These results are preliminary. We believe coding agents should be built around a minimal set of general-purpose tools, giving capable models more freedom to solve tasks. Next, we will test that principle on fresh tasks under matched conditions, varying tools and orchestration one at a time to see whether less context and fewer calls can deliver faster results without sacrificing correctness.

## Evaluation notes

- These measurements evaluate the extracted coding harness, not the hosted product. Separate provider routes and execution infrastructure limit causal and formal-parity claims.[1](#ref-1)
- Solving time excludes setup, pre-trial queueing and verification. Figure 4 shows Reason run 1 only; aggregate figures retain both runs.[1](#ref-1)
- Updated September 14: Reason has 165 passes, 61 failures and no ungraded outcomes across 226 assignments. Kea failed on a fresh attempt after its setup failure; Pwntools passed when its unchanged patch was regraded. No graded failures were replaced. The JSON preserves original outcomes and recovery details. Published scores retain Datacurve’s scoring rules.[1](#ref-1)[2](#ref-2)[3](#ref-3)
- Run results, aggregate scores and solving times are in the results JSON.[1](#ref-1)

## References

<a id="ref-1"></a>
[1] Reason Machines. DeepSWE v1.1. [Run results and aggregate data (JSON)](/research/reason-agent-deepswe/deepswe.json).

<a id="ref-2"></a>
[2] Datacurve. DeepSWE v1.1. [Reporting methodology](https://deepswe.datacurve.ai/blog/deepswe-v1-1) · [published leaderboard](https://deepswe.datacurve.ai/artifacts/v1.1/leaderboard-live.json).

<a id="ref-3"></a>
[3] Datacurve. DeepSWE v1.1 rollout data. [Trial browser](https://deepswe.datacurve.ai/data/v1.1) · [raw trial index](https://deepswe.datacurve.ai/artifacts/v1.1/trials.json).

<a id="ref-4"></a>
[4] Datacurve. DeepSWE: Measuring frontier coding agents. [Evaluation methodology](https://deepswe.datacurve.ai/blog/deepswe#evaluation-harness).

<a id="ref-5"></a>
[5] Datacurve / Harbor. DeepSWE v1.1, 113 tasks. [Dataset](https://hub.harborframework.com/datasets/datacurve/deep-swe-1-1/latest).

<a id="ref-6"></a>
[6] NIST/SEMATECH. e-Handbook of Statistical Methods, §7.2.4.1. [Wilson confidence intervals for a proportion](https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm).
