Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet
4.00T1 sourceAnthropic Engineering
Source record
Published by Anthropic Engineering (T1 source). The original is at https://www.anthropic.com/engineering/swe-bench-sonnet.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryAnthropic details the agent scaffolding built around an upgraded Claude 3.5 Sonnet for the SWE-bench Verified benchmark, where it scored 49% versus the prior 45% state of the art. The post explains how the agent navigates repositories, edits code, and runs tests to resolve real GitHub issues.
Why it mattersFirst-party breakdown of how Anthropic structures a coding agent around Claude, including tooling and evaluation harness details that practitioners can adapt.

Cited by
No citations on record.
