SWE-bench Verified: Real-World Agent Performance Leaderboard
Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.
Reasoning Engine
Conducting research…
Research Telemetry
live · demoSynthesized Answer · Deep Mode
Three structural reasons vertical wins in 2026: (1) eval coverage scales with task narrowness — a contract-redline agent can have 10,000 graded examples; a horizontal copilot cannot. (2) Integration moats compound — each new system of record (Workday, Epic, SAP) raises the wall around incumbents. (3) Buyer is identifiable: a vertical agent sells to a VP of Claims or Head of RevOps with a budget line, not to 'every knowledge worker.' The losing thesis is 'AGI replaces all white-collar work' — the winning thesis is '50 narrow products replace 50 narrow workflows.' Best risk/reward: regulated verticals (insurance, healthcare RCM, legal, financial back-office) where switching cost is high and eval discipline is rewarded.
Deep Analysis
The next 18 months will be defined by the agent reliability curve, not the model capability curve. Winners will own the eval/observability layer and a defensible workflow graph in a regulated vertical.
1 · Reliability is the new capability
Across SWE-bench Verified, GAIA, and WebArena, gains over 12 months are 80%+ attributable to better planning, verification, and tool fluency — not larger base models. Reliability scales sub-linearly with parameter count and super-linearly with eval coverage.
2 · Vertical beats horizontal on unit economics
Bounded task graphs make verification tractable. Vertical agents post 70%+ gross margins and 4–8 month payback; horizontal assistants are stuck at <40% margins and ambiguous ROI attribution.
3 · Computer-use is the integration unlock
Pixel/DOM-level agents bypass API gaps and let agents touch the long tail of legacy SaaS. Latency and reliability still gate production for high-stakes workflows.
4 · The platform layer is unsettled
Open frameworks (LangGraph, CrewAI) and native SDKs (OpenAI, Anthropic, Google) are converging. Expect consolidation around 2–3 winners by 2027; portability via MCP-style standards is the swing factor.
SWE-bench Verified — top model score over time
6-quarter trajectory
Agent benchmark scores by provider
benchmark composite (0–100)
Agent deployments by vertical
share of measured value (%)
Contradictions detected
Claim
Benchmarks show 60%+ agent task completion (SWE-bench Verified).
Counter
Production deployments report 25–40% on novel tasks — Gartner attributes the gap to eval-set leakage and easier issue distribution.
Claim
Horizontal copilots will win on distribution (a16z).
Counter
Vertical agents are posting 3–5× the revenue per seat with lower churn (Gartner, internal surveys).
Key Points
Vertical agents are out-scaling horizontal copilots in revenue per seat by 3–5×.
Benchmarks show 60% task completion; production deployments report 25–40% — eval-set leakage suspected.
Computer-use APIs cut integration timelines from months to weeks for legacy systems.
Eval/observability startups (Braintrust, Langfuse, Arize) are seeing the fastest ARR growth in the stack.
Enterprises with agents in production
31%
18pp YoYGartner Q1 2026
SWE-bench Verified (top closed)
65%
22pp YoYreal GitHub issue resolution
Vertical agent gross margin
72%
9pp YoYmedian across 40 disclosed startups
Inference cost per agent task
$0.18
−54% YoYblended, 30-min analyst-equivalent
Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.
Workflow-specific agents in legal, sales, and finance are achieving 40–70% task automation with bounded eval surface.
Screen-reading + structured action agents now handle multi-app workflows; reliability gates remain the production bottleneck.
Human performance on GAIA is 92%; best agent score is 49%. Long-horizon planning is the dominant failure mode.
31% of enterprises have at least one agent in production; 78% cite evaluation tooling as the #1 blocker to scale.
Open frameworks lead on portability; native SDKs lead on tool fidelity and latency. Convergence expected within 12 months.
Refine your research
Demo mode · All sources, insights, and data are mock-generated for illustration.