SWE-bench Verified: Real-World Agent Performance Leaderboard
Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.
Reasoning Engine
Conducting research…
Research Telemetry
live · demoSynthesized Answer · Deep Mode
Benchmarks across 12 production teams show native SDKs deliver 1.6–2.2× lower median latency on tool-heavy workflows and 30–45% fewer retry loops, driven by tighter coupling with the model's function-calling head. LangGraph closes most of that gap with custom checkpointing and parallel node execution, while retaining a decisive edge in three areas: (1) human-in-the-loop checkpoints, (2) deterministic replay for audit/compliance, and (3) hot-swapping models per node. Expect convergence by late 2026 as MCP standardizes tool surfaces and providers expose graph primitives — until then, regulated industries default to LangGraph, latency-sensitive consumer agents default to native SDKs.
Deep Analysis
The next 18 months will be defined by the agent reliability curve, not the model capability curve. Winners will own the eval/observability layer and a defensible workflow graph in a regulated vertical.
1 · Reliability is the new capability
Across SWE-bench Verified, GAIA, and WebArena, gains over 12 months are 80%+ attributable to better planning, verification, and tool fluency — not larger base models. Reliability scales sub-linearly with parameter count and super-linearly with eval coverage.
2 · Vertical beats horizontal on unit economics
Bounded task graphs make verification tractable. Vertical agents post 70%+ gross margins and 4–8 month payback; horizontal assistants are stuck at <40% margins and ambiguous ROI attribution.
3 · Computer-use is the integration unlock
Pixel/DOM-level agents bypass API gaps and let agents touch the long tail of legacy SaaS. Latency and reliability still gate production for high-stakes workflows.
4 · The platform layer is unsettled
Open frameworks (LangGraph, CrewAI) and native SDKs (OpenAI, Anthropic, Google) are converging. Expect consolidation around 2–3 winners by 2027; portability via MCP-style standards is the swing factor.
SWE-bench Verified — top model score over time
6-quarter trajectory
Agent benchmark scores by provider
benchmark composite (0–100)
Agent deployments by vertical
share of measured value (%)
Contradictions detected
Claim
Benchmarks show 60%+ agent task completion (SWE-bench Verified).
Counter
Production deployments report 25–40% on novel tasks — Gartner attributes the gap to eval-set leakage and easier issue distribution.
Claim
Horizontal copilots will win on distribution (a16z).
Counter
Vertical agents are posting 3–5× the revenue per seat with lower churn (Gartner, internal surveys).
Key Points
Vertical agents are out-scaling horizontal copilots in revenue per seat by 3–5×.
Benchmarks show 60% task completion; production deployments report 25–40% — eval-set leakage suspected.
Computer-use APIs cut integration timelines from months to weeks for legacy systems.
Eval/observability startups (Braintrust, Langfuse, Arize) are seeing the fastest ARR growth in the stack.
Enterprises with agents in production
31%
18pp YoYGartner Q1 2026
SWE-bench Verified (top closed)
65%
22pp YoYreal GitHub issue resolution
Vertical agent gross margin
72%
9pp YoYmedian across 40 disclosed startups
Inference cost per agent task
$0.18
−54% YoYblended, 30-min analyst-equivalent
Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.
Workflow-specific agents in legal, sales, and finance are achieving 40–70% task automation with bounded eval surface.
Screen-reading + structured action agents now handle multi-app workflows; reliability gates remain the production bottleneck.
Human performance on GAIA is 92%; best agent score is 49%. Long-horizon planning is the dominant failure mode.
31% of enterprises have at least one agent in production; 78% cite evaluation tooling as the #1 blocker to scale.
Open frameworks lead on portability; native SDKs lead on tool fidelity and latency. Convergence expected within 12 months.
Refine your research
Demo mode · All sources, insights, and data are mock-generated for illustration.