SWE-bench Verified: Real-World Agent Performance Leaderboard
Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.
Reasoning Engine
Conducting research…
Research Telemetry
live · demoSynthesized Answer · Deep Mode
Three model variants. Bear ($9B 2026 → $28B 2028): enterprise procurement slows, security incidents force per-deal compliance reviews, hyperscaler bundles win mid-market. Base ($16B → $55B): current adoption curve continues; vertical incumbents emerge in 4–6 categories. Bull ($22B → $80B): computer-use APIs unlock long-tail SaaS, eval tooling matures, and regulated verticals scale from pilots to enterprise-wide deployments. Highest-confidence segments through 2028: code (SWE-bench tailwind), customer support (ROI is measurable), and healthcare RCM (labor cost crisis + structured workflow). Lowest-confidence: legal partner-facing work (judgment-heavy, low error tolerance) and complex M&A diligence.
Deep Analysis
The next 18 months will be defined by the agent reliability curve, not the model capability curve. Winners will own the eval/observability layer and a defensible workflow graph in a regulated vertical.
1 · Reliability is the new capability
Across SWE-bench Verified, GAIA, and WebArena, gains over 12 months are 80%+ attributable to better planning, verification, and tool fluency — not larger base models. Reliability scales sub-linearly with parameter count and super-linearly with eval coverage.
2 · Vertical beats horizontal on unit economics
Bounded task graphs make verification tractable. Vertical agents post 70%+ gross margins and 4–8 month payback; horizontal assistants are stuck at <40% margins and ambiguous ROI attribution.
3 · Computer-use is the integration unlock
Pixel/DOM-level agents bypass API gaps and let agents touch the long tail of legacy SaaS. Latency and reliability still gate production for high-stakes workflows.
4 · The platform layer is unsettled
Open frameworks (LangGraph, CrewAI) and native SDKs (OpenAI, Anthropic, Google) are converging. Expect consolidation around 2–3 winners by 2027; portability via MCP-style standards is the swing factor.
SWE-bench Verified — top model score over time
6-quarter trajectory
Agent benchmark scores by provider
benchmark composite (0–100)
Agent deployments by vertical
share of measured value (%)
Contradictions detected
Claim
Benchmarks show 60%+ agent task completion (SWE-bench Verified).
Counter
Production deployments report 25–40% on novel tasks — Gartner attributes the gap to eval-set leakage and easier issue distribution.
Claim
Horizontal copilots will win on distribution (a16z).
Counter
Vertical agents are posting 3–5× the revenue per seat with lower churn (Gartner, internal surveys).
Key Points
Vertical agents are out-scaling horizontal copilots in revenue per seat by 3–5×.
Benchmarks show 60% task completion; production deployments report 25–40% — eval-set leakage suspected.
Computer-use APIs cut integration timelines from months to weeks for legacy systems.
Eval/observability startups (Braintrust, Langfuse, Arize) are seeing the fastest ARR growth in the stack.
Enterprises with agents in production
31%
18pp YoYGartner Q1 2026
SWE-bench Verified (top closed)
65%
22pp YoYreal GitHub issue resolution
Vertical agent gross margin
72%
9pp YoYmedian across 40 disclosed startups
Inference cost per agent task
$0.18
−54% YoYblended, 30-min analyst-equivalent
Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.
Workflow-specific agents in legal, sales, and finance are achieving 40–70% task automation with bounded eval surface.
Screen-reading + structured action agents now handle multi-app workflows; reliability gates remain the production bottleneck.
Human performance on GAIA is 92%; best agent score is 49%. Long-horizon planning is the dominant failure mode.
31% of enterprises have at least one agent in production; 78% cite evaluation tooling as the #1 blocker to scale.
Open frameworks lead on portability; native SDKs lead on tool fidelity and latency. Convergence expected within 12 months.
Refine your research
Demo mode · All sources, insights, and data are mock-generated for illustration.