SWE-bench Verified: Real-World Agent Performance Leaderboard
Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.
Reasoning Engine
Conducting research…
Research Telemetry
live · demoSynthesized Answer · Deep Mode
Mitigation maturity is uneven. Strongest controls today: per-action human-in-the-loop confirmation for irreversible operations, per-domain credential scoping, allow-listed action types per workflow, and content-source provenance tracking. Weakest: detection of subtle prompt injection in rendered HTML/PDF, multi-step exfiltration that splits sensitive data across benign-looking actions, and incident response — most SOCs cannot replay an agent session like they can a user session. Expect a wave of agent-specific security tooling in 2026: runtime sandboxes, action firewalls, and red-team-as-a-service. Regulated industries should require: (a) deterministic action logs with replay, (b) explicit consent for write actions touching $X or N records, (c) credential vaulting with short-lived tokens, (d) red-team coverage on the agent runtime itself, not just the model.
Deep Analysis
The next 18 months will be defined by the agent reliability curve, not the model capability curve. Winners will own the eval/observability layer and a defensible workflow graph in a regulated vertical.
1 · Reliability is the new capability
Across SWE-bench Verified, GAIA, and WebArena, gains over 12 months are 80%+ attributable to better planning, verification, and tool fluency — not larger base models. Reliability scales sub-linearly with parameter count and super-linearly with eval coverage.
2 · Vertical beats horizontal on unit economics
Bounded task graphs make verification tractable. Vertical agents post 70%+ gross margins and 4–8 month payback; horizontal assistants are stuck at <40% margins and ambiguous ROI attribution.
3 · Computer-use is the integration unlock
Pixel/DOM-level agents bypass API gaps and let agents touch the long tail of legacy SaaS. Latency and reliability still gate production for high-stakes workflows.
4 · The platform layer is unsettled
Open frameworks (LangGraph, CrewAI) and native SDKs (OpenAI, Anthropic, Google) are converging. Expect consolidation around 2–3 winners by 2027; portability via MCP-style standards is the swing factor.
SWE-bench Verified — top model score over time
6-quarter trajectory
Agent benchmark scores by provider
benchmark composite (0–100)
Agent deployments by vertical
share of measured value (%)
Contradictions detected
Claim
Benchmarks show 60%+ agent task completion (SWE-bench Verified).
Counter
Production deployments report 25–40% on novel tasks — Gartner attributes the gap to eval-set leakage and easier issue distribution.
Claim
Horizontal copilots will win on distribution (a16z).
Counter
Vertical agents are posting 3–5× the revenue per seat with lower churn (Gartner, internal surveys).
Key Points
Vertical agents are out-scaling horizontal copilots in revenue per seat by 3–5×.
Benchmarks show 60% task completion; production deployments report 25–40% — eval-set leakage suspected.
Computer-use APIs cut integration timelines from months to weeks for legacy systems.
Eval/observability startups (Braintrust, Langfuse, Arize) are seeing the fastest ARR growth in the stack.
Enterprises with agents in production
31%
18pp YoYGartner Q1 2026
SWE-bench Verified (top closed)
65%
22pp YoYreal GitHub issue resolution
Vertical agent gross margin
72%
9pp YoYmedian across 40 disclosed startups
Inference cost per agent task
$0.18
−54% YoYblended, 30-min analyst-equivalent
Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.
Workflow-specific agents in legal, sales, and finance are achieving 40–70% task automation with bounded eval surface.
Screen-reading + structured action agents now handle multi-app workflows; reliability gates remain the production bottleneck.
Human performance on GAIA is 92%; best agent score is 49%. Long-horizon planning is the dominant failure mode.
31% of enterprises have at least one agent in production; 78% cite evaluation tooling as the #1 blocker to scale.
Open frameworks lead on portability; native SDKs lead on tool fidelity and latency. Convergence expected within 12 months.
Refine your research
Demo mode · All sources, insights, and data are mock-generated for illustration.