Reasoning Engine

Conducting research…

Step 1 / 5
  1. Discovering sources
    Identified 6 candidate sources across 4 publication types.
  2. Analyzing sources
    Extracted atomic claims; scored credibility, recency, and bias on each.
  3. Cross-referencing
    Detected contradictions and reconciled overlapping claims.
  4. Synthesizing findings
    Compressed claim graph into structural themes and a working thesis.
  5. Generating intelligence
    Drafted executive brief, evidence map, risks, and recommendations.

Research Telemetry

live · demo
Reasoning
92/100
Confidence
86/100
Evidence
88/100
Depth
94/100
Diversity
82/100

Synthesized Answer · Deep Mode

Project agent ARR by vertical

Three model variants. Bear ($9B 2026 → $28B 2028): enterprise procurement slows, security incidents force per-deal compliance reviews, hyperscaler bundles win mid-market. Base ($16B → $55B): current adoption curve continues; vertical incumbents emerge in 4–6 categories. Bull ($22B → $80B): computer-use APIs unlock long-tail SaaS, eval tooling matures, and regulated verticals scale from pilots to enterprise-wide deployments. Highest-confidence segments through 2028: code (SWE-bench tailwind), customer support (ROI is measurable), and healthcare RCM (labor cost crisis + structured workflow). Lowest-confidence: legal partner-facing work (judgment-heavy, low error tolerance) and complex M&A diligence.

Deep Analysis

86% confidence

The next 18 months will be defined by the agent reliability curve, not the model capability curve. Winners will own the eval/observability layer and a defensible workflow graph in a regulated vertical.

1 · Reliability is the new capability

Across SWE-bench Verified, GAIA, and WebArena, gains over 12 months are 80%+ attributable to better planning, verification, and tool fluency — not larger base models. Reliability scales sub-linearly with parameter count and super-linearly with eval coverage.

2 · Vertical beats horizontal on unit economics

Bounded task graphs make verification tractable. Vertical agents post 70%+ gross margins and 4–8 month payback; horizontal assistants are stuck at <40% margins and ambiguous ROI attribution.

3 · Computer-use is the integration unlock

Pixel/DOM-level agents bypass API gaps and let agents touch the long tail of legacy SaaS. Latency and reliability still gate production for high-stakes workflows.

4 · The platform layer is unsettled

Open frameworks (LangGraph, CrewAI) and native SDKs (OpenAI, Anthropic, Google) are converging. Expect consolidation around 2–3 winners by 2027; portability via MCP-style standards is the swing factor.

SWE-bench Verified — top model score over time

6-quarter trajectory

33%42%49%55%60%65%Q1'25Q2'25Q3'25Q4'25Q1'26Q2'26

Agent benchmark scores by provider

benchmark composite (0–100)

Claude (Anthropic)
65
GPT (OpenAI)
62
Gemini (Google)
58
Llama 4 (Meta)
38
Qwen 3 (Alibaba)
34
Mistral Large
30
Closed Open

Agent deployments by vertical

share of measured value (%)

Software engineering32%
Sales & GTM ops21%
Customer support19%
Legal & compliance14%
Finance & ops14%

Contradictions detected

Claim

Benchmarks show 60%+ agent task completion (SWE-bench Verified).

Counter

Production deployments report 25–40% on novel tasks — Gartner attributes the gap to eval-set leakage and easier issue distribution.

Claim

Horizontal copilots will win on distribution (a16z).

Counter

Vertical agents are posting 3–5× the revenue per seat with lower churn (Gartner, internal surveys).

Key Points

  • 2026 base case: $14–18B global vertical-agent ARR
  • 2028 base case: $48–65B; bull case $80B
  • Top 3 segments: software eng, customer support, healthcare RCM
  • Payback <9 months + 65%+ GM drive the curve
  • Incumbent bundling (Salesforce, Microsoft) is the biggest headwind

Knowledge Graph

10 nodes · 11 edges
topicconceptcompanyentity
AI agentsReliabilityVertical agentsEval / observabilityComputer-use APIsMCP / portabilityAnthropicOpenAILangChain / LangGraphSWE-bench

Auto-generated Insights

Trend

Vertical agents are out-scaling horizontal copilots in revenue per seat by 3–5×.

Contradiction

Benchmarks show 60% task completion; production deployments report 25–40% — eval-set leakage suspected.

Finding

Computer-use APIs cut integration timelines from months to weeks for legacy systems.

Signal

Eval/observability startups (Braintrust, Langfuse, Arize) are seeing the fastest ARR growth in the stack.

Structured Data

Extracted from sources

Enterprises with agents in production

31%

18pp YoY

Gartner Q1 2026

SWE-bench Verified (top closed)

65%

22pp YoY

real GitHub issue resolution

Vertical agent gross margin

72%

9pp YoY

median across 40 disclosed startups

Inference cost per agent task

$0.18

−54% YoY

blended, 30-min analyst-equivalent

Sources6 ranked

Sorted by relevance
S
swebench.com·this week
Research Paper

SWE-bench Verified: Real-World Agent Performance Leaderboard

Top closed models resolve 55–65% of real GitHub issues end-to-end; leading open-weight agents trail at 28–34%.

Cred
96
Auth
94
Fresh
94
Rel
88
Center
Strongevidence
A
a16z.com·this week
Article

The Vertical Agent Thesis

Workflow-specific agents in legal, sales, and finance are achieving 40–70% task automation with bounded eval surface.

Cred
84
Auth
85
Fresh
92
Rel
88
Center
Strongevidence
A
anthropic.com·this week
Blog

Computer Use, One Year In: What Actually Works

Screen-reading + structured action agents now handle multi-app workflows; reliability gates remain the production bottleneck.

Cred
92
Auth
90
Fresh
96
Rel
88
Center
Strongevidence
A
arxiv.org·this week
Research Paper

GAIA: A Benchmark for General AI Assistants

Human performance on GAIA is 92%; best agent score is 49%. Long-horizon planning is the dominant failure mode.

Cred
97
Auth
96
Fresh
80
Rel
88
Center
Strongevidence
G
gartner.com·this week
Report

Enterprise Agent Adoption Survey, Q1 2026

31% of enterprises have at least one agent in production; 78% cite evaluation tooling as the #1 blocker to scale.

Cred
91
Auth
93
Fresh
90
Rel
88
Center
Strongevidence
H
huggingface.co·this week
Blog

LangGraph vs Native SDKs: A Framework Comparison

Open frameworks lead on portability; native SDKs lead on tool fidelity and latency. Convergence expected within 12 months.

Cred
84
Auth
85
Fresh
95
Rel
88
Center
Strongevidence

Refine your research

Demo mode · All sources, insights, and data are mock-generated for illustration.