Why We Label AI Coding Agent Tools as Harnesses, Not Products
The 2026 AI coding market has split into agentic coding stacks — harness, model, billing, and permissions — and we evaluate the stack, not the brand name.
Published 2026-06-29
Why We Label AI Coding Agent Tools as Harnesses, Not Products
TL;DR: We stopped reviewing AI coding agents by brand in 2026 because the harness, model tier, billing layer, and permission model matter more than the product name.
The Context
Tool Crucible started by reviewing products: Cursor, Copilot, Windsurf, Devin Desktop. In 2026, those products became façades for different harness-and-model combinations. Copilot now bundles Fable 5, Opus 4.8, or Sonnet 4.6 depending on tier. Cursor routes Composer/Auto through its own stack or Third-Party API. Devin Desktop runs SWE 1.6 as a desktop agent. Reviewing the façade without the stack underneath produces content that ages in weeks, not years.
What We Tested
| Tool / Stack | Effective Model | Billing | Verdict | Why |
|---|---|---|---|---|
| Copilot + Fable 5 | Claude Fable 5 | AI Credits usage-based | ⚠️ | Capability is strong; billing is opaque until you pull usage logs |
| Claude API direct | Fable 5 / Opus 4.8 / Sonnet 4.6 / Haiku 4.5 | Explicit token ($1–$50 per MTok range) | ✅ | Pricing is transparent; best for teams that need to model cost per task |
| Cursor Composer / Auto | Vendor-managed | Usage split within Teams Premium | ⚠️ | UX is best in class; cost structure requires understanding Composer vs API split |
| Cursor Third-Party API | External provider | Separate usage bucket + possible passthrough | ⚠️ | Adds API cost on top of Cursor subscription |
| Devin Desktop | SWE 1.6 | Unverified | ⚠️ | Desktop agentic layer; pricing model not confirmed from official source this run |
| Claude Code | Vendor-managed model | Free / open CLI | ✅ | Strongest for terminal/CI/deploy; lowest barrier to entry for automation-first teams |
The Pivot Point
We ran a six-week benchmark where we treated each product as a monolith. The headline scores were useful but misleading: Copilot looked expensive, Cursor looked fast, Devin looked promising. When we broke down actual model and billing layers, Copilot’s cost dropped and Cursor’s cost climbed — the opposite of the summary table. We realized our review format was optimizing for shareability, not accuracy.
What We Use Now
Every evaluation starts with harness decomposition:
- Identify the actual model or models in use.
- Map billing to task categories (IDE chat vs API direct vs third-party route).
- Test the same three workflows on equivalent model tiers across products.
- Label synthesis vs first-party testing explicitly.
We no longer publish “best AI coding agent” without also publishing the stack assumptions behind it.
When You’d Choose Differently
If you need a single recommendation for a non-technical buyer, branding still matters for trust and support. We just make sure the buyer knows which stack they are actually buying.
Tool Crucible Rating
Overall / Ease / Value / Support — 1-5 each
- Overall: 4/5
- Ease: 3/5
- Value: 4/5
- Support: 3/5
This is part of our AI coding tool evaluation series. See full comparison: [link]
Last reviewed 2026-06-29. See our methodology and affiliate policy.