Third Shift Labs runs real tasks against your API with every major coding agent. We report where they succeed, where they stall, and where they quietly give up.
If we emailed you, your Agent Readiness Report is already built. Reply "send it" and it's yours.
No form. No call. The report is finished; it just needs an address.
A developer hands your platform to their coding agent. The agent reads your docs, hits a wall in your auth flow, retries a few times, and gives up. The developer does not open a support ticket. They do not post in your Discord. They conclude your product does not work, and they try your competitor, whose quickstart happened to parse cleanly.
All your analytics recorded was a handful of 401s at 2:14 in the morning.
The agent is now the first user of your API. It reads every word of your documentation, takes every error message literally, and churns without saying goodbye. You have never once watched it work.
Not toy prompts. Jobs a working developer would actually delegate: create a metered subscription, migrate a customer between plans, issue a partial refund against a disputed charge. Each task ran across four coding agents, three model tiers, and multiple reasoning settings.
Every transcript is public. Including the embarrassing ones. Especially the embarrassing ones.
| Task | Claude Code | Codex | Cursor | Gemini CLI |
|---|---|---|---|---|
| Create customer + default payment method | ||||
| Create a metered subscription | ||||
| Partial refund on a disputed charge | ||||
| Migrate a customer between plans (proration) | ||||
| Webhook endpoint + signature verification | ||||
| Usage-based invoice with line-item discounts | ||||
| Trial with payment-method collection | ||||
| Checkout session with dynamic tax | ||||
| Failed-payment retry schedule | ||||
| Export payouts + reconcile fees | ||||
| Connected account + funds transfer | ||||
| Test-clock renewal simulation |
Retried the same webhook signature verification eleven times before concluding the API was down.
The API was not down. The docs assume the raw request body; every framework example the agent could find had already parsed it.
The agent was very patient. A prospective customer would not have been.
For each platform, the ten to twelve tasks a developer would most plausibly hand to an agent. Real outcomes with ground truth, not "write hello world."
Every task runs on every agent (Claude Code, Codex, Cursor, Gemini CLI), across model tiers, at multiple reasoning-effort settings. Identical starting conditions each time: API key already in the environment, your docs exactly as they are today, no coaching. The benchmark starts at the first prompt.
Goal completion, verified against ground truth rather than vibes. Cost, in tokens and dollars. Wall-clock time. And friction: retries, wrong turns, and the exact paragraph of documentation or error message where a run died.
Full transcripts, failures included. If a claim in the report cannot be traced to a transcript, it does not go in the report.
If we emailed you, we already ran the full matrix against your platform: four agents, three model tiers, 144 real runs. Inside:
This is the document you forward when you have been saying "our docs need work" for two quarters and getting nowhere. Now you have receipts.
It reads like an audit of the docs, not of you. Most dead-ends we find are things someone inside already suspected and could never prove. Now you can prove it.
144 runs, about $430 in compute and 11 hours of agent wall-clock time. Already spent. The report is free anyway.
Reply "send it" to the email. That's the whole process.
The report is free. No strings, no drip sequence, no "just circling back on this."
That is what we sell. We rewrite the docs, error messages, and examples where agents dead-ended, then re-run the identical benchmark. The before-and-after is the proof, not our word.
What worked last quarter can quietly break. We re-run your benchmark and tell you before your users find out.
Yes, the second and third parts are the business. The report is real either way.
The third shift is the one that works while you sleep. That is when agents integrate against your API: no onboarding call, no sales engineer, nobody watching.
We watch.
Reply "send it" to the email you received. It was written while you slept. Reading it takes about ten minutes.
It reflects this month's models. When the next family ships, these numbers change.
No email from us? Write to [email protected] and we'll put your platform in the queue. We run a few platforms a month; the queue is real.