Updated just now.
Done
- ✅ Found that GLM (Z.AI's coding model) was a real, paid $18/month account that had never run a single job — 0 of 15,727 recorded calls, ever — because the tools that assign AI work never knew it was an option. Fixed and live-tested tonight; it now works correctly.
- ✅ Found the cost meter was pricing DeepSeek at roughly a third to a quarter of its real, published rate. Fixed for that model.
- ✅ Measured, across 7 days: 3,873 cheap-vendor calls, real monthly cost $103–308 (the meter had reported about $65), and 45% of paid dispatches — 98 of 218 — spent real money and produced nothing usable.
In progress
- ◐ None — this report is a completed measurement, not an active build.
Pending
- ⬜ The 45% failure-rate finding: better task instructions and pass/fail checks would recover roughly a fifth of total spend on every vendor at once. Recommended, not yet built.
- ⬜ A second vendor (Qwen's larger model) was found priced off a placeholder far ABOVE its real cost — the opposite error from DeepSeek. Not yet corrected.
- ⬜ Production still runs on a DeepSeek model name that doesn't appear in the vendor's current model list. Still works, but undocumented and could be repriced or retired without notice.
- ⬜ One real DeepSeek invoice, to settle whether a 3x price increase Nick approved the same night is a real increase at all — it rests on an unverified assumption about which pricing tier production was actually billed at.
Finish line
Every dollar the cheap-vendor lane spent, measured and reconciled against 12 lanes' own reports, with the meter itself fixed where it was found to be wrong.