D

EP-016 · Tool deep-dives

MiniMax M2 for agencies: cheap enough to default?

Research Stage · Research Digest v2 drafted Angle undecided

Research digest · v2

Compiled from 2 web-research runs + 3 sources you added · updated Jul 12 · pick an angle below to unlock the brief draft

Finding 1 · Sticker pricing

M2 lists at $0.30 per 1M input tokens and $1.20 per 1M output — roughly 8–12% of frontier-model list pricing12. Replayed against our June client workload (≈38M tokens), the month's model bill drops from about $610 to roughly $705.

Finding 2 · Capability shape

On agent-style work M2 lands within reach of the frontier — ≈69% on SWE-bench Verified and ≈77 on τ²-Bench tool use — but it drops off faster on long multi-step chains, which is exactly the shape of agency deliverable work3.

Finding 3 · What agencies actually spend

Real agency AI spend clusters far below what vendor pricing pages imply: the 2026 survey median for small agencies is $340/month all-in, and model API fees are typically under a third of that — seats and tooling dominate4. Our own spreadsheet agrees: API spend was 28% of the June AI line5.

Finding 4 · The thinking-token catch

M2 bills its interleaved reasoning as output tokens, so real jobs burn 2–3× the output of a non-reasoning model16. Cost per completed task — not per token — is the honest comparison, and it narrows the sticker gap from ~10× to ~4–6×.

  1. 1MiniMax M2 pricing & API docs — web research · Jul 12
  2. 2MiniMax M2 launch post — added by you
  3. 3SWE-bench Verified + τ²-Bench roundup — web research · Jul 12
  4. 4"What agencies pay for AI in 2026" — agency survey · web research · Jul 12
  5. 5Client cost spreadsheet — our real spend — added by you
  6. 6Fireship: MiniMax M2 in 100 seconds — added by you

Three possible angles

None locked — the pick carries into the brief
Angle A — Default-to-cheap: when M2 is enoughAI-recommended A decision rule agencies can run: which client jobs the cheap default genuinely covers — grounded in findings 1 and 3.
Angle B — The hidden costs of cheap models Counter-programming: thinking-token bills, retries and review time eat the sticker discount (finding 4).
Angle C — M2 vs Claude for agency deliverables Head-to-head on three real deliverables — same brief, same refs, every cost shown on screen.

Open questions

3 unresolved
  • Do we have real latency numbers from our own runs?
  • Is the launch pricing locked, or does it step up after the promo window?
  • Where does M2 fail first on client deliverables — loud retries or quiet quality loss?

Refine with AI

Each turn updates the digest and its sources
D
Find real agency spend data, not vendor pricing pages.
Added 4 sources; digest updated — see finding 3.