AI Tool Review · 2026

MiniMax Review (2026): Features, Pricing & Verdict

MiniMax is the model provider most Western developers have used without knowing it — the Shanghai lab whose voices narrate half the AI-generated videos on your feed, whose Hailuo engine renders the clips themselves, and whose M-series text models have quietly become one of the best price-performance stories in the market. That’s the defining fact about MiniMax within this category: it isn’t a text-model company with side projects, it’s a genuinely multimodal house — text, speech, video, music and image APIs under one roof — anchored by an M-series that punches absurdly above its activation size. The architecture is the efficiency play of the cohort: 230 billion total parameters with just 10 billion active per token, which is how every text model from the original M2 (October 2025) through M2.7 (March 2026) to the new M3 flagship bills at the same flat $0.30/$1.20 per million tokens — roughly 5% of Opus-class output rates — while posting real numbers: M2 ranked among the top open-source models for composite intelligence on Artificial Analysis; M2.7 scores 56.2% on SWE-Pro, 57.0% on Terminal-Bench 2 and a category-defining 1495 ELO on GDPval-AA, the benchmark for real-world office work (live debugging, financial modelling, full Word/Excel/PowerPoint document generation); and M3 extends the line to an advertised 1M-token context with frontier coding-and-agent positioning. The honest counterweights are unusually specific: the licence regressed (M2 was true MIT; M2.7 moved to a modified MIT that bars commercial self-hosting without authorisation), output speed is the cohort’s slowest at ~48 tokens/second on standard routing (the 2x-price HighSpeed variant fixes it), the “permanent 50% off” M3 pricing invites the usual promo skepticism, caching bills cache writes above the standard input rate, and the China-residency question applies as it does across this cohort. This review covers the developer platform; Hailuo, the video product, is reviewed separately at 0379.

7.8
Overall Score / 10
The multimodal value house — frontier-adjacent text, leading speech and video, all at single-digit percentages of Western pricing; held back by a licence regression, slow default output and fiddly billing
Best for
Teams that want one cheap provider across modalities — agentic text plus best-in-class TTS and video generation — and cost-driven builders of office-work agents (documents, spreadsheets, debugging) where M2.7’s GDPval lead and $1.20/M output change the economics
Platform
MiniMax Open Platform (platform.minimax.io): M-series text models (M2, M2.7, M3) behind OpenAI-compatible endpoints with interleaved-thinking support; separate speech, video (Hailuo), music and image APIs; pay-as-you-go plus Token Plan subscriptions; weights on Hugging Face (licence varies by version)
Key differentiator
Extreme efficiency plus breadth — 230B-total/10B-active MoE prices every text model at $0.30/$1.20 flat, while the same company ships leading voice synthesis and the Hailuo video engine; M2.7’s 1495 GDPval-AA ELO leads real-world office-task benchmarks
Pricing
All M-series text: $0.30 input / $1.20 output per 1M tokens (M3 lists $0.60/$2.40 with a “permanent 50% off”); cache reads $0.06, cache writes $0.375; M2.7-HighSpeed 2x; M3 Priority tier 1.5x; M3 inputs above 512K bill double. Token Plan subscriptions with 5-hour/weekly quotas. Speech, video, music, image on separate meters
Vendor
MiniMax (Shanghai, founded 2021) — one of China’s original “AI tiger” startups, consumer-proven through the Hailuo apps and Talkie/Xingye companions, now a full-stack multimodal API house
Platform notes (2026): five specifics to handle before production. The licence changed under your feet: M2 shipped under true MIT (commercial self-hosting fine); M2.5/M2.7 moved to a modified MIT that permits free personal use but bars commercial deployment of the weights without MiniMax authorisation — commercial self-hosters must stay on M2, pay for the API, or negotiate; check M3’s terms before assuming anything. Speed needs a decision: standard M2.7 streams at ~48 tokens/second — noticeably slower than DeepSeek or Kimi — and the fix (M2.7-HighSpeed, ~100 tok/s) costs exactly 2x; M3 instead offers a Priority service tier at 1.5x. Caching has a write toll: cache reads are cheap ($0.06/M) but cache writes cost $0.375/M — more than standard input — so caching only pays when prefixes are reused enough times to amortise the population cost. M3’s 1M context is tiered and gated: inputs up to 512K bill $0.30/$1.20; above 512K everything doubles, and the upper band is currently access-limited (contact sales). Preserve reasoning between turns: the M-series uses interleaved thinking, and MiniMax explicitly warns that dropping reasoning content between turns degrades performance — pass reasoning_details back in multi-turn agent loops, and patch tool-schema edge cases before scale.

What Is MiniMax?

MiniMax is the Shanghai AI company, founded in 2021 and one of the original “AI tiger” startups of China’s post-ChatGPT boom, that took the road nobody else in this category took: instead of betting everything on one text model, it built a consumer-proven engine for every modality and then opened the whole stack to developers. Most Western users met MiniMax without the name — through Hailuo, the video generator whose speed-and-physics act we scored highly at 0379; through its speech models, widely regarded as some of the best text-to-speech in the world and the default narrator of a substantial share of AI-generated content; or through its companion apps (Talkie internationally, Xingye in China) that proved the company’s conversational chops at consumer scale. The text line — the M-series — is what earns MiniMax its slot in this category, and its story is an efficiency thesis executed consistently. M2 (23 October 2025) set the template: a 230-billion-parameter Mixture-of-Experts model activating just 10 billion per token, optimised end-to-end for coding and agentic work — multi-file editing, compile-run-fix loops, test-validated repair — with strong SWE-Bench Verified, Multi-SWE-Bench and Terminal-Bench results, competitive agentic scores on BrowseComp and GAIA, and a top-open-source ranking for composite intelligence from Artificial Analysis, all at $0.255–0.30 input/$1.00–1.20 output per million tokens with a ~205K context. The line then iterated fast and stayed cheap: M2 Her (January 2026) spun the conversational-EQ heritage into a dedicated companion model; M2.7 (18 March 2026) became the workhorse — same 230B/10B skeleton, sharper agentic training with multi-agent collaboration, and a benchmark profile aimed squarely at real-world productivity: 56.2% on SWE-Pro, 57.0% on Terminal-Bench 2, and a 1495 ELO on GDPval-AA, the evaluation built around actual office deliverables (live debugging, root-cause analysis, financial modelling, full document generation across Word, Excel and PowerPoint) where M2.7 set the standard for multi-agent systems; and M3, the current flagship, extends to frontier coding-and-agent positioning with an advertised 1M-token window (tiered above 512K) at the same effective $0.30/$1.20 under a labelled-permanent 50% discount. Distribution is broad — an OpenAI-compatible first-party API, a dozen-plus providers on OpenRouter (where M2.7 lists at $0.24/$0.96 and spot markets have driven it lower still), and weights on Hugging Face, though the licence tightened from M2’s true MIT to a modified MIT barring commercial self-hosting from M2.5 onward. Around the text line sit the other meters: speech, video, music and image APIs, each billed on its own terms — making MiniMax the only provider in this category where “add narration and a video clip to the agent’s output” is a same-platform decision. Within our Model Providers & AI Infrastructure category, MiniMax is the multimodal value house: not the price floor (DeepSeek), not the breadth king (Qwen), not the swarm specialist (Moonshot), but the one lab whose answer to text, voice and video is the same invoice.

Core Features

The M-series: office-work agents at 5% of flagship cost

The M-series’ pitch is precision-targeted: it isn’t trying to win abstract reasoning crowns — it’s built to do jobs, and the benchmark that defines it is the one closest to actual employment. GDPval-AA evaluates models on real-world knowledge-work deliverables — debugging live systems, root-cause analysis, financial modelling, generating complete Word documents, Excel workbooks and PowerPoint decks — and M2.7’s 1495 ELO set the standard for multi-agent systems on it, which is a more commercially meaningful claim than another decimal point on a math olympiad. The supporting numbers hold the frame: 56.2% on SWE-Pro (the harder autonomous-engineering benchmark where even flagships live in the 50s–60s), 57.0% on Terminal-Bench 2, an 87.4 GPQA score placing it in the 94th percentile, and M2’s Artificial Analysis composite ranking among the top open-source models — with the agentic fundamentals (long-horizon planning, retrieval, recovery from execution errors on BrowseComp and GAIA) trained in from the start rather than prompted on top. M2.7’s multi-agent collaboration layer echoes Moonshot’s swarm thesis at smaller scale: the model plans, executes and refines across cooperating instances, which is precisely what the document-generation workflows exercise. Now the economics that make it interesting rather than merely good: the 230B-total/10B-active architecture means inference costs like a 10B model, so the entire M-series bills $0.30 input/$1.20 output per million tokens — about 5% of Opus-class output rates, cheaper on output than Kimi K2.6 ($2.50) and undercut on input only by DeepSeek’s V4 Flash — and the version number changes capability, never price. An agent generating a hundred client reports a day, or a support system producing full spreadsheets on demand, runs at costs that would be a rounding error on a Western flagship. The honest capability boundaries: this is frontier-adjacent, not frontier — the hardest single-shot reasoning still belongs to the closed Western models and DeepSeek’s V4 Pro; the ~205K context (M2/M2.7) trails DeepSeek’s and Qwen’s million-token windows, and M3’s advertised 1M is tiered (double price above 512K) and gate-limited at the top band; standard output speed of ~48 tokens/second is the slowest in the Chinese cohort — fine for background agents, awkward for user-facing streaming without the 2x HighSpeed variant; and the interleaved-thinking design imposes a real integration requirement: reasoning content must be passed back between turns (via reasoning_details on routers) or multi-turn performance degrades, a quirk that has tripped more than one naive OpenAI-SDK port, alongside documented tool-schema edge cases. Build for what it is — a tireless, absurdly cheap office worker with a slight stutter — and the M-series is one of the best deals in this review.

The multimodal stack: one vendor, five meters

MiniMax’s structural advantage over every text-first rival in this category is that the text API is one aisle in a larger store — and for a growing class of products, that consolidation is the decision. The inventory: speech — MiniMax’s text-to-speech and voice-cloning models are widely regarded as among the best in the world, with naturalness and multilingual range that made them the de facto narrator of AI-generated video content well before most users knew the vendor’s name; video — Hailuo, reviewed separately at 0379, is the budget speed-and-physics champion of the generative-video market, and its API lives on this same platform; music and image generation round out the creative set, each billed on its own meter (per-character or per-second or per-image, not tokens — read each price page). Why this matters architecturally: the canonical 2026 product — an agent that researches, writes, narrates and illustrates — normally requires stitching three or four vendors (a text API, ElevenLabs-class TTS, a video model, an image model) with separate keys, invoices, rate limits and compliance reviews. On MiniMax it’s one platform, one account, and text-model outputs that were explicitly trained toward document and media production feeding sibling APIs. The consumer provenance is the quality assurance: these aren’t checkbox features — Hailuo and the companion apps (Talkie/Xingye) run at consumer scale, so the speech and video engines have survived contact with millions of users, and M2 Her (January 2026) productised the emotional-conversation heritage into a dedicated 66K-context companion model at the standard $0.30/$1.20 — a niche no other provider in this category serves as a first-class API. The platform machinery beneath: OpenAI-compatible endpoints, two parallel billing systems (pay-as-you-go on standard keys; Token Plan subscriptions on separate keys with fixed quotas managed in 5-hour rolling and weekly windows, unused quota not carried over), a Priority service tier on M3 (1.5x for admission-priority latency), and prompt caching with cheap reads ($0.06/M) but a write cost ($0.375/M) that exceeds standard input — so caching pays only on genuinely repeated prefixes. The honest limits of the breadth claim: the modality APIs are siblings, not a fused any-to-any model — you orchestrate calls across them rather than prompting one omnimodel; the consumer apps (MiniMax Agent, Hailuo) run on entirely separate credit subscriptions that confuse first-time buyers; and Qwen technically lists more model IDs — but Qwen’s media models are catalogue entries, while MiniMax’s voice and video are market leaders in their own reviews. For media-adjacent builders, that’s the distinction that decides.

Licensing, speed and trust: reading the fine print

MiniMax’s fine print carries more signal than most in this cohort, and three clauses deserve the spotlight. First, the licence regression — the most consequential and least advertised change: M2 shipped under a true MIT licence, full commercial self-hosting included, and became a legitimate open-weights citizen on that basis; from M2.5 onward (including M2.7), the weights moved to a modified MIT that permits free personal use but bars commercial deployment without separate authorisation from MiniMax. The practical fork: commercial self-hosters can stay on the older MIT-licensed M2/M2.5-era terms they already hold, route commercial traffic through the paid API (which is what the licence is designed to encourage), or negotiate authorisation — and any architecture that assumed “MiniMax weights are MIT” needs re-reading before the next deployment. Set against DeepSeek’s unqualified MIT and Moonshot’s lightly-modified MIT on their flagships, this is the weakest openness posture of the Chinese big four, and it moves MiniMax’s open-weights story from “portability guarantee” to “evaluation privilege.” Second, speed and its price: ~48 tokens/second standard output is a real product constraint — background agents won’t care, user-facing chat will — and the remedies are explicit line items (M2.7-HighSpeed at exactly 2x; M3 Priority at 1.5x), which means latency-sensitive workloads should be budgeted at the higher rates from day one rather than discovering the markup later. Third, the promo-price pattern: M3’s headline $0.30/$1.20 is documented as a “permanent 50% off” against a $0.60/$2.40 list — and this category’s recent history (Qwen’s expiring flagship promo, the April free-tier purges, versus DeepSeek’s discounts-become-list-price) argues for budgeting at list until “permanent” survives a few quarters. On residency and trust: MiniMax is a China-headquartered vendor, and the standard cohort analysis applies — organisations barring Chinese-parent hosted platforms will bar the first-party API; the mitigation is thinner than peers’ because the modified licence complicates the self-host route for commercial use, leaving Western aggregators and OpenRouter’s dozen-plus providers (M2.7 from $0.24/$0.96, with spot markets lower) as the main non-China inference path. The credibility column is genuine: consumer-scale battle-testing across Hailuo and the companion apps, a text line iterated four times in six months without a price increase, top-percentile independent benchmark placements, and OpenRouter traffic that confirms real production adoption. The posture to take: treat MiniMax as an outstanding API vendor with a leading multimodal stack — and unlike its peers, do not count the open weights as your exit strategy unless your use is personal or your lawyers have read the modified licence.

Scored Categories

Multimodal breadth (leading TTS + Hailuo video + music + image, one platform)

9.2

Price-performance (flat $0.30/$1.20 across the M-series; ~5% of Opus output rates)

9.0

Efficiency architecture (230B total / 10B active; low latency-cost deployment)

8.5

Agentic & coding capability (GDPval-AA 1495; SWE-Pro 56.2; Terminal-Bench 57.0)

8.3

Developer experience (OpenAI-compatible; interleaved-reasoning requirement)

7.5

Enterprise trust & residency (consumer-proven; China HQ; slow default output)

6.8

Billing predictability (“permanent” promo, speed surcharges, cache write toll, dual keys)

6.6

Open weights & licensing (MIT → modified-MIT regression bars commercial self-hosting)

6.5

Pricing

Model / item Price (per 1M tokens) Notes
MiniMax M3 (flagship) $0.30 input / $1.20 output (≤512K) ⚠ Listed $0.60/$2.40 with “permanent 50% off.” Inputs above 512K bill double ($0.60/$2.40) and the top band is currently gated (contact sales). Advertised 1M context. Priority tier = 1.5x
MiniMax M2.7 (workhorse) $0.30 / $1.20 230B MoE, 10B active, 205K context. SWE-Pro 56.2%, Terminal-Bench 2 57.0%, GDPval-AA 1495 ELO. ~48 tok/s standard
M2.7-HighSpeed $0.60 / $2.40 (2x) ~100 tok/s — same weights, faster routing; budget this rate for user-facing streaming
M2 / M2.5 (legacy) & M2 Her $0.30 / $1.20 Same flat rate across the whole M-series — version changes capability, not price. M2 Her: 66K-context companion model. M2 retains the true MIT licence
Prompt caching Reads $0.06 · writes $0.375 ⚠ Cache writes cost more than standard input — caching pays only on well-reused prefixes. Legacy models read at $0.03; M3 reads double above 512K
Third-party hosting From $0.24/$0.96 (OpenRouter) A dozen-plus providers; spot markets have driven M2.7 lower still. Main non-China inference path given the licence terms
Token Plan subscriptions Fixed monthly quotas Separate subscription keys; quota managed in 5-hour rolling + weekly windows, unused quota does not roll over
Speech / video / music / image APIs Separate meters Per-character, per-second or per-asset billing — check each product page. Hailuo and MiniMax Agent consumer apps use their own credit subscriptions
MiniMax’s rate card looks like the simplest in the cohort — one flat text price — and the complications live in the adjectives. One: budget at list. The M3 rate is a labelled-permanent 50% promotion; this category’s track record says plan at $0.60/$2.40 and treat the discount as upside until it has survived a few quarters. Two: decide your speed tier up front. If users watch the tokens stream, you’re realistically a HighSpeed (2x) or Priority (1.5x) customer — price the product on that rate, not the headline. Three: do the caching arithmetic. With writes at $0.375/M against $0.30 standard input, a cached prefix must be read several times before it saves anything — great for stable agent system prompts, a net loss for one-shot workloads. Four: keep the two key systems straight. Pay-as-you-go and Token Plan quotas use separate keys, and subscription quota expires on 5-hour/weekly windows without rollover — meter both, and remember the consumer apps bill separately again. And across everything: pass reasoning back between turns per the interleaved-thinking requirement, or you’ll pay full price for degraded outputs. Verify current rates at platform.minimax.io — this lab reprices as fast as it ships.

Strengths

  • True multimodal house — frontier-adjacent text plus world-class TTS, the Hailuo video engine, music and image on one platform
  • Flat $0.30/$1.20 across the entire M-series — ~5% of Opus-class output rates, cheaper output than Kimi
  • GDPval-AA leadership (1495 ELO) — the best benchmark story in real-world office deliverables
  • Strong agentic engineering numbers — SWE-Pro 56.2%, Terminal-Bench 2 57.0%, top open-source composite intelligence heritage
  • Extreme efficiency — 230B capacity at 10B-active inference cost
  • Consumer-scale battle-testing via Hailuo and companion apps
  • OpenAI-compatible API; a dozen-plus third-party providers with spot pricing below list
  • M2 Her — the only first-class companion/EQ model in this category
  • Fast iteration without price increases — four text releases in six months

Weaknesses

  • Licence regression — M2.5+ weights bar commercial self-hosting without authorisation (M2 was true MIT)
  • Slowest default output in the cohort (~48 tok/s); the fix costs 1.5–2x
  • “Permanent 50% off” flagship pricing invites promo skepticism — budget at list
  • Cache writes ($0.375/M) exceed standard input — caching needs arithmetic, not assumptions
  • M3’s 1M context is tiered (2x above 512K) and gated at the top band
  • Interleaved-thinking requirement and tool-schema edge cases break naive SDK ports
  • 205K context on the workhorse trails DeepSeek’s and Qwen’s 1M windows
  • China HQ residency question, with a thinner self-host mitigation than peers due to the licence
  • Dual billing systems and no-rollover subscription quotas confuse first-time buyers

Verdict: 7.8 / 10 — The Multimodal Value House

MiniMax earns a solid 7.8 as the provider whose whole is its argument: no one else in this category sells frontier-adjacent text, world-class speech, a leading video engine and creative media under one roof at these prices — and for the products 2026 is actually building, agents that produce documents, narration and clips, that consolidation is worth real money and real integration time. The text line justifies its slot on merit alone: the 230B/10B efficiency play delivers a flat $0.30/$1.20 across every M-series model, and M2.7’s benchmark profile — GDPval-AA leadership on real office deliverables, mid-50s scores on the hardest engineering benchmarks — makes it arguably the best pound-for-pound office-work agent in the market, with M3 extending the ceiling. What holds it at 7.8, two tenths under Moonshot and half a point under the DeepSeek/Qwen tier, is a fine-print tax its peers don’t charge: the licence regression from MIT to commercial-restricted modified-MIT weakens the one guarantee that makes Chinese providers safe to depend on — the self-host exit — and forces Western commercial buyers onto the API or aggregators; the slow default output makes the honest price 1.5–2x the headline for anything user-facing; and the billing surface (promotional “permanent” discounts, cache-write tolls, tiered long-context rates, dual key systems with expiring quotas) demands more diligence per dollar than DeepSeek’s brutalist simplicity. The buying logic: if your product spans modalities — and especially if it already uses Hailuo or MiniMax voices — the M-series is the obvious text layer and the platform consolidation is the win. If you’re a pure-text buyer optimising for price, DeepSeek still undercuts it; for swarm-style agents, Moonshot out-specialises it; for open weights you can build a business on, Qwen and DeepSeek keep promises MiniMax has walked back. But as a single vendor for the multimedia agent era — cheap brains, the best voices, and video on the same invoice — MiniMax has no direct competitor at all, and that’s a category of one worth 7.8 and a serious trial.

Frequently Asked Questions

MiniMax vs DeepSeek vs Kimi — which cheap Chinese model should power my agents?

These are the three serious contenders in the cheap-but-capable tier, and the choice resolves cleanly once you name your workload’s shape — price alone won’t separate them, because all three are 90%+ cheaper than Western flagships. Raw cost: DeepSeek wins outright — V4 Flash at $0.14/$0.28 undercuts MiniMax’s $0.30/$1.20 and Kimi’s $0.60/$2.50, its 98%-off automatic caching has no write toll (MiniMax charges $0.375/M to populate cache; DeepSeek charges nothing), and V4 Pro at $0.435/$0.87 handles the hard-reasoning tier with the largest context of the three (1M flat-rate, versus Kimi’s ~260K and MiniMax’s 205K on the workhorse). If your workload is output-heavy text generation or RAG with stable context and you don’t need specialised agent training, V4 Flash is the cheaper hammer, full stop. Agentic office work: MiniMax M2.7’s case — its GDPval-AA leadership (1495 ELO) on real deliverables (debugging, financial models, full Word/Excel/PowerPoint generation), SWE-Pro 56.2% and Terminal-Bench 57.0% reflect training aimed at exactly the “AI employee” workloads, and at $1.20/M output (versus Kimi’s $2.50) it’s the cheapest of the three per generated document. Wide, parallel agent tasks: Kimi’s Agent Swarm — RL-trained orchestration of hundreds of concurrent sub-agents with measured gains on wide-research benchmarks — is a capability neither rival ships, and K2 Thinking’s 300-step runs anchor the deepest research workloads. Multimodal: MiniMax by forfeit — DeepSeek is text-only, Kimi adds vision input, but only MiniMax puts leading TTS, video, music and image on the same platform as its text models. Speed: Kimi and DeepSeek stream comfortably faster than MiniMax’s ~48 tok/s standard tier; MiniMax’s fix costs 2x, which narrows its price lead for user-facing products. Openness (your exit strategy): DeepSeek’s MIT flagship weights are the gold standard, Kimi’s modified-MIT flagship is nearly as good, and MiniMax is the laggard — its post-M2 weights bar commercial self-hosting without authorisation, so its “open” models are evaluation assets, not portability guarantees. Compliance routing is similar for all three (China-parent vendors; Western hosts or aggregators for restricted data), but note the licence makes MiniMax’s non-API paths narrowest. The composite recommendation most teams land on: DeepSeek for text volume and hard single-thread reasoning, MiniMax for document/office agents and anything touching voice or video, Kimi for swarm-shaped research and long autonomous runs — behind one OpenAI-compatible router, because the entire point of this cohort’s compatibility shims is that you never have to choose just one.

Can I self-host MiniMax models commercially?

Only with care, and for anything after M2 the honest default answer is no — this is the single most important fine-print item in the MiniMax ecosystem, and it changed recently enough that much of the internet’s advice is stale. The history: MiniMax M2 (October 2025) shipped its weights under a true MIT licence — use, modify, commercialise, redistribute, no strings — and on that basis entered the open-weights conversation alongside DeepSeek and Qwen, with self-hosted commercial inference fully legitimate. Then the terms tightened: from M2.5 onward, and explicitly including M2.7, the weights ship under a modified MIT licence that permits free personal use but bars commercial deployment without separate authorisation from MiniMax. What that means in practice, case by case. Personal projects, research, internal evaluation: fine — download from Hugging Face, run under vLLM, benchmark against your workload; this is exactly what the licence is designed to allow, and the free trial credits plus personal-use terms make a thorough proof of concept genuinely free. Commercial products on self-hosted M2.7/M3-era weights: not without paper — you need either authorisation from MiniMax (a sales conversation, effectively converting you into a licensing customer) or a different architecture. The three compliant routes: (1) stay on M2 — the original MIT-licensed weights remain under the terms they shipped with, and for many workloads the older model is still excellent value, just mind the capability gap against M2.7; (2) use the paid API — first-party or via OpenRouter’s dozen-plus providers (who host under their own commercial arrangements with MiniMax), which is transparently what the licence change is engineered to encourage, and at $0.30/$1.20 the API is cheap enough that the economics rarely justify fighting; (3) negotiate authorisation if your scale or data constraints genuinely demand on-premises deployment. Two broader implications worth internalising: first, do not carry the “Chinese open weights = residency escape hatch” heuristic from DeepSeek and Qwen over to MiniMax — for commercial use, the self-host mitigation that neutralises the China-HQ question elsewhere in this cohort is contractually narrowed here, leaving Western-hosted API aggregators as the main compliant non-China path; second, check the licence on every new release individually — M3’s terms should be read, not assumed, and a vendor that has moved the line once can move it again in either direction. The one-sentence answer: evaluate freely, self-host commercially only on M2 or with authorisation, and treat the API as the intended front door — which, at these prices, it convincingly is.

Why is MiniMax so cheap, and where do the hidden costs appear?

The cheapness is architectural and strategic — real, not a teaser — but MiniMax’s bill has more moving parts than its one flat headline rate suggests, and knowing the five places costs appear is the difference between the advertised economics and a surprise. Why the price is possible: the M-series activates just 10 billion of its 230 billion parameters per token, so serving costs scale like a small model while capability draws on a large one — the same MoE logic as DeepSeek and Kimi, executed at an even leaner activation ratio — and MiniMax subsidises the text line as the anchor of a multimodal platform whose speech and video meters carry their own margins, with the standard Chinese-cohort strategy of underpricing Western incumbents to buy global developer mindshare. The quality is independently corroborated (top open-source composite-intelligence rankings, 94th-percentile GPQA, GDPval-AA leadership, real OpenRouter production traffic), so you are not paying with capability. You pay, instead, in specifics. One: speed, the honest first surcharge — ~48 tokens/second standard output is fine for background agents and painful for user-facing streaming, and the remedies are priced (M2.7-HighSpeed at exactly 2x, M3 Priority at 1.5x), so latency-sensitive products should treat $0.60/$2.40 as their real rate. Two: the promo asterisk — M3’s $0.30/$1.20 is documented as a “permanent 50% off” against $0.60/$2.40 list; DeepSeek’s precedent (discounts becoming genuine list prices) is encouraging, Qwen’s (promos expiring) is cautionary, and prudent budgeting uses list until “permanent” has a track record. Three: the caching write toll — populating the cache costs $0.375 per million tokens, more than the $0.30 standard input rate, so caching is a wager that each prefix will be read enough times to amortise the entry fee: excellent for a stable agent system prompt hit hundreds of times, a net loss for one-shot or rapidly-changing context (contrast DeepSeek, where caching is free money with no write charge). Four: long-context tiering on M3 — inputs beyond 512K double every rate including cache reads, and the top band is gated behind sales, so the advertised 1M window is really a 512K window with an expensive annex. Five: quota mechanics — Token Plan subscriptions meter through separate keys with 5-hour rolling and weekly windows and no rollover, meaning unused quota simply evaporates; heavy-but-bursty users often do better on pure pay-as-you-go. Add the engineering costs that aren’t invoiced — the interleaved-reasoning requirement (pass reasoning back between turns or quality degrades) and tool-schema edge cases — and the honest summary is: MiniMax’s floor price is genuine and its ceiling surprises are all documented; read the pricing page like a contract, budget at list with your true speed tier, and the economics remain among the best in the market.