AI Tool Review · 2026

Zhipu AI Review (2026): Features, Pricing & Verdict

Zhipu AI — trading internationally as Z.ai — is the Chinese provider that did the thing everyone said open models couldn’t do: it took the coding crown. In May 2026, GLM-5.1 became the first open-weight model ever to top the SWE-bench Pro leaderboard, edging out GPT-5.4 and Claude Opus 4.6; a month later GLM-5.2 (June 2026) extended the lead — a ~753-billion-parameter Mixture-of-Experts model (~40B active) with a 1-million-token context window, vendor-reported scores of 62.1 on SWE-bench Pro and 81.0 on Terminal-Bench 2.1, and, crucially, MIT-licensed weights on Hugging Face with no regional or revenue restrictions whatsoever. Around that flagship, the Beijing lab (a 2019 Tsinghua University spinout, notable for training its models without Nvidia hardware) has built the most developer-friendly platform in the Chinese cohort: the only true Anthropic-compatible API endpoint outside Anthropic itself — making Z.ai a literal base-URL swap for Claude Code — plus an OpenAI-compatible endpoint, three genuinely free API models (GLM-4.7-Flash, GLM-4.5-Flash and the vision GLM-4.6V-Flash bill $0 for input, cache and output), metered flagship access at $1.40/$4.40 per million tokens (about a sixth of GPT-5.5), and the GLM Coding Plan — the flat-rate subscription that went viral as the cheap Claude Code backend, from $18/month list (currently $12.60 on promo). The honest counterweights: Zhipu has the cohort’s worst pricing-stability record (the $3/month early-adopter plan was killed in February 2026, GLM-5 launched with a ~30% price rise, and Coding Plan quota drains at 3x during peak hours), the headline benchmarks are vendor-reported pending independent verification, per-token rates are the highest of the Chinese big four, and the standard China-residency question applies to the first-party endpoints.

8.2
Overall Score / 10
The open coding flagship — the first open-weight model to top SWE-bench Pro, MIT-licensed with a 1M context, behind the friendliest endpoints in the cohort; docked for pricing churn, peak-hour quota games and vendor-reported benchmarks
Best for
Developers who want frontier-tier coding at open-weights prices — especially Claude Code users seeking a drop-in backend at a tenth of the cost — and teams that want an MIT-licensed flagship they can self-host commercially with zero restrictions
Platform
Z.ai open platform: OpenAI-compatible and Anthropic-compatible endpoints; GLM-5.2, GLM-5-Turbo, GLM-4.7 and free Flash models; prompt caching (~1/5 input rate); GLM Coding Plan flat-rate tiers for 20+ coding tools; free chat at chat.z.ai; MIT weights on Hugging Face (zai-org)
Key differentiator
Open-weight coding leadership — GLM-5.1 was the first open model to top SWE-bench Pro; GLM-5.2 (753B MoE, ~40B active, 1M context) posts 62.1 vendor-reported — plus the only Anthropic-protocol endpoint outside Anthropic, making Claude Code migration a base-URL swap
Pricing
GLM-5.2: $1.40 input / $4.40 output per 1M tokens (cached input $0.26); GLM-5-Turbo $1.20/$4.00; GLM-4.7 $0.60/$2.20; Flash models free. OpenRouter from ~$0.93/$3.00 across 25 providers. GLM Coding Plan: $18 / $72 / $160 per month list (30% intro promo; yearly discounts). Self-hosting free (MIT)
Vendor
Zhipu AI (Beijing, 2019, Tsinghua spinout) — China’s academic-lineage frontier lab, notable for Nvidia-free training and the GLM family’s open-weight benchmark coups
Platform notes (2026): four things to price in. The benchmarks are vendor-reported: the June 2026 GLM-5.2 table (SWE-bench Pro 62.1, Terminal-Bench 2.1 81.0/82.7, AIME 99.2, GPQA-D 91.2) was published by Zhipu and awaits neutral-harness verification — GLM-5.1’s leaderboard-topping run gives the claims credibility, but treat exact figures as claims, not settled fact. Coding Plan quota is dynamic: one “prompt” of quota drains at 3x during peak hours and 2x off-peak on GLM-5.2/GLM-5-Turbo (a 1x off-peak promo runs to September 2026); older models drain slower, plans hard-stop at the cap with no overage, and the plan only works inside officially supported coding tools — scripts and production apps must use the metered API. Pricing history argues for skepticism: the $3/month launch plan was withdrawn in February 2026, GLM-5 arrived with a ~30% per-token increase, and tier framing has shifted between quarterly and monthly billing — budget at list prices and re-check before renewals. Two compatibility modes, one better: both OpenAI and Anthropic endpoints work, but community consensus (and our reading) favours the Anthropic mode for reliability in agentic coding tools — it’s also the entire migration story for Claude Code users.

What Is Zhipu AI?

Zhipu AI is the Beijing frontier lab — founded in 2019 as a spinout from Tsinghua University, and internationally rebranded as Z.ai — whose GLM (General Language Model) family completed the most symbolically loaded arc in the 2026 model market: open weights catching, then passing, the closed coding frontier. The receipts are specific. In May 2026, GLM-5.1 scored 58.4% on SWE-bench Pro — the harder, contamination-resistant successor to SWE-bench Verified — making it the first open-weight model to top that leaderboard, ahead of GPT-5.4 (57.7%) and Claude Opus 4.6 (57.3%), while its SWE-bench Verified score (77.8%) sat within four points of Opus. Then on 13 June 2026, GLM-5.2 arrived: a Mixture-of-Experts flagship of roughly 753 billion total parameters with ~40 billion active per token, a 1-million-token context window with up to ~131K output tokens, and a vendor-published benchmark table headlined by 62.1 on SWE-bench Pro (ahead of the GPT-5.5 figure Zhipu cites at 58.6), Terminal-Bench 2.1 at 81.0 (82.7 with the best harness — within a few points of Claude Opus 4.8’s ~85), AIME 99.2 and GPQA-Diamond 91.2 — all self-reported, all pending independent verification, and all backed by the one fact nobody disputes: the weights landed on Hugging Face days later under a clean MIT licence with no regional or revenue restrictions. That licence choice matters doubly here because Zhipu is also a geopolitical curiosity — the lab trains without Nvidia hardware, a differentiator that has only grown sharper as export restrictions tightened through 2026 — and because it makes GLM-5.2 the strongest coding model anyone can legally self-host commercially, full stop. The platform around the models is the most Western-developer-friendly in the Chinese cohort: an OpenAI-compatible endpoint, and — uniquely — a true Anthropic-compatible endpoint (api.z.ai/api/anthropic), which means the same environment variables that point Claude Code at Anthropic can be redirected to GLM with a base-URL swap and zero code changes. That plumbing is what made the GLM Coding Plan the viral budget story of the year: a flat-rate subscription (Lite/Pro/Max at $18/$72/$160 per month list, with a 30% intro promo and yearly discounts to ~$151/$605/$1,344) bundling GLM-5.2, GLM-5-Turbo, GLM-4.7 and GLM-4.5-Air inside Claude Code, Cline, Roo Code and 20-plus other tools — 5-hour and weekly prompt quotas of roughly 80/400/1,600 prompts per five hours by tier, with quota equivalent to an estimated 15–30x the fee at API prices. Metered access prices the flagship at $1.40/$4.40 per million tokens (cached input $0.26, storage currently free), GLM-4.7 at a value-tier $0.60/$2.20, and three Flash models — including a vision model — at literally $0. Within our Model Providers & AI Infrastructure category, Zhipu is the coding specialist with the cleanest exit: not the price floor (DeepSeek), not the breadth king (Qwen), but the lab whose best model is open, whose endpoints speak both Western dialects, and whose subscription undercut the closed coding incumbents hard enough to become a phenomenon.

Core Features

GLM-5.2 and the open coding crown

Zhipu’s claim to this category is concentrated in one model line and one discipline, and the concentration is the strategy: GLM is built to be the best coding and agentic model you can actually own. The architecture follows the cohort’s playbook with bigger numbers — ~753B total parameters, ~40B active per token (a larger activation than DeepSeek’s 32B or MiniMax’s 10B, which buys capability at some serving cost), a 1M-token context window that matches DeepSeek and Qwen at the top of the market, and output headroom around 131K tokens that matters for exactly the whole-file-rewrite and long-plan workloads agentic coding generates. The benchmark story deserves both its trumpet and its asterisk. The trumpet: GLM-5.1’s SWE-bench Pro leaderboard win in May was independently ranked — the first time open weights led the hardest mainstream software-engineering benchmark — and GLM-5.2’s vendor table extends it (62.1 SWE-bench Pro, 81.0–82.7 Terminal-Bench 2.1, 76.8 MCP-Atlas for tool-protocol competence, 54.7 on Humanity’s Last Exam with tools), placing it within a few points of Claude Opus 4.8 on terminal work at roughly a sixth of the price. The asterisk: the 5.2 table is self-reported pending neutral-harness confirmation, and this category’s recent history (Moonshot’s contested Terminal-Bench harness) says wait for third-party numbers before treating decimals as truth — though GLM’s prior, verified leaderboard run makes these claims more credible than most. What the benchmarks translate to in practice, per the now-substantial community record: genuinely strong multi-file refactors, debugging and agentic tool use (the MCP-Atlas score reflects real Model Context Protocol competence), reasoning depth that holds up on math and science probes, and — the differentiator nobody else in the Chinese cohort matches — behaviour tuned to feel native inside Claude Code, because that is transparently the workload Zhipu optimised for. The honest capability boundaries: this is a coding-and-agents specialist, not a multimodal platform — vision exists (the free GLM-4.6V-Flash), but there’s no speech, video or image generation business here (MiniMax’s turf), no swarm orchestration layer (Moonshot’s), and general-assistant polish trails the Western flagships; output speed is respectable but unremarkable; and at $1.40/$4.40 the flagship is the most expensive of the Chinese big four per token — the pitch is frontier coding at a sixth of Western prices, not the absolute floor. For the workload it targets, though, the combination is unique in this review: leaderboard-credible coding, a million tokens of context, and weights you can take home.

The Coding Plan and the Anthropic-endpoint play

Zhipu’s commercial masterstroke isn’t a model — it’s a socket. The Anthropic-compatible endpoint (api.z.ai/api/anthropic) implements the Anthropic messages API contract, tool use and streaming included, which makes Z.ai the only provider anywhere that Claude Code treats as native: swap two environment variables and the world’s most popular agentic coding tool is running on GLM-5.2, no forks, no shims, no code changes. The GLM Coding Plan monetises that socket, and its economics are why it went viral as the budget Claude alternative. The structure: three tiers — Lite at $18/month, Pro at $72, Max at $160 list (a 30% intro promo currently prices them at $12.60/$50.40/$112, and yearly billing lands at roughly $151/$605/$1,344) — all bundling the same model lineup (GLM-5.2, GLM-5-Turbo, GLM-4.7, GLM-4.5-Air) inside 20-plus supported tools: Claude Code, Cline, Roo Code, OpenClaw, Kilo Code, OpenCode, Goose and friends. Quotas are prompt-based, not token-based — roughly 80/400/1,600 prompts per rolling 5 hours and 400/2,000/8,000 per week by tier, plus bundled MCP web-search/reader calls (100/1,000/4,000 monthly) — and Zhipu estimates the bundled quota at 15–30x the subscription fee in API-price terms, which is why heavy daily users almost always beat metered billing on it. Max at $160 undercuts ChatGPT Pro ($200) and Claude Max’s upper tiers while buying the highest prompt quota of any flat-rate agentic-IDE plan currently tracked. Now the fine print, which is where Zhipu’s character shows. Quota drains dynamically: GLM-5.2 and GLM-5-Turbo consume 3x quota during peak hours and 2x off-peak (a promo runs 1x off-peak through September 2026), so a “400 prompts per 5 hours” Pro tier is really ~133 flagship prompts at peak — a mechanism functionally similar to surge pricing that the marketing doesn’t lead with. Hit the cap and the plan simply stops — no overage, no auto-upgrade — and the plan is contractually restricted to supported coding tools: call the API from your own script or production app and you’re on metered billing, full stop. Add the history — the $3/month early-adopter plan withdrawn in February 2026, the ~30% per-token increase that accompanied GLM-5’s launch, tier framing that has shifted between quarterly and monthly — and the pattern is a lab that prices aggressively to acquire developers, then ratchets. None of it changes the bottom line that the Coding Plan is the cheapest credible frontier-coding subscription in the market; all of it says: take the promo, bank the savings, and re-run the maths at every renewal.

MIT weights, free tiers and the trust equation

Zhipu’s trust story is the strongest in the Chinese cohort on the axis that matters most — exit rights — and middling on the axes this cohort always struggles with. Start with the exit: GLM-5.2’s weights ship under a genuine, unmodified MIT licence — commercial use, modification, redistribution, self-hosting, no regional restrictions, no revenue thresholds, no authorisation clauses — which makes it categorically different from MiniMax’s commercially-gated modified-MIT and even a notch cleaner than Moonshot’s lightly-modified terms, and equal to DeepSeek’s gold standard. The consequence compounds with the capability claim: if GLM-5.2 is (or is near) the best open coding model, then the best coding model you can legally run on your own GPUs, for commercial purposes, anywhere on earth, is a Zhipu model — and 25 providers on OpenRouter (from ~$0.93/$3.00, undercutting first-party rates) plus hosts like DeepInfra and Fireworks mean Western-infrastructure inference is a routing decision, not a project. The free layer is unusually generous and worth naming precisely: GLM-4.7-Flash, GLM-4.5-Flash and the vision GLM-4.6V-Flash are listed at $0 for input, cached input and output on the API — genuinely free registered-user models, not trials — and chat.z.ai offers the flagship conversationally for nothing, which together make Z.ai arguably the cheapest serious prototyping platform in this review. Residency and jurisdiction follow the cohort’s standard analysis: Zhipu is a Beijing company, the first-party endpoints are its platform, and organisations that bar Chinese-parent hosted services will bar these too — with the mitigation being the best available anywhere: MIT weights on your infrastructure or a Western host’s, covering the actual flagship rather than a lesser sibling. Two Zhipu-specific trust notes round the picture. First, the Nvidia-free training stack: strategically it insulates GLM’s roadmap from export-control shocks better than any peer’s (a live concern in 2026), though it also means performance-per-watt and serving-cost claims ride on less-familiar silicon. Second, verification culture: Zhipu publishes benchmark tables promptly and in detail, but the June 5.2 numbers remain vendor-reported — the right posture is the one GLM-5.1 earned: credible claims from a lab with a verified leaderboard win, awaiting the neutral harness. The operational hygiene is the cohort’s usual, with emphasis on commercial terms rather than licence terms: prices here move (upward) with little ceremony, promos expire on calendar dates, and quota mechanics get retuned — pin model IDs, budget at list, and let the MIT weights be the insurance policy they genuinely are.

Scored Categories

Open weights & licensing (true MIT flagship, no restrictions, 1M context)

9.4

Coding & agentic capability (first open SWE-bench Pro leader; 5.2 extends it)

9.0

Compatibility & developer experience (only true Anthropic endpoint; OpenAI too)

8.8

Free tier (three $0 API models incl. vision; free flagship chat)

8.6

Price-performance ($1.40/$4.40 ≈ 1/6th of GPT-5.5; Coding Plan value)

8.4

Ecosystem & distribution (25 OpenRouter providers; 20+ Coding Plan tools)

8.2

Enterprise trust & residency (CN endpoints; vendor-reported benchmarks; MIT exit)

6.9

Billing & plan stability (price hikes, promo removals, peak-hour quota multipliers)

6.3

Pricing

Model / plan Price Notes
GLM-5.2 (flagship, MIT open weights) $1.40 / $4.40 per 1M tokens 753B MoE (~40B active), 1M context, ~131K output. Cached input $0.26 (~1/5 rate); cache storage currently free (limited-time). Vendor-reported SWE-bench Pro 62.1
GLM-5-Turbo $1.20 / $4.00 Speed-tuned sibling for latency-sensitive agentic work
GLM-4.7 (value tier) $0.60 / $2.20 Previous flagship — still excellent, and the practical sweet spot for routine coding via the metered API
Flash models Free ($0 input, cache and output) GLM-4.7-Flash, GLM-4.5-Flash and vision GLM-4.6V-Flash — genuinely free for registered users, not trials
Third-party hosting From ~$0.93 / $3.00 (OpenRouter) 25 providers; undercuts first-party list. DeepInfra, Fireworks and others serve the MIT weights on Western infrastructure
GLM Coding Plan — Lite $18/mo list ($12.60 promo; ~$151/yr) ~80 prompts per 5h, ~400/week, 100 MCP calls/mo. Supported coding tools only
GLM Coding Plan — Pro $72/mo list ($50.40 promo; ~$605/yr) ~400 prompts per 5h, ~2,000/week, 1,000 MCP calls/mo — the daily-driver tier
GLM Coding Plan — Max $160/mo list ($112 promo; ~$1,344/yr) ~1,600 prompts per 5h, ~8,000/week, 4,000 MCP calls/mo. Undercuts ChatGPT Pro ($200) and upper Claude Max tiers
Quota mechanics ⚠ 3x peak / 2x off-peak drain GLM-5.2 and 5-Turbo consume multiplied quota (1x off-peak promo to Sept 2026); hard stop at cap, no overage; plan restricted to supported tools — scripts and apps bill via API
Self-hosting Free (MIT) No regional or revenue restrictions — full commercial self-hosting of the actual flagship
Zhipu’s price list is honest line by line and slippery over time — four rules keep it in your favour. One: buy promos, budget list. The 30% Coding Plan intro discount and the 1x off-peak quota promo are real savings today; the $3/month plan’s February burial and GLM-5’s 30% launch bump are your evidence they won’t all survive renewal — plan at $18/$72/$160 and full multipliers. Two: match the plan to the tool. The Coding Plan only meters inside supported coding tools — if your workload is scripts, agents or a product backend, you’re an API customer at $1.40/$4.40 (or $0.60/$2.20 on GLM-4.7, which handles routine work at a fraction of flagship cost). Three: schedule around the multiplier. Peak-hour prompts on GLM-5.2 cost triple quota — batch heavy agentic sessions off-peak, route formatting and boilerplate to GLM-4.7 or the free Flash models, and a Lite plan stretches surprisingly far. Four: remember the two exits. OpenRouter’s 25 providers already undercut first-party rates (~$0.93/$3.00), and the MIT weights make self-hosting the ultimate rate lock. Verify current prices at docs.z.ai — this lab reprices more than any other in the cohort, and rarely downward.

Strengths

  • The open coding crown — GLM-5.1 was the first open-weight model to top SWE-bench Pro; GLM-5.2 extends the claim (62.1 vendor-reported)
  • True MIT licence on the actual flagship — unrestricted commercial self-hosting, the cleanest exit in the cohort alongside DeepSeek
  • The only Anthropic-compatible endpoint outside Anthropic — Claude Code migration is a base-URL swap
  • GLM Coding Plan — the cheapest credible frontier-coding subscription (from $18/mo list; Max undercuts ChatGPT Pro)
  • 1M-token context with ~131K output headroom — top of the market
  • Three genuinely free API models including vision, plus free flagship chat
  • Flagship at ~1/6th of GPT-5.5 pricing; GLM-4.7 value tier at $0.60/$2.20
  • 25 OpenRouter providers undercut first-party rates on Western infrastructure
  • Nvidia-free training stack — unusual insulation from export-control shocks

Weaknesses

  • Worst pricing-stability record in the cohort — $3/mo plan killed Feb 2026, ~30% GLM-5 price rise, shifting tier structures
  • Peak-hour quota multipliers (3x/2x) quietly shrink advertised Coding Plan allowances
  • GLM-5.2 benchmarks are vendor-reported, pending neutral-harness verification
  • Highest per-token flagship rates of the Chinese big four ($1.40/$4.40)
  • Coding Plan restricted to supported tools — no coverage for scripts or production apps
  • Coding-and-agents specialist — no speech, video or media stack; general polish trails Western flagships
  • Hard quota stops with no overage path mid-session
  • Standard China-residency question on first-party endpoints

Verdict: 8.2 / 10 — The Open Coding Flagship

Zhipu AI earns an 8.2 for pulling off the category’s most consequential stunt: making the best-in-class coding model something you can download. GLM-5.1’s verified SWE-bench Pro leaderboard win broke the assumption that open weights must trail the closed frontier, GLM-5.2’s 753B/1M-context follow-up (vendor-reported, credibly) extends it, and the MIT licence converts the achievement into property rights — the strongest coding model legally self-hostable for commercial use anywhere is, right now, a Zhipu model. The commercial packaging is just as sharp: the only Anthropic-protocol endpoint outside Anthropic turns Claude Code into GLM’s distribution channel, the Coding Plan undercuts every closed coding subscription that matters, three free API models make prototyping costless, and 25 third-party hosts already serve the weights on Western infrastructure below first-party rates. What holds it at 8.2 — above Moonshot, below the DeepSeek/Qwen/Llama trio — is character rather than capability: the cohort’s most consistent pattern of pricing churn (the buried $3 plan, the GLM-5 hike, quarterly-to-monthly reshuffles), quota mechanics whose peak-hour multipliers make advertised allowances a best case, flagship benchmarks still awaiting neutral verification, per-token rates that are the cohort’s highest, and a deliberately narrow footprint — coding and agents, brilliantly, and little else. The buying logic: if you live in Claude Code or any agentic IDE and your monthly bill stings, trial the Coding Plan this week — the base-URL swap takes ten minutes, the promo pricing is real, and for the standard majority of coding work the quality clears the bar. If you need an open model to build a commercial product on, GLM-5.2’s MIT weights are now the default recommendation for coding workloads, with DeepSeek and Qwen covering general text. If you need breadth — modalities, media, swarms — look at its siblings in this cohort. And whatever you buy, re-read the price page quarterly: this lab’s models keep getting better on a schedule, and its prices keep doing the same.

Frequently Asked Questions

Is the GLM Coding Plan really a Claude Code replacement?

For most coding workloads, yes — functionally and dramatically on price — with a capability ceiling and quota mechanics you should understand before cancelling anything. The mechanics first, because they’re the magic: Z.ai operates a genuine Anthropic-compatible endpoint that implements the Anthropic messages contract (tool use and streaming included), so Claude Code — which communicates via standard API calls rather than proprietary extensions — runs against GLM-5.2 with nothing more than redirected environment variables. No fork, no wrapper, no behavioural shims: the same applies to Cline, Roo Code, OpenClaw and the 20-plus other tools on the supported list. The economics are the headline: a Pro plan at $72/month list ($50.40 on the current 30% promo) buys roughly 400 prompts per rolling five hours and 2,000 per week — quota Zhipu estimates at 15–30x the fee in API-price terms — against $100–200/month for the closed coding subscriptions it displaces, and the Max tier ($160 list) undercuts ChatGPT Pro outright while carrying the largest prompt allowance of any flat-rate agentic-IDE plan currently tracked. The capability question is the honest hinge. On the benchmarks, GLM-5.2 is at or near the top of the coding leaderboards — its predecessor’s SWE-bench Pro win was independently ranked, and 5.2’s vendor-reported 62.1 leads the table Zhipu publishes, with Terminal-Bench within a few points of Claude Opus 4.8 — and community experience broadly matches: for the standard 80% of work (multi-file edits, test writing, refactors, debugging with tools), the quality difference from the closed incumbents ranges from imperceptible to tolerable. The residual gap shows at the hardest edge — subtle architectural judgement, gnarly multi-system debugging, and general-assistant polish — where the Western flagships retain an advantage worth paying for on the tasks that need it; the pragmatic pattern is a cheap GLM plan for volume plus metered access to a frontier model for escalations. Now the fine print that separates satisfied users from surprised ones: quota drains at 3x during peak hours and 2x off-peak on GLM-5.2/5-Turbo (the 1x off-peak promo ends September 2026), so your “400 prompts” tier delivers ~133 flagship prompts at peak — schedule heavy sessions accordingly, and route boilerplate to GLM-4.7, which drains quota slower; plans hard-stop at the cap with no overage (keep a fallback key configured); and the plan is contractually restricted to supported coding tools — the moment your usage is a script, an agent you built, or a production backend, you’re on the metered API instead. Verdict: as a Claude Code backend for an individual developer or small team, it’s the best value in the market and a ten-minute experiment; as your only model, it asks you to accept the open-frontier trade — 90-something percent of the capability for 10–20% of the money — which, in 2026, an awful lot of developers are accepting.

GLM vs DeepSeek vs Qwen — which open Chinese model wins for coding?

As of mid-2026 the honest answer is GLM for the coding crown itself, with DeepSeek and Qwen winning specific adjacent cases — and the three-way gaps are small enough that harness, price and licence should decide more than leaderboard decimals. The capability picture: GLM-5.2 holds the strongest coding claim — its line produced the first open-weight SWE-bench Pro leader (verified, in May) and the current vendor-reported top score (62.1), with Terminal-Bench results within a few points of the best closed models; DeepSeek’s V4 Pro sits in the same neighbourhood (~80 on SWE-bench Verified, a top-tier reasoning profile, and arguably the edge on hard single-shot logic); Qwen’s open shelf peaks differently — its 397B generalist is a top-three open model overall, and Qwen3-Coder-Next is the standout local coding model, delivering Sonnet-class coding from just 3B active parameters that run on consumer hardware. So: repository-scale agentic coding on serious infrastructure → GLM-5.2; hardest single-thread reasoning and the cheapest hosted tokens → DeepSeek; coding on a laptop or single GPU → Qwen’s Coder-Next, no contest. Licences — all three are genuinely open, with nuance: GLM-5.2 and DeepSeek V4 ship true MIT (unrestricted commercial self-hosting of the actual flagships); Qwen is Apache 2.0 across a vastly broader family but keeps its very best model (Qwen3.7-Max) closed — so “the best model I can own” is a GLM-or-DeepSeek question, while “an open model for every job” is Qwen’s. Price, hosted: DeepSeek is the floor ($0.14/$0.28 Flash; $0.435/$0.87 Pro, with free caching); Qwen’s ladder starts at $0.05 Flash with a $1.25/$3.75 promo flagship; GLM is the premium of the trio at $1.40/$4.40 first-party — though OpenRouter’s 25 hosts cut that to ~$0.93/$3.00, and the Coding Plan changes the unit economics entirely for IDE work, where none of the others has an equivalent this aggressive (Qwen’s coding plan exists but with harsher limits and suspension complaints; DeepSeek sells no subscription at all). Context: GLM and DeepSeek and Qwen’s flagship all reach 1M tokens (GLM’s with the largest output headroom); ecosystem: Qwen’s fine-tune universe is the deepest, DeepSeek’s simplicity is unmatched, GLM’s tool integration (the Anthropic endpoint, MCP-Atlas competence) is the most turnkey for agentic IDEs. The composite most teams land on in practice: GLM-5.2 (via Coding Plan or a Western host) as the daily coding engine; DeepSeek for cheap bulk text and the hardest reasoning escalations within the open world; Qwen weights for local, embedded and multimodal jobs — all three behind one router, because their shared OpenAI compatibility makes monogamy strictly optional.

Can Western businesses rely on Z.ai — and what’s the story with the Nvidia-free training?

The reliance question splits cleanly into the platform (normal Chinese-cohort caveats, plus Zhipu-specific pricing churn) and the models (the safest bet in the cohort, thanks to the licence) — and the Nvidia detail is a genuine strategic footnote rather than marketing. Platform first. Z.ai’s first-party endpoints are operated by a Beijing company, and the standard analysis applies unchanged: organisations whose policies restrict Chinese-parent hosted services — regulated finance, healthcare, government-adjacent, or anywhere counsel has drawn that line — will restrict these endpoints too, and should route accordingly; for everyone else, the practical hygiene is the usual (no secrets or regulated identifiers to any third-party LLM, contractual terms for anything sensitive). Zhipu adds one platform-specific risk its peers exhibit less: commercial churn. The record — a $3/month early-adopter plan withdrawn eleven weeks after going viral, a ~30% per-token increase accompanying GLM-5, billing structures reshuffled between quarterly and monthly, quota multipliers introduced and re-tuned — describes a lab that iterates its pricing as fast as its models. None of it is scandalous; all of it means enterprise buyers should contract prices rather than assume them, and hobbyist budgets should expect renewals to drift upward. The models are the opposite story: because GLM-5.2 ships under an unmodified MIT licence with no regional or revenue restrictions, the dependency risk that usually attends a fast-moving foreign vendor simply doesn’t apply — 25 third-party providers already serve the weights on US/EU infrastructure (below first-party prices), DeepInfra and Fireworks offer them under Western DPAs, and self-hosting for commercial use is unambiguously legal, making the worst-case scenario “migrate inference,” never “lose the model.” That is a materially stronger position than MiniMax (commercially-gated weights) and modestly stronger than Moonshot (lightly modified terms), equal to DeepSeek. The Nvidia-free angle deserves its precise weight: Zhipu trains GLM without Nvidia hardware — domestically-sourced accelerators — which in the 2026 export-control climate insulates its research roadmap from the supply shocks that could, in principle, disrupt peers dependent on restricted silicon; for a buyer, that translates to lower tail-risk on model continuity, not to any difference in the API you call or the weights you download (inference runs fine on whatever hardware your host uses). It also cuts the other way gently: less-proven training silicon means efficiency claims deserve the same wait-for-verification posture as the benchmarks. The bottom line for a Western business: use the Coding Plan or first-party API where your data policy permits and the price is contracted; use Western hosts of the MIT weights where it doesn’t; self-host where control matters most — and in all three configurations, the thing you’re actually depending on, the model, is the one asset in this cohort nobody can take back.