AI Tool Review · 2026

Alibaba Qwen Review (2026): Features, Pricing & Verdict

Alibaba Qwen is the most complete model family in AI — and in 2026 that word “complete” is doing real work. No other provider, Western or Chinese, open or closed, ships top-tier entries across as many fronts at once: Apache 2.0 open weights spanning text, coding, vision and reasoning (the Qwen3.5-397B generalist sits in the open-model top three, and Qwen3-Coder-Next — just 3B active parameters — trades blows with Claude Sonnet-class models on code); a closed flagship, Qwen3.7-Max (May 2026), that scores 80.4 on SWE-bench Verified and cracked the global top ten of the Artificial Analysis Intelligence Index at launch with a 1-million-token context; and a hosted catalogue of over 145 model IDs — chat, coding, vision, embeddings, reranking, speech, image and video generation — behind OpenAI-compatible and Anthropic-protocol endpoints on Alibaba Cloud Model Studio (the platform formerly branded DashScope). The pricing is the other half of the pitch: the flagship costs $1.25/$3.75 per million tokens on its current promotion (list $2.50/$7.50), Qwen-Plus runs $0.40/$1.20, and Qwen-Flash starts at $0.05 input — a fraction of Western flagship rates, with batch calls at half price and roughly 9x input savings versus Opus-class models at the mid tier. The counterweights are real, and this review prices them: billing is the most complicated in the category (tiered rates by prompt length, region-dependent pricing, thinking-mode output billed at 3–10x, promotional prices that expire), the beloved free developer tier died on April 15, 2026, Qwen’s models are famously verbose (which claws back savings), and the flagship’s weights — unlike the rest of the family — stay closed. This review covers the developer platform and model family; the free consumer chat app is reviewed separately at 0013.

8.3
Overall Score / 10
The everything provider — the broadest open-weights family in AI plus a genuine frontier flagship at a fraction of Western prices; held back by the most complicated billing in the category, free-tier churn and a closed flagship
Best for
Developers who want frontier-tier open weights under a true Apache 2.0 licence — especially local and self-hosted coding — plus cost-driven teams using the hosted API for multilingual, coding and long-context work at 6–10x below Western flagship prices
Platform
Alibaba Cloud Model Studio (ex-DashScope): 145+ hosted model IDs across chat, coding, vision, reasoning, embeddings, speech, image and video; OpenAI-compatible and Anthropic-protocol endpoints; international serving from Singapore, cheaper Beijing region; batch mode, context caching, Coding Plan subscriptions
Key differentiator
Breadth nobody matches: Apache 2.0 weights across every modality (Qwen3.5-397B-A17B top-3 open generalist; Qwen3-Coder-Next the standout local coding model) alongside a closed 1T-parameter MoE flagship (Qwen3.7-Max, ~24B active) scoring 80.4 SWE-bench Verified with 1M context
Pricing
Qwen3.7-Max $1.25/$3.75 per 1M tokens (promo; list $2.50/$7.50); Qwen-Plus $0.40/$1.20; Qwen-Flash from $0.05/$0.40; open-weight models from $0.15 hosted or free self-hosted. One-time ~70M-token trial (90 days, Singapore). Coding Plan ~$10–50/month. Free OAuth API tier ended 15 April 2026
Vendor
Alibaba Cloud (Hangzhou) — the e-commerce giant’s cloud arm, which turned Qwen (通义千问) into the world’s most-downloaded open model family and the backbone of thousands of third-party fine-tunes
Platform notes (2026): four things to price in before production. The free API tier is gone: the developer OAuth free tier (originally 1,000 requests/day) and the free Qwen Code CLI allowance were both discontinued on 15 April 2026 — older tutorials promising free Qwen API access are out of date; what remains is a one-time onboarding trial (~1M tokens per model, roughly 70M total, valid 90 days on the Singapore endpoint) and the entirely free consumer chat app. Billing has three structural traps: several models use tiered pricing where the per-token rate rises with the length of a single request (all tokens bill at the tier the request lands in); the same model costs different amounts by region (Beijing is 60–70% cheaper than Singapore but stores data in mainland China); and thinking mode bills its hidden reasoning tokens at 3–10x the standard output rate. The flagship price is a promotion: Qwen3.7-Max’s $1.25/$3.75 is a 50% launch discount against a $2.50/$7.50 list price — budget months ahead at list, not promo. Verbosity is the silent surcharge: Qwen models generate notably more output tokens per task than Western rivals, which offsets part of the headline price advantage — measure real per-task cost, not per-token rates.

What Is Alibaba Qwen?

Qwen (Tongyi Qianwen, 通义千问 — “truth from a thousand questions”) is Alibaba Cloud’s family of AI models and the developer platform that serves them — and by 2026 it has grown into something no single Western provider offers: a full-spectrum catalogue where nearly every layer is available both as a hosted API and as downloadable Apache 2.0 weights. The story arc matters. Launched in 2023 as a credible-but-domestic Chinese model line, Qwen broke out internationally through open weights: the Qwen2.5 generation became the default base model for the world’s fine-tuning community (more derivative models on Hugging Face than any other family, including third-party builds like Rio de Janeiro’s city-government fine-tune), and the Qwen3 generation (2025–26) converted that distribution into capability leadership — Qwen3.5-397B-A17B ranks among the top three open generalists anywhere, and Qwen3-Coder-Next achieved something genuinely new: with just 3 billion active parameters it competes with Claude Sonnet-class models on coding, making it arguably the best local coding model in existence and a laptop-class one at that. Above the open family sits the closed tier — the part of the strategy that mirrors the West. Qwen3.7-Max (20 May 2026) is a roughly 1-trillion-parameter Mixture-of-Experts flagship with ~24B active parameters per forward pass, a 1M-token context window, and results that ended any lingering “fast follower” framing: 80.4 on SWE-bench Verified, 60.6 on the harder SWE-Bench Pro, 56.6 on the Artificial Analysis Intelligence Index — global top ten at launch — at $1.25/$3.75 per million tokens on promotion. Everything is served through Alibaba Cloud Model Studio (formerly DashScope), an OpenAI-compatible platform with an Anthropic-protocol option, batch invocation at 50% off, context caching, and a catalogue exceeding 145 model IDs that stretches well past text: vision-language models with image and video input (a Qwen3.7-Plus sibling takes multimodal input at $0.40/$1.60), embeddings from $0.01 per million, rerankers, speech, and image/video generation. The pricing philosophy is aggressive at every tier — the international rate card runs from $0.05-input Qwen-Flash to the flagship, with the mid-tier roughly 9x cheaper on input than Claude Opus-class rates (one benchmark house ran its full evaluation suite for $483 on a Qwen Plus-tier model versus $4,970 on Opus 4.6) — and paid API traffic is not used for training by default. Within our Model Providers & AI Infrastructure category, Qwen is the breadth champion and DeepSeek’s more corporate sibling: where DeepSeek is two text models and a price floor, Qwen is an entire aisle — and the only provider whose open-weights strategy covers every modality at genuinely competitive quality.

Core Features

The open-weights family: Apache 2.0 across everything

Qwen’s foundation — and the reason it belongs on any open-model shortlist ahead of almost everyone — is the licence and the coverage. The weights ship under Apache 2.0: a real open licence with commercial use, modification, fine-tuning and redistribution, no MAU thresholds (Llama’s catch), no acceptable-use novelties — matched at the frontier-adjacent tier only by DeepSeek’s MIT terms, and unlike DeepSeek, applied across an entire product line rather than two text models. Walk the shelf. Generalists: Qwen3.5-397B-A17B is a top-three open model globally — 397B total parameters, 17B active, runnable in 4-bit quantisation on a 256GB Mac Studio — while the mid-size Qwen3.6-35B-A3B gives multimodal (image-input) capability that runs at 20–25 tokens/second on a single RTX 4090 and costs just $0.15 per million input hosted (cache reads $0.05). Coding: Qwen3-Coder-Next is the family’s headline act — 3B active parameters yet competitive with Claude Sonnet-class models on coding benchmarks, beating the 397B flagship generalist on code while being ~15x smaller in active parameters; it has become the default answer to “what’s the best coding model I can run locally?” Then vision-language (the Qwen-VL line that long led open multimodal leaderboards), math specialists, embeddings, and audio. The strategic consequences compound. Self-hosting is free and unencumbered — from an ollama pull on a gaming PC to production vLLM clusters — which makes Qwen the escape hatch from every API-pricing and data-residency concern in this review: the weights don’t care where Alibaba’s servers are. The third-party market is the deepest in open AI: DeepInfra serves Qwen at 50–80% below official rates, Groq hosts Qwen3-32B on a genuinely free tier (no card required), OpenRouter routes across a dozen hosts — so “using Qwen” rarely means “paying Alibaba.” And the fine-tune ecosystem means the family improves without Alibaba lifting a finger: thousands of derivatives, from government deployments to vertical specialists, all feed reputation and tooling back into the base line. The honest boundary: the very best of Qwen is not open. Qwen3.7-Max — the model with the top-ten benchmark placement — is closed weights, full stop, and Alibaba has held every Max-tier flagship closed while open-sourcing the tiers below. If your requirement is “frontier capability AND weights I control,” Qwen’s open shelf gets closer than anyone’s, but the last 10% stays behind the API — a philosophical inversion of DeepSeek, which opens its best and hosts it cheaply, and the single most important nuance in choosing between the two.

Qwen3.7-Max and the hosted platform

The hosted side of Qwen is anchored by a flagship that ended the fast-follower era and a platform whose breadth is its own argument. Qwen3.7-Max first: launched 20 May 2026 as a ~1T-parameter MoE with roughly 24B active parameters, it carries a 1M-token context window and posted numbers that placed it in the global top ten at launch — 80.4 on SWE-bench Verified (level with DeepSeek’s V4 Pro line and within range of Western flagships), 60.6 on SWE-Bench Pro’s harder autonomous-engineering tasks, 56.6 on the Artificial Analysis Intelligence Index — at $1.25/$3.75 per million tokens on promotion, roughly one-sixth the per-token cost of comparable closed alternatives (one caveat: an earlier Qwen3-Max variant capped context at 262K, so verify the window on the exact model ID you deploy). Below it, the ladder is deep and cheap: Qwen3.5-Plus at $0.40/$2.40 is the value-tier workhorse (about 9x cheaper on input than Opus-class rates — the $483-versus-$4,970 full-benchmark-run comparison was made on this tier’s successor), Qwen3.6 Flash at $0.19/$1.13 and Qwen-Flash from $0.05/$0.40 handle volume work, the multimodal Qwen3.7-Plus sibling takes image and video input at $0.40/$1.60, and hosted open-weight models start at $0.15. The platform machinery is genuinely capable: OpenAI-compatible endpoints (plus an Anthropic-protocol option, so both major SDK ecosystems migrate with a base-URL change), batch invocation at 50% off input and output on supported models, context caching that discounts repeated prefixes (the two discounts don’t stack), function calling, JSON mode, and optional thinking mode (enable_thinking: true) for chain-of-thought reasoning on supported models. For flat-rate buyers, the Coding Plan (~$10/month Lite, ~$50/month Pro with up to 90K requests) converts token anxiety into a subscription for editor-based coding workloads. Now the friction, because Qwen’s platform demands more spreadsheet literacy than any rival’s. Pricing is three-dimensional: rates vary by model, by region (the Beijing endpoint is 60–70% cheaper than Singapore but processes data in mainland China with no free quota), and — on several models including the coding line — by request length, with tiered brackets where a long prompt re-prices every token in the request at the higher tier. Thinking mode bills hidden reasoning tokens at 3–10x standard output rates. The flagship’s headline price is a time-limited 50% promotion against a $2.50/$7.50 list. And Qwen’s well-documented verbosity — more output tokens per task than Western peers — quietly narrows the real-world gap between rate card and invoice. None of this is disqualifying; all of it is homework the OpenAI and Anthropic platforms don’t assign.

Trust, residency and the post-free-tier reality

Qwen’s enterprise story is stronger than DeepSeek’s on paper and still carries the same fundamental asterisk — plus a fresh scar from April 2026 that buyers should read as a character reference. Residency first: the international API serves from Singapore (Alibaba Cloud’s ap-southeast-1), not mainland China, and Model Studio’s terms state that paid API traffic is not used for model training by default — two genuine differentiators over DeepSeek’s first-party endpoint, and enough for many international businesses. Some Qwen models are additionally available through AWS Bedrock, Azure and Google Vertex AI in other regions, and enterprise contracts add SLAs, explicit no-training clauses and data-residency commitments. But Singapore is not the West, Alibaba is a Chinese company subject to Chinese jurisdiction, and for the same regulated-industry and government-adjacent buyers who ban DeepSeek’s first-party API, Qwen’s hosted platform typically falls under the same policy — the difference is that Qwen’s mitigation path is wider: Apache 2.0 weights self-hosted on your infrastructure, or served by Western hosts (DeepInfra with private-deployment options, Groq, Together and the OpenRouter constellation), with only the closed Max flagship forcing you back to Alibaba’s endpoints. The Beijing-region discount deserves its own flag: 60–70% cheaper is tempting, and it moves your data into mainland China — that trade should be a compliance decision, never a cost optimisation that happens by default. Then there’s April 15, 2026 — the day Alibaba killed the free developer tier. The OAuth API allowance (originally 1,000 requests/day, already cut to 100) and the Qwen Code CLI’s free 2,000-requests/day coding tier both ended simultaneously, converting overnight a large community of free-tier developers into either paying customers or — as the forums showed vividly (“ngl i just subscribed to Claude”) — ex-users. The episode cuts both ways. Charitably: it’s the normal maturation of a loss-leader, the consumer chat app remains genuinely free, and the replacement trial (~70M tokens across models, 90 days) is a real onramp. Less charitably: it’s the second data point (after promotional flagship pricing) that Qwen’s generosity is tactical and revocable, and teams building on the hosted platform should architect with the same assume-churn posture this category keeps teaching — pin model IDs, watch deprecation notices, and keep the open weights as your leverage. That leverage is the real trust story: with Qwen, unlike any closed Western provider, the worst case is never lock-in — it’s migration to your own GPUs.

Scored Categories

Open-weights breadth (Apache 2.0 across text, code, vision, reasoning)

9.5

Price-performance (flagship $1.25/$3.75; mid-tier ~9x under Opus-class)

9.1

Model catalogue breadth (145+ IDs; VL, embeddings, speech, image, video)

9.0

Flagship capability (SWE-V 80.4; AA Index top-10; 1M context)

8.7

Ecosystem & third-party hosting (most fine-tuned family; Groq free tier)

8.5

Developer experience (OpenAI + Anthropic compatible; Model Studio complexity)

8.0

Billing predictability (length tiers, region pricing, thinking multipliers, promos)

6.9

Enterprise trust & stability (Singapore residency helps; free-tier churn; CN parent)

6.7

Pricing

Model / item Price (per 1M tokens) Notes
Qwen3.7-Max (flagship, closed) $1.25 input / $3.75 output ⚠ 50% launch promotion — list price $2.50/$7.50. ~1T MoE (~24B active), 1M context, SWE-bench Verified 80.4. Budget at list price
Qwen3.5-Plus $0.40 / $2.40 Value-tier workhorse — roughly 9x cheaper on input than Claude Opus-class rates; Qwen-Plus line from $0.40/$1.20
Qwen3.7-Plus multimodal sibling $0.40 / $1.60 Image and video input (June 2026)
Qwen3.6 Flash / Qwen-Flash $0.19/$1.13 · from $0.05/$0.40 Volume tier for extraction, classification, routing. Qwen-Turbo deprecated in favour of Flash
Hosted open-weight models From $0.15 input (Qwen3.6-35B-A3B) Same Apache 2.0 models you can self-host free; cache reads $0.05. Third parties (DeepInfra) run 50–80% cheaper; Groq hosts Qwen3-32B free
Embeddings From $0.01 combined Part of a 145+ model catalogue including rerankers, speech, image and video generation
Batch & caching Batch = 50% off; caching discounts repeat prefixes The two discounts cannot be combined on the same request
Coding Plan ~$10/mo Lite · ~$50/mo Pro Flat-rate coding subscription (Pro up to ~90K requests/month) — the answer to April’s free-tier removal
Free access One-time trial ~70M tokens ~1M tokens per model, 90 days, Singapore endpoint only. OAuth free tier and free Qwen Code CLI ended 15 April 2026. Consumer chat app remains free
Qwen’s rate card is the cheapest broad catalogue in the market and the easiest to misread — four rules keep the invoice honest. One: price by region deliberately. The published international rates are Singapore; the Beijing endpoint is 60–70% cheaper but processes data in mainland China with no free quota — treat that as a compliance decision made by counsel, not a toggle flipped by an engineer chasing savings. Two: respect the length tiers. On tiered models (including the coding line), one long request re-prices every token in it at the higher bracket — keep requests under tier thresholds where you can, and split monster contexts. Three: treat thinking mode like a premium SKU. enable_thinking: true bills hidden reasoning tokens at 3–10x standard output — enable it per-task, never globally, and cap outputs; Qwen’s natural verbosity already pads output bills. Four: budget at list, not promo. The flagship’s $1.25/$3.75 is a time-limited 50% discount, and April 2026 proved free things here can end abruptly — plan at $2.50/$7.50 and let the promotion be upside. Route most traffic to Flash/Plus tiers, reserve Max for measured failures, use batch mode for anything overnight-able, and remember the ultimate discount: the Apache 2.0 weights, self-hosted, cost per token exactly nothing. Verify current rates in the Model Studio console — Alibaba’s prices move often, and usually down.

Strengths

  • The most complete open-weights family in AI — Apache 2.0 across text, coding, vision and reasoning
  • Qwen3-Coder-Next: Sonnet-class coding at 3B active parameters — the best local coding model of 2026
  • Genuine frontier flagship — Qwen3.7-Max top-10 globally at launch, 80.4 SWE-bench Verified, 1M context
  • Aggressive pricing at every tier — mid-tier ~9x under Opus-class input rates; flagship at ~1/6th comparable closed cost
  • 145+ hosted model IDs — the widest catalogue of any provider (VL, embeddings, speech, image, video)
  • OpenAI-compatible and Anthropic-protocol endpoints — trivial migration from either ecosystem
  • International serving from Singapore; paid API not used for training by default
  • Deepest third-party ecosystem in open AI — DeepInfra 50–80% cheaper, Groq free tier, most-fine-tuned family on Hugging Face
  • Batch at 50% off, context caching, flat-rate Coding Plan option

Weaknesses

  • Most complicated billing in the category — length tiers, region pricing, thinking multipliers, promo rates
  • Free developer API tier and free Qwen Code CLI killed abruptly on 15 April 2026
  • The flagship stays closed — the best Qwen is never the open Qwen
  • Famous verbosity inflates output bills, narrowing the headline price advantage
  • Chinese parent company — hosted platform fails the same enterprise policies that bar DeepSeek, Singapore endpoint notwithstanding
  • Flagship price is a time-limited promotion (list $2.50/$7.50)
  • Model Studio’s console and 145-ID catalogue are harder to navigate than Western platforms
  • Earlier Max variant capped context at 262K — verify the window on your exact model ID

Verdict: 8.3 / 10 — The Everything Provider

Alibaba Qwen earns a strong 8.3 as the provider whose answer to nearly every question is “yes, and there’s an open-weights version.” No rival matches the shape of the offer: a genuine frontier flagship (Qwen3.7-Max, top-ten globally at launch, 80.4 SWE-bench Verified, 1M context) at a sixth of comparable closed pricing; an Apache 2.0 family beneath it that owns the open-model conversation — the top-three-anywhere 397B generalist, the astonishing Qwen3-Coder-Next that puts Sonnet-class coding on a laptop, vision-language models, the lot; and a 145-model hosted catalogue with both OpenAI and Anthropic compatibility, all priced with the aggression that has made Chinese providers the market’s disciplinarians. For open-weights buyers specifically, Qwen is now the default recommendation of this entire site: broader than Llama, more permissively licensed, and better at code. What holds it at 8.3 — level with DeepSeek and Meta Llama, a notch under Mistral — is the tax the breadth collects: billing so multidimensional (length tiers, region arbitrage, thinking-mode multipliers, promotional list prices) that it needs a spreadsheet before a proof of concept; the April 2026 free-tier shutdown, which converted goodwill into churn overnight and stands as a warning about how quickly generosity here can be repriced; the verbosity that quietly pads invoices; and the strategic asymmetry that the family’s very best model is the one you can’t download — the exact inverse of DeepSeek’s open-flagship posture, and the deciding line between them for many teams. The buying logic: if you want the best open models to run yourself, Qwen is the answer, full stop, and no residency concern survives self-hosting. If you want the cheapest hosted frontier-adjacent tokens and your data can route through Singapore (or a Western host of the open weights), Qwen’s ladder from $0.05 Flash to the Max flagship is as good a value stack as exists. If your compliance posture bars Chinese-parent platforms entirely, take the weights to Western infrastructure and leave the hosted flagship to others. Either way, learn the rate card before you scale — Qwen rewards diligence with the best breadth-per-dollar in AI, and punishes assumptions with the category’s most surprising invoices.

Frequently Asked Questions

Qwen or DeepSeek — which Chinese model family should I build on?

They’re the two poles of the same price revolution, and the decision usually resolves on three axes: what’s open, what modalities you need, and which billing model you can live with. Openness first, because the philosophies invert. DeepSeek open-sources its very best — the V4 flagship weights are MIT-licensed, so the top of its range is fully yours to self-host or buy from Western providers. Qwen open-sources everything except its best — the Apache 2.0 shelf is vastly broader (text, coding, vision, reasoning, at every size from laptop to data-centre), but the Qwen3.7-Max flagship stays closed behind Alibaba’s API. So: if your requirement is “the strongest possible model under my own control,” DeepSeek wins — its open flagship outguns Qwen’s open ceiling. If your requirement is “an open model for each job” — a local coding model (Qwen3-Coder-Next is the category’s best), a vision-language model, a 4090-friendly multimodal mid-size, embeddings — Qwen’s shelf has no rival, DeepSeek included, because DeepSeek ships two text models and nothing else. Modality is the second axis and it’s one-sided: DeepSeek is text-only with no hosted tools; Qwen’s hosted catalogue spans vision input, image and video generation, speech and embeddings across 145+ model IDs. Any multimodal requirement decides for Qwen immediately. Billing is the third axis and it favours DeepSeek: its pricing is brutally simple (two models, flat rates, automatic caching, no tiers, no regions) and its floor is lower — V4 Flash at $0.14/$0.28 undercuts everything in Qwen’s ladder at comparable capability, and its “discounts” become permanent list prices rather than expiring promotions. Qwen’s rate card is richer and trickier: length-tiered rates, region-dependent pricing, thinking-mode multipliers, and a flagship promo that will end. Residency is roughly a wash with an edge to Qwen — its international endpoint serves from Singapore with a no-training default on paid traffic, versus DeepSeek’s first-party processing in mainland China — though enterprises that ban Chinese-parent platforms typically ban both hosted services and route to the open weights on Western hosts either way, which both families support well (Qwen’s third-party market is deeper; DeepSeek’s open flagship is stronger). Benchmarks at the top are effectively tied — both flagships post ~80 on SWE-bench Verified. The pragmatic pattern we see most in 2026: Qwen weights for local and self-hosted work (especially coding), DeepSeek’s API for cheap hosted text volume, and a Western flagship for the hardest 1% — all three behind an OpenAI-compatible router, which every one of these platforms supports precisely so you never have to marry any of them.

Is Qwen really free to use?

Parts of it are genuinely, permanently free; the part everyone used to mean by “free Qwen API” is dead — and knowing which is which saves you from following a 2025 tutorial off a cliff. What is truly free, with no expiry: the open weights. Most of the Qwen family ships under Apache 2.0, so downloading Qwen3-Coder-Next or a 35B multimodal model and running it locally — ollama pull on a gaming PC, vLLM in production — costs nothing, forever, with full commercial rights, no API key and no rate limits beyond your own hardware (a 4090 runs the 35B-A3B at 20–25 tokens/second; a 256GB Mac Studio fits the 397B flagship-class generalist in 4-bit). Also genuinely free: the consumer Qwen Chat app (web, iOS, Android, macOS) — Alibaba has every commercial incentive to keep its consumer front door open, and it’s the right tool for casual evaluation; and Groq’s hosted free tier for Qwen3-32B, which requires no credit card and is the fastest zero-cost way to hit a Qwen model over an API (rate limits apply). What is no longer free: the developer API. On 15 April 2026 Alibaba discontinued both the OAuth free tier (originally 1,000 requests/day, already trimmed to 100 before the end) and the Qwen Code CLI’s free 2,000-requests/day coding allowance — simultaneously, with real community fallout. The replacement is a one-time onboarding trial: new Alibaba Cloud accounts on the international (Singapore) endpoint receive roughly 1 million tokens per model — about 70 million tokens total across the catalogue — valid for 90 days, which is a substantial prototyping budget but categorically not a standing free tier. After the trial, it’s pay-as-you-go from $0.05 per million input on Qwen-Flash, or the flat-rate Coding Plan (~$10/month Lite, ~$50/month Pro) that Alibaba positioned as the free CLI tier’s successor for editor-based coding. The practical decision tree: evaluating casually → free chat app; want free API calls today → Groq’s Qwen3-32B tier; building anything local or privacy-sensitive → the Apache 2.0 weights, which are the best free tier in this review because they never expire; building on the hosted platform → burn the 90-day trial deliberately on your real workload, measure true per-task cost (Qwen’s verbosity means per-token rates understate it), and budget at list prices — April 2026 is your evidence that free things here end without much ceremony.

Can enterprises trust Alibaba’s hosted Qwen API with their data?

Qwen’s hosted trust story is materially better than DeepSeek’s first-party one and still ends at the same wall for the strictest buyers — so the honest answer is “many can, some structurally can’t, and the open weights mean nobody has to.” The genuine positives, stated fairly: the international API serves from Alibaba Cloud’s Singapore region (ap-southeast-1), not mainland China; Model Studio’s terms state paid API requests are not used for model training by default; enterprise contracts add explicit no-training clauses, SLAs and data-residency commitments; selected Qwen models are also available through AWS Bedrock, Azure and Google Vertex AI, which lets you consume Qwen under a Western hyperscaler’s data-processing agreement entirely; and the third-party market (DeepInfra with private deployments, Together, and Groq) offers the open-weight tiers on US/EU infrastructure with their own compliance artefacts. For international businesses without jurisdiction-specific restrictions, that package — Singapore residency, no-training default, hyperscaler distribution — clears procurement comfortably and does so more easily than DeepSeek’s China-based first-party endpoint. The structural limits, stated equally fairly: Alibaba is a Chinese company subject to Chinese jurisdiction regardless of where a given server rack sits, and organisations whose policies restrict Chinese-parent platforms — defence-adjacent, government, parts of regulated finance and healthcare, or anyone whose counsel has drawn that line — will typically bar the hosted platform on that basis alone, Singapore notwithstanding. Two operational cautions belong in every assessment. First, the Beijing-region endpoint: it’s 60–70% cheaper and it processes data in mainland China — make region selection an explicit, documented compliance decision, because the cost delta creates quiet pressure in exactly the wrong direction. Second, platform stability as a trust signal: the abrupt 15 April 2026 termination of the free developer tier and free coding CLI wasn’t a data issue, but it demonstrated how quickly terms here can change — pin model IDs, contract for what matters, and architect for churn. The mitigation that resolves everything else is the one unique to open families: the Apache 2.0 weights. Self-hosted Qwen on your own infrastructure involves Alibaba in precisely nothing — no endpoint, no jurisdiction, no terms — and covers every tier except the closed Max flagship. So the enterprise decision procedure: if your policy permits Singapore-resident Chinese-parent platforms, the hosted API is a legitimate, well-papered choice — take the enterprise contract, stay out of the Beijing region, and never send secrets or regulated identifiers you wouldn’t send any third-party LLM. If it doesn’t, consume Qwen through Bedrock/Azure/Vertex or self-host the open weights — and accept that the closed flagship simply isn’t for you, which is the one trade-off no amount of paperwork removes.