Qwen 3.8 Max: The Open-Weight 2.4T Flagship That Resets Qwen API Pricing
Qwen 3.8 Max launches at $2/$6 per million tokens with open weights next week. Here is how the new Qwen API pricing changes your agent budget.
On August 3, 2026, Alibaba released Qwen 3.8-Max, the most capable model in the Qwen family to date. The 2.4-trillion-parameter mixture-of-experts flagship lands at $2 per million input tokens and $6 per million output tokens. For heavy AI users, the headline is less about the benchmark scores and more about what this does to Qwen API pricing across the board, especially the promise of open weights arriving next week.
This is the first time Alibaba has open-sourced a Qwen-Max-class model. The open weights, alongside a much smaller Qwen3.8-27B, are scheduled for release next week. That combination, a frontier-grade hosted API at a budget price plus a downloadable version of the same weights, is the strongest signal yet that the AI price war is now a race to zero on open-weight flagship performance.
What Qwen 3.8 Max costs and what you get
Qwen 3.8 Max is priced to compete head-on with Western flagships, not to undercut them into the ground. The exact per-token rates on Qwen Cloud are $2 per million input tokens and $6 per million output tokens. That positions it well below Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30, and roughly level with the mid-tier GPT-5.6 Terra at $2.50/$15 on input while beating it clearly on output.
The caching discounts are where heavy users save real money. Implicit cache reads cost $0.25 per million tokens and explicit cache reads just $0.17 per million. For agentic workloads that repeatedly send the same system prompt, tool definitions, and repository context, those rates turn long sessions into a fraction of their nominal cost. Context tops out at 1 million tokens, with a 262K reasoning budget, so the model can hold an entire large codebase in a single conversation.
It ships with a 2 million tokens-per-minute limit and 15,000 requests per minute, figures that comfortably support dozens of concurrent agents before you hit a wall. For a heavy user who has been throttled by tighter rate limits elsewhere, that headroom matters as much as the sticker price.

The open-weight release changes the math
The biggest news for cost-sensitive users is not the API price. It is that Qwen 3.8-Max and Qwen3.8-27B are going open weight next week. Qwen3.6-27B was already one of the most popular local models on the market, widely praised as the best thing you could run on consumer hardware before the models got too large. Its successor should inherit that role.
When a frontier-class model becomes downloadable, heavy users gain an option that removes the API meter entirely for the workloads they can fit locally. The trade-offs are the usual ones: you need the hardware, you own the operational overhead, and you lose the provider’s uptime guarantees. But for steady, high-volume workloads, self-hosting an open-weight flagship at a one-time infrastructure cost changes the pay-per-token economics completely. The break-even analysis that applied to earlier open-weight models now applies to a model close to the top of the frontier.

Controls that actually lower your bill
Alibaba has also shipped something heavy users have been asking for across every provider: a direct dial on reasoning effort. Qwen 3.8 Max supports a reasoning_effort parameter with three settings: xhigh for complex tasks, medium for balanced work, and low for cheap, fast reasoning. Because reasoning tokens are billed as output, cutting from xhigh to medium or low can reduce the expensive half of your bill without necessarily harming the tasks where you do not need the full reasoning depth.
That is more honest than the recent trend of burying smaller or weaker models behind a paywall. Here the provider gives you the same model and a real way to control how much it thinks, letting you match spend to task difficulty instead of forcing a model downgrade.
For a heavy AI user running agent pipelines, this pairs naturally with the long-horizon capabilities. Qwen is marketing the model as able to autonomously code and deliver complete projects spanning ten or more days, handling specialized work across legal, financial, and design domains with production-grade output. Whether or not the marketing fully holds up, the combination of a reasoning dial, deep context, and a cheap cache tier gives you the levers to keep an extended agent run inside budget.

A new Qwen coding plan and what to watch
Alongside the model, Alibaba is pushing a subscription model via Qwen Cloud that bundles multiple models, including Qwen 3.8 Max, access-harness tools, and concurrency tiers starting at Lite, Standard, and Pro plans that let you run from two up to eight agents at once. If you are already paying for Cursor, Claude Code, or Codex, this is a direct competitor on the monthly-plan front. The individual and team tiers roughly mirror what the Western coding tools charge, so the value proposition rests on whether Qwen’s models deliver in your actual task mix.
The open-weight release next week is the date to watch. If Qwen 3.8 Max’s downloadable version matches the hosted API on real workloads, the argument for any heavy user staying locked into a $200-per-month Western subscription weakens considerably. Budget the hosted API now to validate the model, then decide after the weights drop whether to move the high-volume jobs local.
For now, the practical guidance is simple. Sample Qwen 3.8 Max on your own longest-running tasks. Turn down the reasoning effort wherever you can. Lean on the $0.17 cache reads for anything with a stable preamble. And mark next week on the calendar for the open weights, because that is where the real cost relief for heavy AI users likely arrives.
Now available
Stop guessing your AI limits
The Mac app and web dashboard watch your Claude, ChatGPT, Gemini and more, and warn you before quotas hit.