Qwen3.8-Max vs. DeepSeek V4-Flash: Two Rival Strategies for AI Supremacy.
Alibaba and DeepSeek Are Racing China's AI Models Toward Zero Cost:
Qwen3.8-Max claims the size crown while DeepSeek's V4-Flash undercuts nearly every rival on price — and the real lesson is that sticker price isn't the same as task cost.
2.4T: Total parameters in Alibaba's Qwen3.8-Max
$0.14: V4-Flash's price per million input tokens
3¢: DeepSeek V4-Flash's average cost per benchmark test
1: Alibaba Goes Bigger With Qwen3.8-Max:
China's AI labs are no longer just chasing capability — they're racing each other on cost, and the latest releases show two very different ways to win that race.
Alibaba has launched Qwen3.8-Max, its largest AI model to date, at the same time DeepSeek's new V4-Flash model is drawing attention for inference pricing that undercuts several major competitors.
Qwen3.8-Max carries 2.4 trillion total parameters and runs on a mixture-of-experts architecture, activating only part of the model — around 95 billion parameters — for any given request, which Alibaba says cuts both cost and response delay compared with running the full model on every query.
The model can process text, images, and video, supports up to one million tokens of context, and Alibaba says it completed a software engineering project spanning 16 days. Its scale puts it close to Moonshot AI's Kimi K3, released in July at 2.8 trillion total parameters with roughly 104 billion active during inference.
Since its release, Qwen3.8-Max has moved to the top spot among Chinese text models on the crowdsourced Arena.AI leaderboard, though it still trails several Anthropic models overall, and ranks second — behind an Anthropic Claude Fable 5 variant — on Arena.AI's leaderboard for models that analyze images and other visual material.
2: DeepSeek Undercuts the Field on Price:
Rather than compete on raw scale, DeepSeek is competing on what it costs to actually use the model.*
mixture-of-experts approach to Qwen3.8-Max, but at a smaller scale — Artificial Analysis lists it at 284 billion total parameters, with only 13 billion active during inference. What sets it apart is pricing: $0.14 per million input tokens and $0.28 per million output tokens, a fraction of what most widely used systems charge. Artificial Analysis also lists cache-hit pricing for the model's Max Effort version at just $0.003 per million tokens — 98% below its standard input rate — for context that's been processed before and can be reused across requests.
That pricing carried through to real benchmark testing. Artificial Analysis estimated V4-Flash's average cost at three cents per test, compared with 86 cents for Kimi K3, $1.86 for OpenAI's GPT-5.6 Sol, and $3.15 for Anthropic's Claude Fable 5. On Artificial Analysis' Intelligence Index, the Max Effort reasoning version of V4-Flash scored 40, with an output rate of about 118 tokens per second during testing.
3: Why Token Price Isn't the Whole Story:
A model's advertised API rate and what it actually costs to finish a task can be two very different numbers.

The Hidden AI War
Nobody Is Telling You About
Our latest documentary deep-dive into the geopolitical struggle for machine intelligence dominance. Explore the two paths of AI development: open source vs. closed architecture.
Kimi K3 makes the point clearly. Moonshot AI prices it at $3 per million input tokens and $15 per million output tokens, with cached input at $0.30 per million tokens — but on Artificial Analysis' AA-Briefcase benchmark for agentic knowledge work, Kimi K3 averaged $10.57 per task, generating around 120,000 output tokens across an average of 83 turns.
Artificial Analysis attributed that cost to the combination of token pricing, output volume, and the sheer number of model interactions required — repeated calls and longer outputs can drive total cost well past what the headline rate suggests.
● Kimi K3 recorded the second-highest overall score on AA-Briefcase, behind Claude Fable 5, and scored 57 on the broader Intelligence Index.

Meta's Next Big Bet: This New App Lets You Build Games Simply by Typing a Prompt
● Architecture, active parameter count, and token consumption all shape real-world cost — not model size alone.
Support our research
Independent analysis fueled by you.
● The number of calls required to complete a task can matter as much as the per-token price itself.
“Many business workflows do not need the industry's very best model. They need models that are good enough, affordable, transparent and accessible, and open-weight models help meet that demand.”
— Lian Jye Su, Chief Analyst, Omdia
4: Open Weights Widen the Deployment Options:
Cost isn't only about API pricing — it's also about how much control developers have over where and how a model runs.
Alibaba, DeepSeek, and Moonshot AI have all continued supporting open-weight releases alongside their hosted APIs, giving developers more flexibility in deployment. Artificial Analysis lists DeepSeek V4-Flash as an open-weight model under an MIT license, with weights available on Hugging Face; Kimi K3 is open-weight as well, under Moonshot AI's own license.
That openness lets developers run the models on their own infrastructure or through third-party providers rather than depending solely on a single hosted API — though deployment costs still hinge on the hardware and infrastructure a business chooses to run it on. It's a notably different approach from OpenAI, Anthropic, and Google, which generally keep their flagship model weights closed.
Taken together, the moves from Alibaba and DeepSeek point to a market that's maturing past the pure capability race. Bigger and cheaper are no longer opposing strategies — they're both viable paths to the same goal: making powerful AI affordable enough for everyday business workloads, not just headline benchmarks.
You Don't Need the Industry's Priciest Model — You Need the Right One.
The lesson from Alibaba and DeepSeek's price war isn't just about model size or token rates — it's that the cheapest sticker price and the lowest real-world cost aren't the same thing, and most businesses don't need frontier-scale compute to get frontier-level results.
Otherworlds AI's Agent+ Business AI Platform is built the same way: matching the right model and the right workflow to your actual task, so you're not paying for capability you'll never use. That's how automation stays affordable as you scale, not just impressive in a demo.
Find the right-sized AI solution for your business at otherworldsai.com








