Qwen 3.8-Max launches with 2.4 trillion parameters. Don't trust the benchmarks just yet

Alibaba launched its large AI model Qwen3.8-Max with many parameters. This model supports extensive context windows and multimodal inputs for complex tasks. OpenAI also hinted at its next model, Astra, for advanced problem-solving.

Alibaba says the model supports a one million token context window
Alibaba just launched a 2.4 trillion-parameter AI model. But that's not the biggest story. Alibaba has unveiled Qwen3.8-Max, its largest AI model yet with 2.4 trillion parameters, making it the second-largest publicly announced model globally after Moonshot AI's Kimi K3. The headline number grabbed attention and even pushed Alibaba's stock higher. But the parameter count isn't what matters most.

The bigger questions are simple: How good is the model really? And should anyone trust the benchmarks yet?

A giant model that doesn't use all of itself

Despite having 2.4 trillion parameters, Qwen3.8-Max isn't activating the entire network every time it generates text. It's a Mixture-of-Experts (MoE) model, meaning only about 95 billion parameters are active per token. That keeps inference costs significantly lower than running a traditional dense model of comparable size while still delivering frontier-level capabilities.


Alibaba says the model supports a one million token context window (roughly 750,000 words, allowing it to process entire books or large codebases in a single prompt), multimodal inputs (meaning it can understand text, images, documents and videos together), and can generate outputs of up to 131,072 tokens (around 100,000 words) in one response. The company has also promised to release open weights next week, allowing developers to download and run the model themselves instead of accessing it only through Alibaba's cloud, alongside a much smaller 27-billion-parameter (27B) version that is likely to be far more practical for enterprises to deploy on their own infrastructure.

The benchmarks come with a big disclaimer


Alibaba published an extensive benchmark comparison against models from OpenAI and Anthropic. The problem? Alibaba ran every single benchmark itself. That means it chose the prompts, sampling methods, retry policies and evaluation settings. At launch, there were no independent scores from platforms such as Artificial Analysis or community leaderboards.

ADVERTISEMENT
So every benchmark should be treated as a vendor claim rather than an independently verified result.
Interestingly, Alibaba didn't try to paint a flawless picture.

The company shows Qwen leading on benchmarks such as PaperBench and IFBench, suggesting improvements in long-horizon reasoning and instruction following. But it also openly shows the model trailing Anthropic's Claude on SWE-bench Pro, one of the industry's most important software engineering evaluations, as well as Humanity's Last Exam, a benchmark designed to measure advanced reasoning.

The real pitch isn't benchmarks

Alibaba appears to be positioning Qwen3.8-Max less as a chatbot and more as an autonomous AI worker.
The company demonstrated the model spending 16 days building a command-line project on its own, producing hundreds of commits, pull requests and issues. In another demonstration, it reproduced an academic machine learning paper over several days, while a third experiment placed the model in a 24-hour data science competition where it reportedly outperformed most participating human teams.

ADVERTISEMENT
Unlike many AI demos, Alibaba has published the repositories, allowing developers to inspect what the model actually did rather than simply watching a polished video.

Why Indian developers should care

For engineering teams, GCCs and AI startups in India, pricing could matter far more than benchmark rankings.
Alibaba has priced Qwen significantly below premium Western models, with additional discounts for cached inputs. That's particularly relevant for agentic workloads, where long-running AI systems repeatedly reuse the same context, making cache costs a major part of overall inference expenses.
ADVERTISEMENT

The caveat is reasoning. The model supports an enormous reasoning budget, but those reasoning tokens are billed as output tokens. Teams that leave deep reasoning enabled for routine workloads could end up paying substantially more than expected.

Pricing is another area where Alibaba is trying to stand out. According to its official Model Studio rate card, Qwen3.8-Max costs $2 per million input tokens (the text you send to the model) and $6 per million output tokens (the text the model generates). Cached reads, where the model reuses previously processed context instead of reading it from scratch, cost just $0.25 per million tokens using implicit caching and $0.17 with explicit caching. More importantly, Alibaba applies the same pricing across the model's entire one-million-token context window, without charging extra for very long prompts. Most frontier AI models increase costs as prompts become longer, making this flat pricing structure an unusual and potentially significant advantage for developers building long-running AI agents.

OpenAI quietly revealed its next move too

Almost unnoticed, OpenAI also hinted at its next flagship model. Instead of a product launch, the company mentioned Astra in the third paragraph of a mathematics research blog, claiming the model had generated new results across several long-standing problems in mathematics and theoretical computer science.

Like Alibaba, OpenAI's evidence is impressive but largely self-produced. While the mathematical proofs are publicly available for verification, researchers have questioned how many problems were attempted, what role humans played, and how the experiments were conducted.

The bigger trend

Looking past the headlines, both Alibaba and OpenAI are pushing the industry toward the same destination.
Neither company is talking about faster chatbots anymore. Instead, they're trying to convince developers that their models can work independently for hours or even days on complex tasks. The challenge is that both companies are also asking the industry to trust evidence they've generated themselves.

Independent evaluations for Qwen are expected soon, while Astra may take much longer to assess. Until then, the safest approach for enterprises remains the same: ignore the marketing tables, wait for third-party benchmarks, and test the models on your own workloads before making deployment decisions.
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › Magazines › Panache › Qwen 3.8-Max launches with 2.4 trillion parameters. Don't trust the benchmarks just yet
Text Size:AAA
Success
This article has been saved

*

+