The short answer

The usable lesson from Apsara 2026 is a sorting method: separate listed prices from stage promises. Qwen3.8-Flash bills cache hits at 0.1 RMB per million tokens, published and effective today. Before chasing the next chip, measure what share of your AI bill is the machine re-reading what it already read.

0.1 RMB Cache-hit price per million tokens — Qwen3.8-Flash (listed on Aliyun)
20 GW Global data-center target by 2032 — a stage promise, not built capacity
3x Zhenwu V900 performance vs the M890 — the manufacturer's own claim

Between one polished demo and the next, Alibaba CEO Wu Yongming (Eddie Wu) said the opposite of what the room expected. The application the whole industry cheers today, AI coding, is nothing better than a light bulb from 1882. He didn't prove that with a benchmark table. He told the story of children with progeria, a rare aging disorder.

The opening speech at the Apsara Conference on September 22, 2026 is worth your time not for the superintelligence talk. It's worth reading because it sorts itself cleanly into two baskets: listed numbers you can act on, and stage promises you can't yet.

TL;DR

  • Wu renamed the game. Not "AI that behaves like people" but machine intelligence, a different species that handles under 3% of all human thinking now and aims for a thousandfold more. That's a strategic analogy, and no measurement method was published with it.
  • The number an owner can use today: the cache-hit price of Qwen3.8-Flash, 0.1 RMB per million tokens. Machine thinking at length is getting cheaper in the literal, invoice-level sense.
  • The job this week isn't re-platforming around the V900 chip or a planned five-to-ten-trillion-parameter model. It's measuring how much of your AI spend covers a machine reading the same material twice.

Why a Chinese keynote deserves your attention

You probably don't sell anything to Alibaba, and you may not run a Qwen model anywhere. Doesn't matter. Your monthly AI bill, the content drafting, the customer-reply bots, the automation runs, gets metered in exactly the thing this speech is about: the unit price of one machine thought.

When that unit price falls fast, the border between "work worth delegating to AI" and "work that isn't" moves under your feet. The progeria story later in this piece is the lens for figuring out which side of that border your own backlog sits on.

"Machine intelligence" — swap the ruler before the product

What sets this year's speech apart from ordinary AI keynotes: Wu refuses the question itself. Is AI like a human? Wrong frame, he says. The machine isn't a copy of us. It's a different species.

He borrows the argument from the history of engines. When steam first appeared, it mostly replaced people and horses pulling loads, so everyone measured it in horsepower. Then trains, planes and rockets did things no horse ever did. The machine was never an imitation of the horse. It ran on a different mechanism entirely.

Wu applies that same frame to AI:

What machines do today, helping people write code and reports, is only the early stage of machine intelligence.

Two numbers came with the argument, and both need their conditions read alongside them:

Reference The claim Evidence status
Physical work Machines already carry 99.9% of it Historical analogy, not a 2026 measurement
Thinking work today Machines do under 3% of total human thought His figure; no measurement method published
The target Machines at 1,000 times human thinking Strategic extrapolation, not a forecast with error bars

Our advice: use this set of numbers the way you'd use an investor's thesis, for direction, and don't paste any of it into a financial model. Wu's own arithmetic, 1,000 divided by 0.03, hands him "tens of thousands of times of room left." That looks clean on a slide. As a methodology it's fragile, and nobody claimed it was data.

The progeria question: how many important jobs never got a budget?

This is the passage that made us stop and reread.

Wu's example: infants born with progeria, the early-aging syndrome, number only a few dozen new cases worldwide each year. The total patient record, patients ever documented, runs to a few hundred people. No pharmaceutical company will spend more than a decade and a fortune developing a treatment for that many patients. Not because the medicine would be worthless. Because it can't clear a P&L.

He then draws the conclusion that's more interesting than anything in his technical section. Once AI reaches field-expert research quality and compute gets cheap enough, "every small domain can mobilize millions of agents to research, simulate and verify continuously." The rarest kind of thinking moves from luxury good to bulk commodity.

For a business owner, that gives you a selection test better than any superintelligence slogan: stop automating only the work that already has a budget. The real value zone for cheap machine thought sits in tasks that share three traits. The sample is too small, the cycle is too long, and the return is too thin against an expert's hour. The cost of nobody doing those jobs, society keeps paying it anyway, in slow dribbles.

We spot that same pattern in our own work. The defect catalogue for one rare line of furniture, which nobody ever compiled. On-chain data hypotheses too thin for any fund to staff. Legal playbooks for niche industries that no one has standardized. Humans skip them because they lose money. A machine priced by the token isn't so sure.

Vibe coding is the 1882 light bulb — a warning, not a compliment

Here's Wu's circuit, wired step by step. In 1882, Edison's Pearl Street station lit roughly 400 bulbs nearby. Back then, the utility sold you the bulb and threw in the electricity. That's uncomfortably close to AI platforms today that hand you free tokens when you buy the agent.

Then came 1902 and air conditioning, and after it the washing machine, the refrigerator. The first electronic computer arrived in 1946, more than sixty years after the bulb. And this is the detail that lands hardest: electric appliances had entered homes by 1900, yet the world's entire annual electricity output of that year would cover roughly two hours of our consumption today.

The takeaway Wu pulls from it, the anti-hype line in a conference stuffed with demos:

AI coding today may be like that early lamp in the first days of machine intelligence: it replaces work that already existed. Replacing old jobs alone never creates a new era.

Which means the product that will actually define this period, the refrigerator the bulb-holder couldn't have imagined, hasn't appeared yet. The practical read cuts both ways. Sellers and buyers alike: don't mistake one more week of coding productivity for owning the era's defining product. Equally, don't rush to call it a bubble, because standing in "the year 1900 of intelligence," the infrastructure always looks oversized relative to demand you can see from here.

Three cornerstones: which numbers enter the ledger, which stay on the slide

Wu closed by positioning Alibaba's entire investment plan in one line:

Tokens are the electricity of the AI era. Chips are the generators. Cloud is the grid that carries tokens to every outlet.

All three layers, with our confidence read on each:

Layer The stated number Can you use it yet?
Models Qwen3.8-Flash list price: 0.8 RMB input, 2.7 output, 0.1 cache hit per million tokens Yes, today — published on the Aliyun platform
Models RSI direction (recursive self-improvement), training a five-to-ten-trillion-parameter model Intent signal, not weights you can download and run
Chips Zhenwu V900: 3x the M890's performance, a cluster architecture scaling to 500,000 cards, mass production slated for Q1 2027 Stage roadmap, no independent benchmark; the 500,000 is a design ceiling, not a built cluster
Cloud More than 20 GW of global data centers by 2032 Hyperscaler-scale ambition, not capacity under construction

Keep the red flags visible: "the strongest AI chip in China" is the manufacturer's own announcement. Numbers circulating around the conference like "the model ran 33 self-directed iterations and cut module area 42%" or "cost down to 8% of comparable commercial models" exist as stage claims or secondary reporting, and no independent test suite has been published for any of them.

Worth noting on its own: Qwen's cache pricing never appears in the keynote itself. The rate card was published on Aliyun the same day, and that's exactly the kind of detail you should read separately from the speech. The broader collapse in the unit price of machine thought is the pattern we've been tracking here, including DeepSeek's permanent 75% price cut and what it means for AI democratization.

The 5ac read: take two things from the speech, leave the rest

We have no internal data to validate RSI or the V900. What we can do today is compare listed prices against our own operations, and that alone is worth an afternoon.

In the agent systems we run daily, cost has stopped being a question of which model is smarter. The competition moved to who can run real work cheaply, the same front our AI price war piece on GPT-5.5 and Vietnamese enterprises documented with numbers. Most tokens in our working day are the machine re-reading what it already read: the same knowledge base, the same toolset, the same process, over and over. That's precisely the spend a cache rate card attacks, and mixing DeepSeek and Gemini to optimize agentic workloads targets the same line item. If your agent pipeline holds a cache-hit rate above 60-70%, every future dollar of your AI bill rides this pricing curve down. You don't need to wait for superintelligence to collect the discount.

The second layer gets said less often. When the unit price of thinking drops this far, choosing the job matters more than choosing the model. For comparing specific model candidates the ROI-based routing framework we laid out for DeepSeek, Claude Code and Codex still holds up. But the list of work worth delegating stops being "what are we doing by hand," and becomes "what has never had anyone hired to do it."

Frequently asked questions about the Apsara 2026 keynote

What did the Apsara 2026 keynote actually say?

Alibaba CEO Wu Yongming's opening speech at the Apsara Conference on 22 September 2026 renamed the frame from "AI that behaves like people" to machine intelligence: a different species that handles under 3% of all human thinking today and aims for a thousandfold more. He used the progeria story and the 1882 light bulb to argue that AI coding today only replaces work that already existed.

How much does a Qwen3.8-Flash cache hit cost?

A Qwen3.8-Flash cache hit costs 0.1 RMB per million tokens, roughly 360 Vietnamese dong. That is one eighth of the 0.8 RMB standard input rate, and it is published on the Aliyun platform rather than promised on stage. If your agent pipeline holds a cache-hit rate above 60-70%, every future dollar of your AI bill rides this pricing curve down.

What is the Zhenwu V900 chip?

The Zhenwu V900 is the AI chip announced at Apsara 2026, quoted at three times the M890 performance with a cluster architecture scaling to 500,000 cards and mass production slated for Q1 2027. It is a stage roadmap with no independent benchmark published; the 500,000 figure is a design ceiling, not a built cluster. Don't re-platform on the strength of that number.

What does the 1882 light bulb comparison mean for vibe coding?

Wu compared AI coding today to a light bulb from 1882: it replaces work that already existed, the way the bulb replaced oil lamps, and that alone never creates a new era. The product that will define the machine intelligence period, the refrigerator a bulb-holder could not have imagined, has not appeared yet. It is a warning to both sellers and buyers of AI.

Sources

Editorial note: the opening figures (3%, 99.9%, the thousand-fold target) are strategic analogies presented by Alibaba's CEO, with no published measurement methodology behind them. This article deliberately separates stage claims from listed API pricing; technical decisions should rest on the second group only.