China’s ninth annual World AI Conference (WAIC) was held from July 17–20 in Shanghai. Escalating points of tension in the U.S.-China AI competition, including tit-for-tat cybersecurity measures aimed at frontier models and U.S. labs’ reiterated emphasis of the risks posed by China’s AI ascendancy, meant this year’s gathering would be a crucial venue for Beijing to pitch an alternate model for global AI development. Top Chinese leader Xi Jinping’s decision to give the WAIC’s keynote address further emphasized the Chinese government’s focus on the issue. DigiChina invited a group of specialists to weigh in on the implications of Xi’s speech, a new initiative released by the Cyberspace Administration of China and the National Development and Reform Commission, the establishment of the Shanghai-based World AI Cooperation Organization, and the seemingly coordinated release of Moonshot AI’s Kimi K3 model. — Read More
Tag Archives: China AI
Who’s Afraid of Chinese Models?
There’s a story I tell about my first day in STRT-431 at Kellogg School of Management, the introductory class that every first-year MBA was required to take; I leafed through the readings and case studies and was dismayed that there weren’t any tech companies on the docket. Me being me, I spoke to the professor after class wondering why, and was told that the goal of the course was not to necessarily learn about specific industries, but rather to uncover broadly applicable universal principles that could be applied to any company in any industry.
I did not, as I usually tell the story, find this very satisfactory: to me the nature of tech, particularly the fact that software and distribution had zero marginal costs (and zero transaction costs), was something fundamentally different; putting in zeroes in formulas tends to wreak havoc! I soon realized, however, that that was my opportunity. The fundamental insight undergirding Aggregation Theory is that zero marginal costs leads to fundamentally different value chains than people once expected from the Internet: centralization and scale in a world where controlling demand mattered more than distributing supply.
What is fascinating about AI, however, is the extent to which those old universal principles are coming back to the forefront. That was never more apparent than this past weekend, when arguments raged on X about the implications of Kimi K3, another open weights model out of China, approaching the state-of-the-art in terms of capabilities. The long and short of it is this: marginal costs are back in a big way, both in terms of short-term implications of state-of-the-art free models, and in terms of the long-term structure of the industry. — Read More
While Neuralink drills into skulls, China’s BrainCo is betting brain tech will be something you wear
The most visible race in brain-computer interfaces involves surgery. But one of China’s most valuable neurotech firms is deliberately not competing in it, CNBC reports.
BrainCo, based in Hangzhou, builds devices that read the brain from outside the skull. Headbands and caps pick up electrical signals through the scalp, with no operating theatre involved. — Read More
Beijing is looking at curbing overseas access to China’s top AI models
Chinese authorities have held meetings with top tech firms over the past month about potentially restricting overseas access to China’s most advanced AI models, including those yet to be released, three people familiar with the discussions said.
The talks follow a number of steps by Beijing to keep homegrown AI within the country and underscore how China, like the U.S., is now treating cutting-edge artificial intelligence as a critical national asset that needs controls. — Read More
DeepSeek open sources DSpark, a new framework to speed up LLM inference by up to 85%
Even as the geopolitical conversation around AI continues to grow more fraught following the U.S. government’s actions to limit the new models from Anthropic and OpenAI, Chinese open source darling DeepSeek is back with yet another open release that could once again change AI development around the globe.
Over the weekend, the firm released DSpark, a new, MIT-Licensed system designed to make large language models answer faster without changing what the underlying model is trying to say.
… DeepSeek published the work with a technical paper, model checkpoints and DeepSpec, a codebase for training and evaluating speculative decoding systems. The release is available through DeepSeek’s public GitHub and Hugging Face pages, both under the permissive, friendly, commonplace MIT license, making the new technique broadly usable by developers, researchers and commercial enterprise operations that want to study or adapt the approach. — Read More
GLM-5.2: Built for Long-Horizon Tasks
GLM-5.2 is Chinese AI lab Z AI’s latest flagship model for long-horizon tasks.
Supporting long-horizon tasks starts with making long context engineering-usable: the model must maintain quality across long, messy coding-agent trajectories, not just accept more tokens. A 1M context is easy to claim, but much harder to keep reliable under real engineering pressure. To this end, we substantially expanded 1M-context training for coding-agent scenarios, covering large-scale implementation, automated research, performance optimization, and complex debugging. The result is a long-context system that is not only wide in scope, but solid in execution: a practical substrate for sustained engineering work. — Read More
China’s Xiaomi MiMo Is Now 15X Faster Than ChatGPT and Claude
Most people know Xiaomi as the Chinese phone brand. The one that makes cheap electric scooters and air purifiers. Not exactly the company you’d expect to break a major AI inference speed record on a Monday morning.
And yet. Xiaomi just released MiMo-V2.5-Pro-UltraSpeed, a serving mode for its trillion-parameter flagship that hits over 1,000 tokens per second—peaking near 1,200 in demos. — Read More
Huawei looks beyond Moore’s Law
Outside of China, Alibaba is mostly known as an e-commerce titan.
But inside the country, the company is obsessed over catching up to DeepSeek on its development of AI models, and catching up to Huawei on the chips that power them.
When Alibaba’s chip design unit T-Head unveiled its latest AI chip, the Zhenwu M890, last week, it also outlined a multi-year chip roadmap showing how the M890’s future successors would deliver massive performance gains in the next few years. Less than a year ago, Huawei had laid out a similar timeline that ran until 2028. — Read More
Notes from inside China’s AI labs
The Chinese companies building language models are set up as the perfect fast-followers for the technology, building on long-standing cultural traditions in education and work, along with subtly different approaches to building technology companies. When you look at the outputs, the latest, biggest models enabling agentic workflows, and the ingredients, excellent scientists, large-scale data, and accelerated computing, the Chinese and American labs look largely similar. The lasting differences emerge in how these are organized and conditioned.
long thought that a reason that the Chinese labs are so good at catching up and keeping up with the frontier is that they’re culturally aligned for this task, but without talking to people directly I felt like it wasn’t my place to attribute substantial influence to this hunch. Speaking with many wonderful, humble, and open scientists at the leading Chinese labs has crystallized a lot of my beliefs. — Read More
DeepSeek V4—almost on the frontier, a fraction of the price
Chinese AI lab DeepSeek’s last model release was V3.2 (and V3.2 Speciale) last December. They just dropped the first of their hotly anticipated V4 series in the shape of two preview models, DeepSeek-V4-Pro and DeepSeek-V4-Flash.
Both models are 1 million token context Mixture of Experts. Pro is 1.6T total parameters, 49B active. Flash is 284B total, 13B active. They’re using the standard MIT license.
I think this makes DeepSeek-V4-Pro the new largest open weights model. — Read More