Short answer: SemiAnalysis counted 857 model releases from nine leading Chinese AI developers, from 2021 to September 15, 2026, and found that only 31 of them (3.6%) ever had a safety result published by the developer. Just 9 (1.1%) had that result available at or before release. The other 813 releases, 94.9% of the total, have no safety disclosure at all. The report, Beijing Will Not Pace the Frontier by Mark Chen, Doug, and Dylan Patel, was published on October 8, 2026 and is partly paywalled; this post covers the public portion.
The census lands in the middle of a live argument. On September 12, Dario Amodei published an essay arguing that frontier labs should deliberately slow capability progress; we covered the plan in Dario Amodei wants to "pace the frontier". SemiAnalysis asks the obvious next question: if the US side slowed down, would anyone in China follow? Its answer is no, and the release census is its evidence.
A clipboard with checkmarks and a green shield, representing published safety evaluation results for AI models
TL;DR: the census in one table
| Question | Answer |
|---|---|
| What was counted? | 857 releases (741 product, 116 research) from nine Chinese labs, 2021 to Sept 15, 2026 |
| Which labs? | ByteDance, Alibaba, Tencent, Baidu, DeepSeek, Moonshot, Zhipu Z.ai, MiniMax, StepFun |
| Releases with any safety result | 31 (3.6%) |
| Result available at or before launch | 9 (1.1%) |
| Documented after launch | 16, median lag 42 days, maximum 349 days (DeepSeek-R1) |
| Timing or model match unclear | 6 |
| Evaluation claim with no figures | 10 |
| Known only from press or investor accounts | 3 |
| No safety disclosure | 813 (94.9%) |
| Reasoning models | 93% undocumented |
How did SemiAnalysis define a safety result?
The definition is strict, and it drives the headline number. A release counts only if the developer published a quantitative or substantive finding on harmful output, jailbreaks, toxicity, privacy, refusal, or dangerous capability, tied to the named model. A sentence saying a model was "safety-trained" or "evaluated" does not count.
The researchers checked each release against its model card, release notes, and technical report. They are explicit about the limits: per-company rates are indicative rather than a ranking, and "not found" means not found in the materials checked, not "not tested." Hold that caveat in mind for the rest of this post. The census measures what a reader can verify from public documents. It cannot see what happens inside a lab.
What does the lab-by-lab picture look like?
The public portion of the report names several examples.
Zhipu is the outlier. SemiAnalysis says it is the only developer with a result in every year since 2022 and the only lab where safety is a named strategic commitment. Of nine leader statements that express concern and propose something, six come from Zhipu alone. CEO Zhang Peng is quoted as saying "defense is always harder than attack," and founder Tang Jie wrote that "the stronger the capability, the more robust the safety constraints must be." For the product side of the same lab, see our post on Zhipu's GLM-5.3 vision work and Jie Tang.
DeepSeek documented V3 at launch but not R1, V3.1, or V4. The R1 safety results arrived 349 days after release, the longest delay in the census. DeepSeek founder Liang Wenfeng has no public safety statement, per the report. A DeepSeek researcher, Liu Shengyu, wrote that he has "no choice but to join the cruel arms race." The V4 launch itself is covered in DeepSeek V4 official release and peak pricing.
Alibaba has 7 of 238 releases with any result, 3 of them at launch. SemiAnalysis notes that the keynote by Eddie Wu does not contain the word "safety" or "risk."
Tencent has 1 of 133 releases with a result. ByteDance has 2 of 120. Baidu's Robin Li has said AI "could indeed develop in directions unfavorable to humanity," and Moonshot's Yang Zhilin said "we should not give up its development."
Across the nine labs, at five of them the founder or CEO has said nothing on frontier safety, based on the 65 leader statements SemiAnalysis compiled, of which 15 were frontier-safety statements.
A bridge between two landmasses with green dots crossing, standing for the gap between Chinese and US AI safety practice
Does China have AI safety rules?
Yes, and that is the interesting tension in the report. SemiAnalysis calls China's official texts the most safety-forward of any major AI power outside Europe. It cites Framework 3.0 from TC260 and the Cyberspace Administration, dated September 14, 2026, which warns about recursive self-improvement and adds a risk category for behavior deviating from expectations.
But the same framework lists "promoting AI innovation and development as the first priority" as its first principle. And the comprehensive AI Law that appeared in the 2023 and 2024 legislative plans was shelved in 2025, four months after what the authors call the DeepSeek moment. China regulates content and applications, such as labeling, minors, agents, and AI companions, but has no frontier-risk duties tied to training compute or capability.
The report contrasts this with other regimes. The EU attaches systemic-risk duties from 10^25 FLOP, and California's SB 53 requires frontier frameworks and 15-day incident reporting from 10^26 operations. The authors' conclusion: a Chinese lab can satisfy every rule on the list without running a dangerous-capability evaluation.
Other Chinese institutions are active on paper. Shanghai AI Lab's Frontier AI Risk Management Framework (version 2.0, July 2026) sets 13 red lines, and CAICT's AI Safety Benchmark 2.0 (February 2026) covers deception, loss of control, and agent risk. SemiAnalysis points out that CnAISDA, the network meant to coordinate this work, has no staff, budget, or mandate to test models. In its words, these bodies produce frameworks that no one must follow.
What do Chinese experts say?
The authors coded 102 texts: 51 official or legal documents and 51 technical ones. Frontier or loss-of-control risk appears in 86% of the technical texts but only 22% of the law-scholar texts. Twenty-two of 42 technical texts argue that safety must precede or gate development, and no legal or public-policy text takes that position. Thirteen texts call for binding frontier duties such as compute-threshold registration, pre-release safety cases, third-party audits, or shutdown authority. None has been adopted.
There is also open dissent from the other side. Zhu Songchun, speaking at WAIC 2026, called the human-extinction narrative a technical misjudgment packaged for capital, according to the report.
Why does SemiAnalysis say Beijing will not slow down?
The argument has three parts.
- Priority: The state's AI+ plan targets 70% agent and intelligent-terminal penetration by 2027 and 90% by 2030. That is a deployment target, not a pause.
- Structure: Official language has shifted from safety to security, and Framework 3.0 lists export controls as a supply-chain security risk. Safety talk is folded into the competition with the US.
- Reaction: After the Amodei essay, Beijing's state-run Global Times called it a "Cold War playbook." Sam Altman and Elon Musk backed the call to slow down; Trump wrote on September 14 that fears of AI destroying humanity are "a HOAX."
The report places this inside a diplomatic timeline: a US-China AI dialogue on September 20, a Trump-Xi summit that formalized a "US-China Super Intelligence Dialogue," and on September 29 the White House Accord on Super Intelligence, a one-page voluntary pledge signed by Trump and six executives. For the political side, see Trump officials plan AI risk talks with China and the White House task force and Bessent's "slow down".
The implicit conclusion in the public text is that US labs cannot rely on Chinese restraint to justify slowing down, since "nobody paces alone." The recommendations for American labs, investors, and policymakers sit behind the paywall, along with dated tests over the next six months that the authors say would change their conclusions.
How should you read the numbers?
Treat 3.6% as a lower bound on disclosure and an unknown on practice. A few points help.
The denominator is generous. 857 releases include 741 product releases such as point updates, fine-tuned variants, and regional builds, where a fresh safety report is rarely expected. A stricter denominator of flagship models would give a higher percentage. The report's more telling number is that 93% of reasoning models are undocumented, and that no Chinese frontier text model has shipped with a dangerous-capability evaluation across the domains named in the IDAIS statements.
Disclosure norms differ. Western labs publish system cards partly because of voluntary commitments and state laws. The same census run on US labs would be a useful control. SemiAnalysis does not present one in the public portion, so we cannot say how large the gap really is.
Late results count differently. Sixteen results came after launch with a median 42-day lag. A result 349 days after release (DeepSeek-R1) is not much use to a team choosing a model in week one.
Counting is judgment. Ten releases carry an evaluation claim with no figures, and three are known only from press or investor accounts. Different coding choices would move the percentage by a point or two, not change the conclusion.
For context on the compute side of this same competition, SemiAnalysis's data center count is covered in China has 24 GW of data center capacity versus 56 GW in the US, and the model-theft dispute is in China's countermeasures to US accusations of AI model theft.
What does this mean for builders?
If you build on Chinese open-weight models, and many teams do for cost reasons, the practical consequence is that you often cannot read a safety result before you ship.
- Run your own evaluations. For any model with no published result, test jailbreak resistance, refusal behavior, and data handling on your own prompts before production. Treat the absence of a result as the absence of evidence.
- Pin versions and re-check on updates. A launch with no safety result today may get one 42 days later; a point release may change behavior.
- Constrain agents regardless of model origin. Safety documentation does not replace runtime controls. If your agents act on real systems, put policy checks between the model and its tools; AgentBeam, the agent security platform from the explainx.ai team, stops AI agents before they take dangerous actions.
- Read launch posts for numbers, not adjectives. "Safety-trained" without a result is exactly what the census does not count.
Open questions
- Will the Chinese labs publish results for flagship releases in the next six months, the tests the authors say would change their conclusions?
- Does the same census on US and European labs show a gap as wide as 3.6% versus the Western norm?
- Will Zhipu's practice spread, or remain the exception?
- Does a binding frontier-risk rule appear, given that thirteen expert texts call for one and none is adopted?
- How do reactions in Washington change if the Chinese side begins publishing evaluations at launch?
For the wider debate on slowing down, also read the AI employees' pacing the frontier letter, Bengio's call for safety-minded researchers to leave frontier companies, and our AI policy timeline on export controls, distillation and open weights.
A rounded folder with a green flag, representing a safety report folder for Chinese AI model releases
Figures are accurate as of October 10, 2026 and come from the public portion of the SemiAnalysis report; the rest is paywalled. Earlier SemiAnalysis coverage of the Chinese buildout is in The Chinese AI Infrastructure Boom.
Related reading
- Dario Amodei wants to "pace the frontier"
- Trump officials plan AI risk talks with China
- White House AI task force and Bessent's "slow down"
- China has 24 GW of data center capacity versus 56 GW in the US
- Bengio: leave frontier AI companies if you prioritize safety
- DeepSeek V4 official release and peak pricing
- AI policy timeline 2026
