Study Finds Limited Public Safety Testing for Chinese AI Releases
A review of 857 releases finds model-specific public safety tests for only 3.6% of models in its sample.
Reuters reported October 9 on a SemiAnalysis review of 857 Chinese AI model releases from 2021 to September 2026. The research identified published evaluations for only 31 releases, roughly 3.6 percent, with just nine available at or before launch. Its criteria required tests tied to an identifiable model.
The study measured disclosure, not necessarily whether developers had performed private safety testing. It would be inaccurate to treat a lack of published evaluations as conclusive evidence of an unsafe release. Different jurisdictions also use different obligations for models, providers and downstream deployments.
For enterprise purchasers, disclosure affects the ability to compare model risks, reproducibility, oversight and suitability for tasks with external consequences. Procurement reviews should seek version-specific testing, known failure modes, abuse-prevention measures and evidence that serious incidents are escalated.
The key follow-up is whether developers publish more independently reviewable model reports and whether disclosed evaluations actually predict real-world behavior. Comparisons should use similar test definitions across US, European and Chinese suppliers rather than relying only on broad safety branding.
Reporting sources & references
These links identify the reporting or public materials on which the article is based; they do not imply our newsroom witnessed the events.