Only 17 to 41 percent of answers by Alibaba Qwen, DeepSeek and Moonshot AI Kimi to politically sensitive questions were rated as balanced in a benchmark built by Aleph Alpha, with the rest repeating state doctrine, deflecting or refusing. The test covered 967 hand-picked taboo topics such as Tiananmen, Taiwan and Xinjiang. For companies selecting models for customer-facing use, the result frames model choice as a compliance and reputation decision.

Aleph Alpha test finds Chinese models repeat doctrine on taboo topics

How Aleph Alpha tested Qwen, DeepSeek and Kimi

Aleph Alpha developed its own benchmark and scored responses with its own AI scoring system. The subjects were models from Alibaba, DeepSeek and Moonshot AI, while the question set focused on topics treated as taboo in China. The reported outcome was a wide but low band of balanced answers, from 17% to 41% depending on the model. The remainder fell into three behaviors: repetition of official positions, deflection to other subjects, or refusal to answer. The design matters for buyers because it measures behavior on a defined sensitive set rather than general quality.

The mechanism described in the report has two layers: regulation and training data. China requires public-facing models to reflect "socialist core values", so refusal and doctrinal wording are predictable product behavior rather than accidents. Aleph Alpha also points to training-data lineage: Nvidia Nemotron Cascade 2 showed party-line patterns in 17% of responses, which the report links to roughly 3,500 of its 9.3 million training examples generated with DeepSeek and Qwen. When asked to draft a speech for recognition of Taiwan, that model refused and produced a patriotic defense of the One-China principle. Distilled outputs can thus transfer political alignment into otherwise unrelated systems.

The findings fit earlier audits and recurring anecdotal reports rather than standing alone. An earlier study by the Central European Institute of Asian Studies found a similar spillover effect when terms such as human rights, opposition or surveillance appeared. In those cases models often returned Beijing talking points such as the "principle of non-interference in internal affairs" and a "community with a shared future for mankind". A concrete example cited is Qwen 3.6 answering a question about censorship in the United States and then closing with a defense of China on global internet governance. The commercial backdrop is direct: Aleph Alpha positions itself with Cohere as sovereign AI for governments, while Nvidia pushes its models into government and enterprise.

What this means for companies choosing AI models

For businesses, the first consequence is geographic and use-case screening of models. A system that answers correctly on product questions can still insert political framing into answers about rights, governance, censorship or history, which is visible to customers, partners and regulators. Large organizations with operations in several jurisdictions face higher exposure because the same assistant may serve users with different expectations of neutrality. Smaller firms using a single off-the-shelf model have fewer checkpoints, so the choice of base model and its system instructions carries more weight. Procurement should therefore test sensitive and adjacent prompts before deployment, not only domain tasks.

The second issue is evaluation discipline and vendor questioning. Aleph Alpha has a commercial interest in differentiating its models from Chinese competitors, and it used its own scoring system, so buyers should treat the 17 to 41% figures as one vendor benchmark rather than an independent standard. What the news does not establish by itself is how a given model behaves after enterprise fine-tuning, retrieval grounding, or strict system prompts. Before signing, ask which training sources were distilled from third-party models, how refusal lists are maintained, and whether political-topic behavior can be audited separately from safety filters. A pilot with logged edge-case prompts gives more signal than marketing claims about neutrality.

The marker to watch is whether European sovereign models close the performance gap enough to win broader adoption, or whether buyers accept a choice between US and Chinese value systems embedded in models. Follow enterprise and government tenders involving Aleph Alpha, Cohere and Nvidia, plus independent replications of the 967-topic test. If independent audits confirm the spillover into unrelated answers, content controls and model provenance will become standard procurement clauses.