concept Updated 2026-07-24 Tags: Ai, Censorship, Governance, China

AI Model Censorship

AI model censorship is the pattern in A hawk who flew on political winds: Lindsey Graham where models refuse, deflect, or give party-line answers on politically sensitive subjects. The episode’s main case is Chinese models answering questions about Tibet, Taiwan, and Tiananmen, which the source calls the “three T’s test.”

The concept is not limited to one refusal message. The source says censorship can enter through post-training, where answers are rated as good or bad, and through language-specific training data drawn from a controlled internet. That links model censorship to Language-Dependent AI Bias rather than treating it only as a visible safety filter.

Key Claims

  • Refusal behavior can hide what a model has learned from the corpus without erasing that underlying information.
  • Post-training can make some answers politically unacceptable even when the base model has relevant knowledge.
  • A censored public internet can shape the training distribution before any explicit refusal rule is added.
  • Censorship tests reveal political alignment and governance pressure, not just technical capability.

Connections