Language-Dependent AI Bias
Language-dependent AI bias is the pattern in A hawk who flew on political winds: Lindsey Graham where a model’s answers vary depending on the language used to ask the question. The episode links this to the training corpus: Chinese-language data comes from a more controlled internet, while English-language and other corpora carry different cultural and political distributions.
The concept extends AI Model Value Surveying because model values are not only attached to a model name. A single model may show different defaults, refusals, examples, and moral language across languages, making model evaluation harder for multilingual users and regulators.
Key Claims
- The same question can produce different model behavior when asked in different languages.
- Language is not a neutral translation layer; it carries corpus history, censorship conditions, idioms, and cultural examples.
- Bias can appear before post-training if the available text in one language has already been filtered or politically constrained.
- Multilingual model evaluation needs to test language routes separately rather than assume one benchmark result generalizes.
Connections
- AI Model Value Surveying - survey method this concept complicates.
- AI Model Censorship - censorship route through controlled language data.
- Taki AI - historical-corpus example of a bounded text world.
- World Values Survey - cross-cultural comparison frame.
- Human Judgment Under AI - user responsibility to notice context-sensitive answers.