Model Value Embedding / 模型价值观嵌入
Model value embedding is the claim that a large model’s outputs reflect the data, rewards, post-training choices, product policies, and institutional contexts that shaped it. In 174. 我们还能给算法当多久的品味老师?|对谈亚马逊AGI查晟, 查晟 / Cha Sheng argues that language models are built from human symbols rather than direct access to reality, so their apparent neutrality is a product surface over selected data and reward signals.
The concept complements AI Model Value Surveying. Surveying measures output tendencies after the fact; value embedding asks how those tendencies enter the model through training and governance. The episode uses Anthropic and Claude as a public example where personality, constitution, and value choices are explicitly discussed rather than hidden.
Key Claims
- Model values can enter through pretraining data, filtering, reward design, post-training, product policy, and deployment context.
- Neutrality should be treated as an interface claim, not as proof that no values were selected.
- Value embedding matters more when models advise users, represent companies, or become national information infrastructure.
- Enterprise and national model ownership are partly value-control strategies, not only cost or accuracy strategies.
- Alignment governance has to govern the people and institutions shaping model values, not only the final output behavior.
Connections
- AI Model Value Surveying - output-side measurement of model value tendencies.
- AI Advice Moral Outsourcing, AI Model Censorship, and Language-Dependent AI Bias - downstream problems when model defaults enter advice or public information.
- AI Alignment Governance - institutional governance frame for value selection.
- Enterprise Owned Models and Sovereign AI Models / 主权AI模型 - ownership routes for controlling data, value expression, and representation.
- Anthropic and Claude - example model provider and model named in the source.