concept Updated 2026-08-07 Tags: Ai, Values, Governance, Training

Model Value Embedding / 模型价值观嵌入

Model value embedding is the claim that a large model’s outputs reflect the data, rewards, post-training choices, product policies, and institutional contexts that shaped it. In 174. 我们还能给算法当多久的品味老师?|对谈亚马逊AGI查晟, 查晟 / Cha Sheng argues that language models are built from human symbols rather than direct access to reality, so their apparent neutrality is a product surface over selected data and reward signals.

The concept complements AI Model Value Surveying. Surveying measures output tendencies after the fact; value embedding asks how those tendencies enter the model through training and governance. The episode uses Anthropic and Claude as a public example where personality, constitution, and value choices are explicitly discussed rather than hidden.

Key Claims

  • Model values can enter through pretraining data, filtering, reward design, post-training, product policy, and deployment context.
  • Neutrality should be treated as an interface claim, not as proof that no values were selected.
  • Value embedding matters more when models advise users, represent companies, or become national information infrastructure.
  • Enterprise and national model ownership are partly value-control strategies, not only cost or accuracy strategies.
  • Alignment governance has to govern the people and institutions shaping model values, not only the final output behavior.

Connections