论文arxiv cs.CL · 1mo ago需要关注
Small edits, large models: How Wikipedia advocacy shapes LLM values
分类释义:学术论文 / 技术报告
TL;DR
arXiv:2606.24890v1 Announce Type: new Abstract: Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in nearly every major language model training dataset and is weighted more heavily than web-crawled text. The Pro-Animal Wikipedians (PAW), a group of advocates who add sourced animal welfare content to relevant articles, have made 125 edits across 115 pages. Using gradient-based data attribution (Bergson;
关键要点
- 01arXiv:2606.24890v1 Announce Type: new Abstract: Can a small group of volunteers shape how AI systems discuss animal welfare。
- 02just by editing Wikipedia? We show that they can. Wikipedia appears in nearly every major language model training dataset and is weighted more heavily than web-crawled text. The Pro-Animal Wikipedians (PAW)。
- 03a group of advocates who add sourced animal welfare content to relevant articles。
- 04have made 125 edits across 115 pages. Using gradient-based data attribution (Bergson;。
为什么值得关注
对你的工程实践意味着什么
LLM 实时生成MiniMax-M2.7缓存命中
| 角色 | 你应该做什么 |
|---|---|
| Tech Lead | 审查训练数据来源的多元化策略,避免单一内容源过度影响模型价值观 |
| 应用工程师 | 评估依赖Wikipedia内容的RAG或微调方案,建立输出偏见检测机制 |
| 运维 / 平台 | 梳理数据管道中Wikipedia内容的引用比例,建立数据血缘追踪能力 |
| 产品 / 业务 | 了解模型观点可能受少数编辑者影响,在敏感话题上不盲信模型输出 |
同类资讯
本页 TL;DR 与「为什么」由 LLM 生成 · 模型:MiniMax-M2.7 / Claude Haiku 4.5