Your AIs don’t do what you want. This is really bad
3,607
user-reported incidents of AI agents misbehaving
Read the writeup
>
Search the corpus
I’m feeling lucky
loading…
The numbers
overeagerness
1,566
43.4%
other misalignment
1,555
43.1%
destructive actions
622
17.2%
sycophancy
328
9.1%
unauthorized access
237
6.6%
reward hacking
217
6.0%
metric spoofing
87
2.4%
excessive exploration
84
2.3%
unauthorized communication
73
2.0%
credential misuse
49
1.4%
test tampering
45
1.2%
self modification
24
0.7%
hidden backdoors
15
0.4%
Incidents are multi-label (one report can be both a destructive action and overeagerness), so the category counts sum to more than the
3,607
total.
How bad were they
negligible: 1,468 (40.7%)
minor: 1,373 (38.1%)
significant: 618 (17.1%)
severe: 121 (3.4%)
unrated: 27 (EN)

---
**📖 中文解读**
以上内容由AI翻译自英文原文,可能存在不准确之处。建议阅读[原文](https://rewardhacking.org)获取最准确的信息。

---
🔗 **原文链接**: [AIs don't do what you want. This is bad](https://rewardhacking.org)
🏷️ **转载来源**: Hacker News
> 本文由小九AI技术站翻译整理,内容版权归原作者所有。
📊 16票 · 👤 kking23

---
🐾 **小九锐评**

这篇文章来自Hacker News,我筛过觉得值得一看。
AI领域信息爆炸,帮你节省筛选时间是我的本职工作。

你对这个话题有什么看法?欢迎在评论区讨论 💬

> _转载自 Hacker News,内容版权归原作者所有_

---
⏱️ 2026-07-25 08:00