Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant code. We study vulnerability localization: given a weakness class and an unfamiliar repository, identify the implementation files associated with that weakness.

---
**📖 中文解读**
以上内容由AI翻译自英文原文,可能存在不准确之处。建议阅读[原文](https://arxiv.org/abs/2609.15939v1)获取最准确的信息。

---
🔗 **原文链接**: [Vulnerability Localization Benchmark: Measuring Agentic Secu](https://arxiv.org/abs/2609.15939v1)
🏷️ **转载来源**: ArXiv cs.AI
> 本文由小九AI技术站翻译整理,内容版权归原作者所有。
👤 作者: Aman Priyanshu, Supriti Vijay, Kimia Majd, 何旭红, Fraser Burch, Takahiro Matsumoto, Jianliang He, Baturay Saglam, Arthur Goldblatt, Zhuoran Yang, Amin Karbasi

---
🐾 **小九锐评**

这篇论文来自arXiv预印本,虽然还没有经过同行评审,但选题方向值得关注。
建议先读中文摘要判断是否相关,再看全文细节。
Agent是2026年最卷的方向,没有之一。这篇文章的实操经验够硬。
建议收藏,做Agent开发的时候拿出来翻翻。
AI安全不是遥远的哲学问题,正在变成每个AI开发者都要面对的工程实践。
这篇文章不贩卖焦虑,讲的东西很实在。
Benchmark看多了容易麻木——跑分好不一定产品好用。这篇文章好在对分差有分析,不只是贴数据。

你对这个话题有什么看法?欢迎在评论区讨论 💬

> _转载自 ArXiv cs.AI,内容版权归原作者所有_

---
⏱️ 2026-09-15 14:02