跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-04-22精选AI 评分73

SWE-chat:来自真实用户与代码智能体在真实环境中的交互数据集

AI 导读

SWE-chat是首个从开源开发者真实工作环境中收集的大规模代码智能体交互数据集,目前已包含6000个会话、超过6.3万条用户提示和35.5万次智能体工具调用。研究发现代码编写模式呈现双峰分布:41%的会话中智能体几乎编写了所有提交代码,而23%的会话完全由人工编写。尽管能力快速提升,代码智能体在自然场景中效率仍不理想,仅44%的智能体生成代码最终被用户采纳,且其代码比人工代码引入更多安全漏洞。用户在44%的交互轮次中会对智能体输出进行修正、报告失败或中断操作。该数据集通过完整记录人机代码归属的交互轨迹,为超越人工基准测试、基于实证理解AI智能体在真实开发流程中的表现提供了基础。

推荐理由

这是目前唯一来自真实开发者的编码代理交互数据集,揭示代理产出的代码存活率仅44%,且引入更多安全漏洞,对于所有推代理的团队都是必读的现实检验。

正文

View PDF HTML (experimental)

Abstract:AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful in practice. We present SWE-chat, the first large-scale dataset of real coding agent sessions collected from open-source developers in the wild. The dataset currently contains almost 18,000 sessions, comprising more than 229,000 user prompts and 2 million agent tool calls. SWE-chat is a living dataset; our collection pipeline automatically and continually discovers and processes sessions from public repositories. Leveraging SWE-chat, we provide an initial empirical characterization of real-world coding agent usage and failure modes. We find that coding patterns are bimodal: in 41% of sessions, agents author virtually all committed code ("vibe coding"), while in 25%, humans write all code themselves. Despite rapidly improving capabilities, coding agents remain inefficient in natural settings. Only 59% of all agent-produced code survives into user commits, and agent-written code introduces more security vulnerabilities than code authored by humans. Furthermore, users push back against agent outputs - through corrections, failure reports, and interruptions - in 50% of all turns. By capturing complete interaction traces with human vs. agent code authorship attribution, SWE-chat provides an empirical foundation for moving beyond curated benchmarks towards an evidence-based understanding of how AI agents perform in real developer workflows.
Comments: Accepted at COLM 2026
Subjects: Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Software Engineering (cs.SE)
Cite as: arXiv:2604.20779 [cs.AI]
  (or arXiv:2604.20779v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2604.20779

arXiv-issued DOI via DataCite

Submission history

From: Joachim Baumann [view email]
[v1] Wed, 22 Apr 2026 17:08:19 UTC (3,861 KB)
[v2] Thu, 1 Oct 2026 17:58:03 UTC (2,107 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org