HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分43
BIABench:评估 AI 智能体在真实生物图像分析任务上的表现
AI 导读
研究者推出 BIABench,一个由 16 项已发表生物学研究重建而成的基准,用于端到端评估 AI 智能体完成真实生物图像分析的能力,任务覆盖 11 类分析子任务,从 H&E 组织学到单分子定位显微成像。
正文
Abstract:Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so an agent must choose and run an analysis through code, specialized software and rendered views. Published studies make this capability testable, because each pairs raw images with a peer-reviewed result. We introduce BIABench, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth. The tasks span eleven analysis subtasks and modalities from H&E histology to single-molecule localization microscopy. Each submission receives an outcome score, which compares the output files with the ground truth using field-standard metrics, and a process score, in which a vision-language model judges method choice and quality control against an expert-written rubric. We evaluated general-purpose and biology-specific agents across several language models, with repeated runs of every task. Routine two-dimensional tasks were solved well, but on some tasks that added a third dimension or a time axis no agent scored above 0.19. Neither biological specialization, stronger models nor detailed expert instructions closed this gap. The agents were also unreliable, with scores varying more between repeated runs of one agent than between different agents, and without ground truth a correct run could not be told from a wrong one by its process score or by the time spent. Released openly with its data and code, BIABench provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
| Comments: | 41 pages, 6 figures, 11 tables |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.34274 [cs.AI] |
| (or arXiv:2609.34274v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34274 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zixuan Pan [view email]
[v1]
Mon, 28 Sep 2026 04:22:51 UTC (12,510 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org