跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分42

TabFM-Auto:为表格基础模型自进化数据流水线

AI 导读

TabFM-Auto 将表格基础模型 TabFM 与语言模型智能体结合,通过迭代优化数据清洗、特征工程、上下文选择和后处理来降低 TabFM 的误差。在 TabArena 全部 51 个数据集上,五种不同智能体与语言模型配置的 TabFM-Auto 包揽前五名,最佳配置将 TabFM 的 Elo 从 1785 提升至 2013。

正文

View PDF HTML (experimental)

Abstract:Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, task descriptions, and auxiliary files that carry dataset semantics. Meanwhile, self-evolving machine learning engineering (MLE) agents train models from scratch on each dataset, yet jointly searching over features, architectures, and hyperparameters is noisy and prone to overfitting. We introduce TabFM-Auto, which pairs a tabular foundation model, TabFM, with a language model agent that evolves the data pipeline around it. Guided by dataset metadata and validation feedback, TabFM-Auto iteratively refines data cleaning, feature engineering, context selection, and post-processing to reduce TabFM's error. Across all 51 datasets of the TabArena benchmark, five TabFM-Auto configurations with different agents and language models take the top five overall positions, and the best raises TabFM from 1785 to 2013 Elo. The discovered pipelines also transfer to other frozen tabular foundation models (+69 to +143 Elo) with no further search. On the 8 tabular competitions of MLE-Bench, TabFM-Auto ranks first overall among MLE agents.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.37989 [cs.LG]
  (or arXiv:2609.37989v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.37989

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Deqing Fu [view email]
[v1] Tue, 29 Sep 2026 16:50:56 UTC (916 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org