跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分53

TabFM:400M 参数的零样本表格数据基础模型

AI 导读

论文提出 TabFM,一个 400M 参数的表格基础模型,将监督式表格预测建模为 in-context learning,单次前向传播即可产生校准的零样本预测,无需任务特定调参。

正文

View PDF HTML (experimental)

Abstract:Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representations that transfer zero-shot to real-world tasks. Across all 51 benchmark datasets in TabArena (38 classification and 13 regression), zero-shot TabFM ranks first among default tabular foundation models and outperforms tuned AutoML pipelines. Two extensions over the same frozen weights improve performance further on both tracks: multi-view feature expansion with ensembling and post-hoc calibration (TabFM+), and LLM-guided, dataset-specific data processing and feature engineering (TabFM-Auto).
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.37959 [cs.LG]
  (or arXiv:2609.37959v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.37959

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Deqing Fu [view email]
[v1] Tue, 29 Sep 2026 16:30:44 UTC (556 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org