Memory Efficient Tabular Foundation Models
FMSD @ ICML 2026 · workshop paper · arXiv:2607.27546v1
为什么读:它把表格基础模型从排行榜拉回实际部署成本,直接测量压缩、显存与性能之间的交换。
重要材料按首次公开日期倒序排列;版本更新时间单独标记。每篇先进入自己的资料页,集中书目、P11 判断、证据边界、阅读路径与便笺;针对这份资料形成的技术拆解继续放在它的子页面下。核心模型材料另有专门的 Discord 反馈子区。
FMSD @ ICML 2026 · workshop paper · arXiv:2607.27546v1
为什么读:它把表格基础模型从排行榜拉回实际部署成本,直接测量压缩、显存与性能之间的交换。
arXiv · arXiv preprint v1 · arXiv:2607.26628v1
为什么读:P11 的 Unit token 也依赖可见 evidence;这篇论文正面研究 context 数量、覆盖度与随机抽样如何改变稳定性和精度。
arXiv · arXiv preprint v1 · arXiv:2607.25532v1
为什么读:它给出一个非常接近 P11 风险面的负例:同一变量表示内部的 causal 与 spurious subspace 会被 ICL 按当前 context 相关性错误路由。
Google Research · official model and code release; no archival paper located at cutoff · official release
为什么读:TabFM 是当前合成先验、row/column attention 与 row compression 路线的重要系统参照,也迫使 P11 区分模型发布与可审计论文证据。
arXiv · BeyondArena / Data Foundry · arXiv preprint v1 · arXiv:2606.30410v1
为什么读:它扩展到 grouped、temporal、large 与 high-dimensional regimes,是检验“foundation”是否超出 tiny/small IID 的关键反证面。
43rd ICML Workshop on Foundation Models for Structured Data · workshop paper · arXiv:2606.26467v1
为什么读:TabPFN-CFM 直接把 structure、outcome、intervention 与 counterfactual queries 放进同一 causal foundation model,是 P11/DiscoSCM 必须认真比较的最近工作。
ICML 2026 · peer-reviewed paper · arXiv:2606.04485v2
为什么读:它把精度—效率差距定位到 scalar tokenization 的低秩通道与 attention routing,是 P11 value-token 设计最直接的架构压力。
arXiv · arXiv preprint v1 · arXiv:2605.21288v1
为什么读:它不再只比 accuracy,而是干预 TabPFN、TabICLv2 与 Mitra 的表示和 readout;row token、similarity vote 与 invariance 都因此有了可检查前例。
arXiv · technical report v2 · arXiv:2605.13986v2
为什么读:这是 TabPFN 当前尺度、test-time compute、关系表、时间序列与文本扩展的总入口,决定 P11 不能只和早期 small-data TabPFN 比。
arXiv · arXiv preprint v1 · arXiv:2603.26611v1
为什么读:它把 TFM 从 point prediction 推到完整 conditional density,并直接测 calibration,是 P11 typed predictive distribution 的重要评测邻居。
ICML 2026 · open preprint and model release · arXiv:2602.11139v1
为什么读:TabICLv2 把 synthetic generator、scalable attention 和训练 protocol 一起升级,并提供开放权重,是 P11 公开、可执行、可审计的强基线。
arXiv · arXiv preprint v1 · arXiv:2511.09665v1
为什么读:它追问跨表泛化究竟需要多少物理 tables,直接冲击把 table count 当作 pretraining scale 的简单叙事。
arXiv · technical report v2 · arXiv:2511.08667v2
为什么读:它连接 TabPFN v2 与 v3,扩展表格规模并引入 distillation,是理解当前产品级 TabPFN 路线的重要版本节点。
NeurIPS 2025 · peer-reviewed paper · arXiv:2510.21204v1
为什么读:Mitra 把 synthetic prior design 本身变成研究对象,系统比较 SCM 与 tree-based prior 的多样性和混合。
arXiv · technical report v2 · arXiv:2509.03505v2
为什么读:LimiX 把 prediction、imputation、generation 与 retrieval 统一为 masked joint-distribution query,是 P11 当前最重要的上游架构与对象对手。
arXiv · arXiv preprint v1 · arXiv:2507.03971v1
为什么读:它直接比较纯 synthetic pretraining 与少量 curated real tables 的 continued pretraining,暴露真实数据价值和 contamination 风险。
NeurIPS 2025 Datasets and Benchmarks · Spotlight · peer-reviewed living benchmark · arXiv:2506.16791v4
为什么读:TabArena 是当前模型排名最常引用的 living benchmark;validation、tuning budget 与 ensembling 足以改变结论。
arXiv · position paper v2 · arXiv:2505.19825v2
为什么读:它明确指出单表缺少 operational、declarative 与 procedural context,是 P11/DiscoSCM 不能只谈 schema token 的最近立场对手。
NeurIPS 2025 · peer-reviewed paper · arXiv:2505.18125v2
为什么读:TabSTAR 用 target-aware textual representations 处理带文本字段的表,是 P11 从 description 生成 semantic feature token 的直接语义邻居。
arXiv · TARTE · arXiv preprint v2; final publication status unconfirmed · arXiv:2505.14415v2
为什么读:TARTE 把 column names 与 strings 的世界知识压进可复用表示,是 P11 semantic feature construction 的重要前史。
ICML 2025 · peer-reviewed paper · arXiv:2502.05564v2
为什么读:TabICL 首先把 features 压成 fixed-dimensional row embedding,再沿 rows 做 ICL,是 P11 Unit-token 位置与长表扩展的关键架构对照。
Nature 637, 319–326 · peer-reviewed journal article · DOI:10.1038/s41586-024-08328-6
为什么读:这是 TabPFN v2 在 small-data prediction 上形成领域转折的正式论文,也是 synthetic prior + one-forward inference 叙事的基准入口。
NeurIPS 2025 · peer-reviewed paper · arXiv:2410.18164v3
为什么读:TabDPT 用 123 个真实 OpenML tables、retrieval 与 self-supervision 研究 real-data scaling,是 synthetic-only 路线的重要对照。
ICML 2024 · peer-reviewed paper · arXiv:2402.16785v2
为什么读:CARTE 用 column names、entry strings 与 graph attention 实现无 schema matching 的 transfer,是 P11 semantic feature tokens 最重要的早期对照之一。