TabPFN and TabICL vs. tuned XGBoost: the model that doesn't train won 14/14
EfrainGaray
12 points
5 comments
September 28, 2026
Related Discussions
Found 5 related stories in 72.7ms across 7,833 title embeddings via pgvector HNSW
- Choosing an AI model: one prompt, 11 models, different results toddmorey · 192 pts · August 13, 2026 · 49% similar
- Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models adam_rida · 310 pts · July 23, 2026 · 44% similar
- Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models joozio · 71 pts · September 16, 2026 · 43% similar
- GLM-5.3 Artificial Analysis Benchmarks apitman · 114 pts · August 18, 2026 · 43% similar
- Exploring Claude/GPT Knowledge Cutoffs and Pre-Training Timelines sshh12 · 23 pts · August 10, 2026 · 43% similar
Discussion Highlights (3 comments)
rdedev
Tabular foundation models are one of those things where when you first look into it, it does not make sense as to why they would work so well but it does. In drug property prediction domain, tabular foundation models coupled with another foundation model for molecules are pretty close to being the state of art. Btw the article makes heavy use of AI or is written in that way A lot of unnecessary dramatic flair that gets very tiring
3eb7988a1663
XGBoost’s search optimized accuracy, and afterwards I also compare by area under the curve. Which means the “fourteen of fourteen on AUC” is against a boosting model that was not tuned for that metric. Tuning it for AUC would probably improve it there; I did not measure that. So, not a fair test? I also did not get why Xgboost had to count its training time for the inference. You only train once. I guess in some scenario, where someone says, "I need the best model now, you have five minutes on this singular dataset", but I have never been in that situation. I would feel better if scripts were released, because I am fairly dubious. I take it as a given that a tabular model has been pre-trained on all of the public benchmark datasets, but that is what it is. The slop was so meandering, I am not sure what is truth or not.
lyelibi
I have never seen tabular transformer models beat xgboost/catboost in industrial context where datasets is gigantic. They most produce these results on relatively small datasets, clearly not in the tens of millions of rows.