{"id":3407,"date":"2026-08-04T12:10:29","date_gmt":"2026-08-04T16:10:29","guid":{"rendered":"https:\/\/ozer.gt\/log\/?p=3407"},"modified":"2026-08-04T12:10:29","modified_gmt":"2026-08-04T16:10:29","slug":"in-context-learning-for-tabular-data","status":"publish","type":"post","link":"https:\/\/ozer.gt\/log\/2026\/08\/04\/in-context-learning-for-tabular-data\/","title":{"rendered":"In-context learning for tabular data"},"content":{"rendered":"<p>For data scientists working on tabular data, which is still most of the data science world, Google&#8217;s TabFM is a nice solution using transformers (thankfully not LLMs).<\/p>\n<p>TabFM promises to break the tedious cycle of endless hyperparameter tuning and feature engineering in ensemble methods like XGBoost. Instead of a manual training loop, TabFM treats tabular prediction as in-context learning: it ingests the entire dataset as a single prompt and predicts outcomes in a single forward pass. Clearly, the context window becomes the primary bottleneck as the dataset grows.<\/p>\n<p>My day-to-day is mostly causal inference, with some projects using predictive models. Also interesting to me here is how the team solved the training problem: they pretrained the model on hundreds of millions of synthetic tables generated using Structural Causal Models.<\/p>\n<p>This aligns neatly with our posts at <a href=\"https:\/\/www.dataduets.com\">Data Duets<\/a>, specifically around the value of <a href=\"https:\/\/www.dataduets.com\/2026\/02\/using-generative-models-well-to-generate-data.html\">using transformers to generate datasets<\/a> and our broader <a href=\"https:\/\/www.dataduets.com\/2026\/03\/augmented-data-science-human-intent-ai-execution.html\">Augmented Data Science framework<\/a>.<\/p>\n<p><a href=\"https:\/\/research.google\/blog\/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data\/\">Source<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>For data scientists working on tabular data, which is still most of the data science world, Google&#8217;s TabFM is a nice solution using transformers (thankfully not LLMs). TabFM promises to break the tedious cycle of endless hyperparameter tuning and feature engineering in ensemble methods like XGBoost. Instead of a manual training loop, TabFM treats tabular [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3408,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"cybocfi_hide_featured_image":"","footnotes":""},"categories":[1],"tags":[],"class_list":["post-3407","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/posts\/3407","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/comments?post=3407"}],"version-history":[{"count":1,"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/posts\/3407\/revisions"}],"predecessor-version":[{"id":3409,"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/posts\/3407\/revisions\/3409"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/media\/3408"}],"wp:attachment":[{"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/media?parent=3407"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/categories?post=3407"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/ozer.gt\/log\/wp-json\/wp\/v2\/tags?post=3407"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}