EXAONE Tabular

EXAONE Tabular

by LG AI Research · Current; released for research and education

A compact LG AI Research foundation model for tabular classification and regression that uses labeled examples as context instead of dataset-specific parameter updates. It supports missing-value scenarios, offers separate classification and regression checkpoints, and is available for research and education with no published hosted API pricing.

Reasoning Coding
EXAONE Tabular brings foundation-model techniques to structured data rather than text, images, or audio. Its separate classification and regression checkpoints use the Cross-Axis Summary Transformer architecture and are pretrained entirely on synthetic tabular datasets. The result is a low-retraining workflow for applying a pretrained model to practical datasets that may include missing values, with no published hosted API pricing or commercial license for the released weights.
Model profile

Performance characteristics

5/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family EXAONE Tabular
Model type Other
Context window tokens
Maximum output tokens
Release date 2026-09-09
Status Current; released for research and education
Model notes

EXAONE Tabular is a model family-level release containing separate EXAONETabularClassifier and EXAONETabularRegressor checkpoints. It uses the Cross-Axis Summary Transformer architecture and in-context learning with labeled support rows and unlabeled query rows. Classification has approximately 20.8 million parameters and regression approximately 21.1 million. The native classification head supports up to 10 classes; larger class spaces use ECOC-based decomposition. The inference runtime is BSD-3-Clause-LG AI Research licensed, but the released weights are separately licensed for non-commercial research and education under the EXAONE AI Model License Agreement 1.2 - NC. No hosted API pricing, token context window, maximum output-token limit, or knowledge cutoff is published for this release.

Model guide

EXAONE Tabular: In-Context Classification and Regression for Structured Data

EXAONE Tabular is LG AI Research’s compact foundation model for tabular classification and regression. It predicts outcomes from labeled support rows and unlabeled query rows at inference time, using in-context learning instead of dataset-specific parameter updates or retraining.

What is EXAONE Tabular?

EXAONE Tabular is a tabular foundation model from LG AI Research for predicting values from structured datasets. It supports two main tasks: classification, where the model selects a category, and regression, where it predicts a numeric value. Separate checkpoints are provided for these tasks: an EXAONETabularClassifier and an EXAONETabularRegressor.

The model uses in-context learning. Instead of first fine-tuning model weights on each new dataset, an application supplies labeled support rows as examples and then supplies unlabeled query rows for prediction. The model uses the support examples as context during inference and returns predictions in a forward pass. This can reduce the need for dataset-specific parameter updates, although the quality of the result still depends on how representative the support rows are and how the data is prepared.

EXAONE Tabular is therefore not a general-purpose chatbot or a text-generation model. Its purpose is structured-data prediction, such as estimating a maintenance outcome, assigning a risk category, or predicting a continuous measurement from field data.

Where EXAONE Tabular fits in the EXAONE lineup

EXAONE Tabular extends LG AI Research’s EXAONE portfolio beyond language and vision models. It is aimed at structured, row-and-column data rather than natural-language prompts or visual inputs. This positioning makes it relevant to organizations that have substantial operational, manufacturing, financial, scientific, or clinical datasets but want to avoid building and retraining a separate neural network for every prediction task.

It should not be confused with LG AI Research’s general language and vision-language releases. EXAONE Tabular does not provide text generation, image understanding, speech processing, or general-purpose agent behavior. Its narrower scope is also its main practical distinction: the model is designed around tabular examples and predictions.

How the model works

EXAONE Tabular is built with the Cross-Axis Summary Transformer, or CAST. In simple terms, the architecture examines tabular data in two complementary directions. It can compare features within an individual row, while also comparing support examples within an individual feature. Item-summary and feature-summary tokens help the model exchange information between these two views without discarding cell-level representations.

The published configuration uses a 192-dimensional embedding space, six attention heads, twelve transformer layers, four-times feed-forward expansion, three item-summary tokens per item, and 32 feature-summary tokens per feature. The classification checkpoint has approximately 20.8 million parameters, and the regression checkpoint has approximately 21.1 million parameters.

The model was pretrained entirely on synthetically generated tabular data. LG AI Research describes synthetic data generation based on structural causal models, causal graphs, nonlinear mechanisms, categorical transformations, noise patterns, and missing-value scenarios. The stated purpose is to expose the model to a broad range of data-generating structures so it can transfer to datasets that were not part of a single industry-specific training corpus.

Supported tasks and input behavior

The primary supported tasks are tabular classification, point regression, and predictive-distribution tasks. The runtime provides scikit-learn-style fit, predict, and, for classification, predict_proba interfaces. In this setting, fit refers to preparing the support and query data for inference rather than conventional training that updates the model’s parameters.

The model is intended to work with structured rows containing numerical and categorical information. The research describes support for missing-value scenarios, which is useful for field data where sensors, forms, or business records are incomplete. The supplied material does not specify a universal maximum number of rows, features, or tokens. Instead, practical capacity can depend on the configured inference limits, available memory, and the way the runtime chunks query data.

Classification natively supports up to ten classes. For larger class spaces, the implementation uses an error-correcting output-code decomposition that performs multiple binary predictions. This can extend applicability to larger label sets, but it also increases inference work compared with the native classification path.

Performance and efficiency

LG AI Research reports results on TabArena, BCCO, TALENT, and ScoringBench. On the reported TabArena configuration, the overall Elo score was 1,755, with classification Elo of 1,759 and regression Elo of 1,883. The reported median prediction cost was 0.605 seconds per 1,000 samples. These are provider-reported benchmark results, not independent editorial measurements, and actual performance can vary with hardware, dataset shape, support-set size, and inference settings.

The model’s approximately 21-million-parameter size is relatively compact compared with larger tabular foundation models. That can make it attractive when inference efficiency and reduced retraining work matter more than using a much larger model. However, the supplied research does not establish that it will be the fastest or most accurate option for every dataset. Conventional gradient-boosting or purpose-built statistical models may remain more appropriate for some small, stable, or highly specialized datasets.

The runtime requires Python 3.11 or newer and PyTorch 2.6 or newer. A CUDA-capable GPU is strongly recommended because the implementation uses fused attention kernels and half precision. CPU inference is supported, but the expected trade-off is lower speed. When query data is processed in chunks, support representations may be recomputed for each chunk by the default wrapper, which can add work for larger prediction jobs.

Main strengths and limitations

Strengths

  • Low-retraining workflow: labeled support rows can be supplied at inference time instead of updating model parameters for every dataset.
  • Broad synthetic pretraining: the training approach covers varied causal structures, nonlinear relationships, categorical data, noise, and missing values.
  • Compact architecture: the classification and regression checkpoints contain approximately 20.8 million and 21.1 million parameters respectively.
  • Two important tabular tasks: separate checkpoints address both categorical prediction and numeric prediction.
  • Practical runtime interface: the scikit-learn-style API can be familiar to Python users working with machine-learning pipelines.
  • Research coverage: LG AI Research reports evaluation on several tabular benchmarks rather than presenting the model only as a conceptual release.

Limitations

  • Specialized scope: it does not handle unstructured text, images, audio, or video and is not a generative language model.
  • No general-purpose generation: it does not provide text, image, speech, or music generation, embeddings, web search, or general-purpose tool calling.
  • Class-count constraint: the native classifier supports up to ten classes; larger label spaces require the more expensive ECOC-based approach.
  • Hardware expectations: GPU acceleration is strongly recommended, while CPU use may be considerably slower.
  • Inference-size considerations: large support sets may be subsampled when they exceed configured limits, and chunked query processing can repeat support computation.
  • Limited service information: no hosted API price, token context window, maximum output-token limit, or standard managed endpoint is published in the supplied research.
  • Commercial restrictions: the released weights are limited to non-commercial research and education use under the EXAONE AI Model License Agreement 1.2 - NC.

Modalities, reasoning, coding, and tools

EXAONE Tabular accepts structured tabular data and produces prediction outputs. It has no documented text, image, audio, or video input mode, and its output is not direct image, audio, video, or natural-language generation. The appropriate output types are class predictions, class probabilities where supported, regression values, and predictive-distribution results.

It performs task-specific statistical reasoning over relationships among rows and features, but it should not be described as a general reasoning model in the language-model sense. The supplied research also does not document code generation or code execution. A Python runtime is required to use the implementation, but that is a deployment requirement rather than a model capability. Likewise, there is no documented function-calling, MCP, web-browsing, or agent-tool interface for EXAONE Tabular.

Licensing, pricing, and availability

The runtime is released under the BSD-3-Clause-LG AI Research License. The model weights are governed separately by the EXAONE AI Model License Agreement 1.2 - NC. According to the supplied research, the weights are available for non-commercial research and education. Commercial deployment requires separate licensing or permission from LG AI Research.

The runtime and checkpoints are available through the official EXAONE Tabular GitHub repository and the Hugging Face model repository. No hosted API pricing is published for this release, so there is no verified per-request, per-token, monthly, or usage-based price to report. Users should also account for their own GPU, hosting, storage, and engineering costs when evaluating deployment.

When to choose EXAONE Tabular

EXAONE Tabular is a reasonable choice when the main problem is classification or regression on structured data and the team wants to provide representative labeled examples at inference time rather than maintain a separate fine-tuning process for every dataset. It may be especially useful for research, experimentation, and field-data applications involving missing values, changing datasets, or many related prediction tasks.

It is less suitable when the workload requires text or image understanding, generative reporting, web-connected agents, or a managed commercial API. It may also be a poor fit when a CPU-only environment must process large volumes with strict latency targets, when the class count is very high, or when commercial licensing of the released weights is essential. In those cases, a conventional tabular model, a commercial hosted service, or another model designed for the required modality and licensing conditions may be more appropriate.

Overall, EXAONE Tabular’s practical value is its combination of compact size and in-context tabular prediction. Its strongest use case is not replacing every established machine-learning method, but providing a pretrained research model for structured-data tasks where reducing dataset-specific retraining is important and the licensing and hardware requirements can be met.


Answers to Frequently Asked Questions

What is EXAONE Tabular used for?
EXAONE Tabular is a tabular foundation model from LG AI Research for structured-data prediction. It supports classification, regression, and predictive-distribution tasks using numerical and categorical rows.
What are the main limitations of EXAONE Tabular?
The model is specialized for tabular data and does not support text, image, audio, or video understanding or generation. Native classification supports up to ten classes, GPU acceleration is strongly recommended, and large support sets or chunked queries can increase inference costs.
Is EXAONE Tabular available for commercial use?
The runtime is released under the BSD-3-Clause-LG AI Research License, while the model weights are governed by the EXAONE AI Model License Agreement 1.2 - NC. The supplied research states that the weights are available for non-commercial research and education; commercial deployment requires separate licensing or permission from LG AI Research.
What hardware and software are required to run EXAONE Tabular?
The runtime requires Python 3.11 or newer and PyTorch 2.6 or newer. A CUDA-capable GPU is strongly recommended because the implementation uses fused attention kernels and half precision, although CPU inference is supported at lower expected speed.
How does EXAONE Tabular use in-context learning?
An application provides labeled support rows as examples and unlabeled query rows for prediction. EXAONE Tabular uses the support rows as context during inference, so model parameters do not need to be fine-tuned for every new dataset.


Sources 4
Provider

About LG AI Research