What is EXAONE-Path-2.0-rev-EGFR?
EXAONE-Path-2.0-rev-EGFR is a specialized pathology AI model provided by LG AI Research. Its purpose is to predict the likelihood that a lung adenocarcinoma tissue sample carries an epidermal growth factor receptor (EGFR) mutation, using a digitized hematoxylin-and-eosin (H&E) whole-slide image.
H&E whole-slide images are high-resolution digital scans of tissue sections prepared with a commonly used pathology stain. The model does not read a medical report or conduct a conversation. Instead, it examines visual patterns distributed across the tissue image and returns an EGFR mutation probability score.
The model is a downstream application of the EXAONE Path 2.0 pathology foundation model. In practical terms, EXAONE Path 2.0 supplies the pathology-focused visual representation, while EXAONE-Path-2.0-rev-EGFR adds a task-specific classifier for EGFR prediction. This makes the current model narrower than a general pathology foundation model, but more directly aligned with this particular biomarker-prediction workflow.
How the prediction pipeline works
The released implementation uses a multi-stage workflow suited to images that are too large to process as a single ordinary image. First, tissue regions are identified and the whole-slide image is divided into smaller fixed-size patches. A patch encoder then converts each tissue patch into a numerical feature representation. These patch-level features are aggregated by a slide encoder into a representation of the entire slide. Finally, a classification head produces an EGFR mutation probability.
- Tissue segmentation: relevant tissue regions are separated from empty background.
- Patch creation: the tissue is divided into smaller image regions for processing.
- Patch feature extraction: a pretrained visual encoder converts each patch into features.
- Slide-level aggregation: information from many patches is combined into a whole-slide representation.
- EGFR classification: the model outputs a probability score associated with EGFR mutation status.
This design is important because a pathology slide may contain a very large number of local visual regions. Rather than treating the scan like a conventional photograph, the pipeline collects evidence from patches and combines it at slide level.
Supported inputs and outputs
The documented input is a lung adenocarcinoma whole-slide pathology image, generally based on H&E-stained tissue. The documented output is an EGFR mutation probability score. It does not produce text, images, audio, video, embeddings, or a conversational response as its primary output.
| Capability | Verified status |
|---|---|
| Whole-slide image input | Supported for the intended pathology workflow |
| Image understanding | Yes, specialized for digital pathology images |
| Text input or generation | Not documented as a supported function |
| Output | EGFR mutation probability score |
| Audio or video | Not supported |
| Tool calling or function calling | Not documented |
| Streaming | Not documented |
Because this is a classifier rather than a generative model, concepts such as context window, maximum output tokens, JSON mode, and text reasoning level are not applicable in the usual language-model sense. No context-length or output-token limit is provided in the supplied documentation.
Deployment and access requirements
EXAONE-Path-2.0-rev-EGFR is presented as a locally deployable research model rather than a broadly documented hosted API. The implementation requires an NVIDIA GPU with at least 12 GB of VRAM and a compatible CUDA-enabled PyTorch environment. This requirement makes it more suitable for a research workstation, laboratory server, or institutional computing environment than for casual browser-based use.
The model files are hosted in a gated Hugging Face repository. Users must accept the applicable access conditions before downloading the repository contents. The supplied research does not identify a conventional subscription plan, public per-image inference endpoint, token price, or hosted service price for this model. Therefore, the direct software access may be available for eligible research use under the repository terms, but infrastructure, storage, engineering, and validation costs still apply to an actual deployment.
Strengths and practical value
The model's main strength is task specificity. It is not trying to solve every pathology or language task; it is organized around EGFR prediction from lung adenocarcinoma whole-slide images. That focus can make it a useful research component for studies of computational biomarkers, precision oncology, and digital pathology workflows.
Another advantage is the use of slide-level processing. EGFR-related evidence may be distributed across multiple regions rather than appearing in a single obvious location. The patch-and-aggregation architecture is designed to preserve information from many tissue areas while producing one slide-level prediction.
The model also benefits from its position within EXAONE Path 2.0. According to the supplied research, EXAONE Path 2.0 is a pathology foundation model trained for slide-level pathology representation and provides the foundation for specialized downstream tasks. The current model therefore fits into a broader research strategy in which a pathology model can be adapted to specific biomarkers or clinical research questions.
These are practical strengths supported by the model's design and stated use case. They should not be interpreted as a guarantee of clinical accuracy across hospitals, scanners, populations, or staining protocols.
Limitations and safety considerations
EXAONE-Path-2.0-rev-EGFR should not be treated as a standalone clinical diagnostic device. A probability score is not the same as a confirmed molecular test, and the model should not replace laboratory testing, pathologist review, clinical judgment, or institutional validation.
Performance can vary when the deployment data differs from the data used to develop or evaluate the model. Relevant sources of variation include staining procedures, scanner characteristics, tissue preparation, slide quality, patient populations, and differences in clinical sampling. A research team should evaluate calibration, sensitivity, specificity, and failure cases on data that represents its intended setting before relying on predictions for downstream research decisions.
The model is also narrow in scope. It is not documented as a general cancer classifier, a complete pathology assistant, or a system for predicting every genomic alteration. Its output is centered on EGFR mutation probability in the stated lung adenocarcinoma workflow.
Operationally, the gated repository and local GPU requirement add friction compared with a hosted image-analysis service. Teams must manage the software environment, model files, image preprocessing, compute capacity, data governance, and validation process themselves. No official hosted API pricing, context window, maximum output size, or service-level availability is documented.
Reasoning, coding, and tool support
This model should not be evaluated using the usual language-model feature checklist. It does not provide a conversational reasoning mode, code generation, web search, function calling, or general-purpose tool use. Its computation is the fixed image-processing and classification pipeline described in the implementation.
Researchers may of course write code around the model to prepare slides, run inference, collect scores, or evaluate results. That surrounding software is separate from a model capability. The model itself is not documented as generating code or taking actions through external tools.
Speed, cost, and infrastructure trade-offs
The supplied specification rates the model's speed as 7 on the relevant internal scale, but this is an editorial or catalog evaluation rather than a provider-published benchmark. No verified inference-time benchmark, throughput figure, or per-slide cost is provided.
In practice, processing whole-slide images requires more computation than handling a small ordinary image because the slide may be divided into many patches. The local deployment model avoids dependence on a per-token language API, but it shifts costs toward GPU hardware, storage, preprocessing, engineering, and maintenance. Actual speed will depend on slide dimensions, the number of tissue patches, batch settings, GPU hardware, and the surrounding pipeline.
When to choose this model
EXAONE-Path-2.0-rev-EGFR is a reasonable choice when the project has all or most of the following characteristics:
- The research question concerns EGFR mutation prediction in lung adenocarcinoma.
- The available data consists of suitable digitized whole-slide H&E pathology images.
- The team can run a CUDA-enabled PyTorch environment on an NVIDIA GPU with at least 12 GB of VRAM.
- The project requires local or controlled processing rather than a public hosted API.
- The output will be treated as a research prediction requiring validation, not as a final clinical diagnosis.
A different option may be more appropriate for general pathology exploration, multiple biomarker tasks, or broader slide representation. In those situations, a pathology foundation model such as the related EXAONE Path 2.0 repository may offer a more general starting point, although the supplied materials do not establish that it provides the same EGFR classifier directly. A hosted medical-imaging platform may also be preferable when a team lacks GPU infrastructure or needs managed deployment, but no specific alternative service is documented here.
Availability and current positioning
LG AI Research released EXAONE-Path-2.0-rev-EGFR on July 9, 2025, as a research-oriented model associated with EXAONE Path 2.0. Its exact identifier distinguishes it from the general EXAONE-Path-2.0 foundation model and from the broader EXAONE-Path-EGFR listing.
The model's current position is therefore best understood as a gated, specialized downstream pathology release: more focused than a general foundation model, and less productized than a managed clinical or commercial diagnostic platform. Its value lies in providing a starting point for validated computational oncology research, not in offering a ready-made replacement for molecular testing or clinical decision-making.

