Protenix

Protenix

by ByteDance Seed · Active open-source biomolecular structure prediction project with multiple model variants, including Protenix-v1 and Protenix-v2

Protenix is ByteDance Seed’s open-source model family for predicting three-dimensional biomolecular complexes. It supports research involving proteins, nucleic acids, ligands, ions, and related molecular components, with v1, v2, base, mini, tiny, constraint, and other variants. The project offers pretrained checkpoints, inference tools, training documentation, and fine-tuning workflows, but no standard hosted pricing or language-model context window. It is best suited to computational biology researchers who can manage self-hosted scientific inference.

Reasoning Coding
Protenix is a ByteDance Seed AI-for-Science project focused on biomolecular structure prediction rather than text generation or conversational assistance. Its open-source implementation provides pretrained checkpoints, inference tools, training instructions, and multiple model variants for predicting molecular structures from structured inputs. The project is intended for computational biology, molecular design, and scientific research workflows that can support self-hosted inference.
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
1/10 Coding
5/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Protenix
Model type Other
Context window tokens
Maximum output tokens
Release date 2025-01-08
Status Active open-source biomolecular structure prediction project with multiple model variants, including Protenix-v1 and Protenix-v2
Knowledge cutoff notes

There is no single knowledge cutoff for the Protenix project because different checkpoints use different training-data cutoffs. Official documentation lists 2021-09-30 for several v1 and v0.5.0 variants and 2025-06-30 for protenix_base_20250630_v1.0.0. The cutoff for the umbrella Protenix identity is therefore not applicable.

Model notes

Protenix is used as the project and model-family identity rather than a single token-based model endpoint. The official project includes concrete variants such as protenix-v2, protenix_base_default_v1.0.0, protenix_base_20250630_v1.0.0, protenix_base_default_v0.5.0, constraint variants, mini variants, and tiny variants. The project predicts biomolecular structures from structured molecular inputs and produces structural outputs such as coordinates and confidence-related results. It is primarily intended for self-hosted or research deployment, so no canonical provider-hosted token pricing, context window, or maximum text-output-token limit applies to the umbrella identity. The initial public Protenix publication was posted in January 2025; later model-specific releases have separate release dates and training-data cutoffs.

Model guide

Protenix: ByteDance Seed’s Open-Source Model for Biomolecular Structure Prediction

Protenix is ByteDance Seed’s open-source model family for predicting three-dimensional structures of biomolecular complexes. It supports research involving proteins, nucleic acids, ligands, ions, and related molecular components, and includes Protenix-v1, Protenix-v2, base, mini, tiny, constraint, and other variants.

What is Protenix?

Protenix is an open-source biomolecular structure prediction project from ByteDance Seed. Its purpose is to estimate the three-dimensional arrangement of molecular components in complexes involving proteins, nucleic acids, ligands, ions, and related biological structures. In practical terms, it takes structured descriptions of molecules as input and produces structural results such as predicted coordinates and confidence-related outputs.

This makes Protenix fundamentally different from a general-purpose language model. It does not primarily generate text, answer questions, write software, or operate as a chat assistant. Its output is scientific structure data intended for analysis, visualization, computational biology, and molecular design workflows.

The project was publicly released in January 2025 and is maintained as a model family rather than a single fixed endpoint. The official repository includes pretrained checkpoints and documentation for installation, inference, training, and fine-tuning. Because the project is open source, users can inspect the implementation and run supported variants in their own research environment instead of relying on a single provider-hosted API.

Where Protenix fits in ByteDance Seed’s lineup

ByteDance Seed’s wider portfolio includes general-purpose, creative, multimodal, audio, robotics, and scientific AI research. Protenix belongs specifically to the AI-for-Science part of that portfolio. It represents ByteDance Seed’s work on biomolecular modeling and generative molecular design, rather than the company’s consumer-facing assistants or image and video products.

That positioning matters when evaluating access and capabilities. Protenix is not presented as a consumer subscription product with a chat interface, recurring plan, or standard token-based API. It is a research project distributed through its official GitHub repository and associated scientific publications. Users should therefore evaluate it as software and a collection of model checkpoints, not as a conventional hosted model service.

The project contains several supported variants. The catalog includes Protenix-v2 and Protenix-v1, along with base, mini, tiny, constraint, and ESM-related variants. These versions are intended to provide different trade-offs in supported features, model size, and computational requirements. The exact capabilities and training-data cutoff can vary by checkpoint, so “Protenix” should not be treated as having one universal specification sheet.

What Protenix can predict

Protenix is designed for biomolecular complexes rather than isolated text or images. Its stated scope includes:

  • Proteins and protein-containing complexes
  • Nucleic acids
  • Ligands
  • Ions
  • Other related molecular components represented in supported inputs

The model’s central task is structure prediction: estimating how the components of a molecular system are arranged in three dimensions. The resulting coordinates can be used as a starting point for computational analysis, structural inspection, or downstream molecular design. Confidence-related results are also produced to help researchers assess how reliable different parts of a prediction may be.

Protenix should not be described as a universal simulator or as a replacement for experimental validation. A predicted structure is a computational result whose usefulness depends on the input, the selected checkpoint, the biological system, and the limitations of the underlying model. Experimental work and additional computational analysis may still be required for high-stakes scientific conclusions.

Inputs, outputs, and modalities

The available documentation describes Protenix in terms of structured molecular inputs and structural outputs. It is not a text-in/text-out model, and the supplied research does not identify ordinary text, image, audio, or video generation as supported output modes.

For the umbrella Protenix identity, the available model record lists no general-purpose text input or text output, no image, audio, video, music, speech, or embedding output, and no multimodal consumer interface. The relevant input is molecular and structured: the user provides information describing the molecules or complex to be modeled, using formats and schemas supported by the repository and selected checkpoint.

Outputs are likewise scientific rather than conversational. They can include predicted molecular coordinates and confidence-related values. The precise output files, required fields, and preprocessing steps depend on the implementation and model variant selected. Researchers should follow the repository’s current inference documentation rather than assume that every Protenix checkpoint accepts exactly the same inputs.

Variants and checkpoint differences

One of Protenix’s practical strengths is that it is distributed as a family of checkpoints rather than only one large configuration. The official supported-model documentation lists Protenix-v2, Protenix-v1, base variants, mini and tiny variants, constraint variants, and ESM-related options. These variants allow researchers to select a configuration that better matches their hardware, use case, and desired feature coverage.

Larger or more feature-rich configurations may be preferable when the research problem benefits from the fullest supported modeling capability and the user has sufficient computational resources. Smaller mini or tiny configurations may be more suitable for experimentation, development, or environments where memory and processing time are constrained. The research supplied here does not provide a single benchmark table that would justify ranking every variant by accuracy or speed.

Training-data cutoffs also differ. Official documentation identifies September 30, 2021 for several v1 and v0.5.0 variants, while a checkpoint named protenix_base_20250630_v1.0.0 is associated with a June 30, 2025 cutoff. These dates apply to specific checkpoints, not to Protenix as a whole. A user should record the exact checkpoint name when documenting a scientific workflow.

Technical capabilities and limitations

Protenix’s main technical capability is structure prediction for biomolecular systems. The project also supports training and fine-tuning workflows according to its official documentation, making it more adaptable than a closed hosted endpoint for teams with the expertise and data needed to modify or evaluate a model locally.

There is no meaningful conventional context-window or maximum-output-token specification for the umbrella project. Protenix does not produce a sequence of language-model tokens as its primary result, and the supplied research does not identify a single provider-hosted context limit or maximum text-output limit. Resource requirements instead depend on the molecular input, model variant, implementation, and available hardware.

The project also has no verified public token pricing in the supplied material. Running it may involve infrastructure, storage, GPU, engineering, and research costs, but those costs are determined by the user’s deployment environment rather than by a published Protenix per-token rate. This can make the software attractive for organizations that need control over data and execution, while making it less convenient for users seeking a simple pay-as-you-go web service.

Tool use, function calling, web search, streaming responses, and conversational memory are not the relevant interfaces for Protenix. The model record lists these general assistant features as unavailable or not applicable to the project. Its documented workflow is closer to scientific software execution: prepare supported molecular inputs, select a checkpoint, run inference, inspect structural outputs, and evaluate confidence and biological plausibility.

Protenix’s main strengths

  • Specialized scientific focus: Protenix is built for biomolecular structure prediction rather than adapted from a general conversational system.
  • Broad molecular scope: Its stated target domain includes proteins, nucleic acids, ligands, ions, and related complex components.
  • Open-source availability: The implementation, inference instructions, training documentation, and supported-model information are available through the official repository.
  • Checkpoint choice: Base, mini, tiny, constraint, v1, v2, and related variants provide options for different research and hardware requirements.
  • Self-hosted research potential: Teams can run the project in their own environment, which may be important for sensitive molecular data, reproducibility, or customized experimentation.
  • Fine-tuning support: The supplied research identifies fine-tuning and training workflows, giving experienced teams a path beyond fixed pretrained inference.

These strengths are most relevant to researchers who can manage scientific software and computational infrastructure. They are less relevant to someone looking for an immediately accessible assistant or a simple browser-based prediction tool.

Limitations and trade-offs

The largest limitation is operational complexity. Protenix is a research codebase and model family, not a fully managed consumer application. Users may need to install dependencies, prepare molecular input files, select an appropriate checkpoint, provide suitable hardware, and interpret scientific outputs.

There is also no single Protenix specification that applies equally to every variant. Supported molecular features, computational requirements, training-data cutoff, and practical performance can differ between versions. Comparing results therefore requires naming the exact checkpoint and documenting the inference configuration.

Protenix is also not suitable for general-purpose AI tasks. It should not be selected for drafting documents, answering ordinary questions, generating software, browsing the web, creating images, or processing audio and video. A general language or multimodal model is more appropriate for those requirements.

Finally, predicted structures should not be treated as experimentally confirmed facts. Confidence outputs can help prioritize or inspect results, but they do not remove the need for domain expertise, validation, and awareness of the biological question being studied.

When to choose Protenix

Choose Protenix when the primary task is biomolecular structure prediction and your team can work with an open-source scientific model. It is a strong candidate for:

  • Academic or industrial computational biology research
  • Prediction of protein, nucleic-acid, ligand, or ion-containing complexes
  • Self-hosted molecular modeling where data-control requirements matter
  • Investigations that require access to multiple checkpoint sizes or configurations
  • Research programs involving training, fine-tuning, or reproducible local inference
  • Early-stage molecular design workflows that need predicted structures for downstream analysis

Another option may be more appropriate when you need a managed API, predictable hosted pricing, a graphical consumer interface, general-purpose reasoning, text generation, or tightly integrated laboratory and data platforms. A smaller Protenix variant may be preferable when speed and hardware cost are more important than using a larger configuration, while a larger or newer checkpoint may be worth evaluating when broader supported features or more recent training data are important.

Pricing and availability

Protenix is available as an open-source project through ByteDance’s official GitHub repository. The supplied research does not identify a recurring subscription, hosted API price, per-request fee, or token-based billing schedule for the Protenix project. Consequently, the direct model price is unverified rather than zero: users may still incur costs for GPUs, cloud compute, storage, maintenance, and technical support.

Availability should also be understood at the checkpoint level. The repository provides supported model information and instructions, but the practical ability to run a given variant depends on compatible software, hardware, and molecular inputs. Researchers should consult the current repository documentation before beginning a production or publication-critical workflow.

Bottom line

Protenix is ByteDance Seed’s open-source answer to a specialized scientific problem: predicting three-dimensional biomolecular structures from structured molecular inputs. Its value lies in domain focus, open implementation, multiple checkpoints, and the possibility of self-hosted research and fine-tuning. Its trade-offs are equally clear: it is not a conversational model, it has no standard token pricing or context window, and using it requires scientific and engineering expertise.

For researchers evaluating molecular structure prediction software, Protenix is best considered as a configurable research platform whose exact behavior depends on the selected checkpoint. For users seeking general AI assistance or a managed service, a different type of model will be a better fit.


Answers to Frequently Asked Questions

What is Protenix?
Protenix is an open-source biomolecular structure prediction project from ByteDance Seed. It predicts the three-dimensional arrangement of complexes containing proteins, nucleic acids, ligands, ions, and related molecular components.
Which Protenix model variants and checkpoints are available?
The Protenix model family includes Protenix-v2, Protenix-v1, base, mini, tiny, constraint, and ESM-related variants. Checkpoints differ in supported features, computational requirements, and training-data cutoffs, so users should record the exact checkpoint used.
Who should use Protenix?
Protenix is best suited to academic or industrial computational biology teams that need biomolecular structure prediction, self-hosted research, checkpoint flexibility, or training and fine-tuning capabilities. It is not intended for general-purpose chat, text generation, image creation, web browsing, or audio and video processing.
Is Protenix free to use, and does it have API pricing?
Protenix is available as open-source software through ByteDance’s official GitHub repository, and no recurring subscription, hosted API price, per-request fee, or token-based billing schedule is identified. Users may still incur costs for GPUs, cloud computing, storage, maintenance, and technical support.
What inputs and outputs does Protenix support?
Protenix uses structured molecular inputs defined by the selected checkpoint and repository-supported formats. Its outputs can include predicted molecular coordinates and confidence-related values rather than text, images, audio, or conversational responses.


Sources 5
Provider

About ByteDance Seed