Llama-3.1-TFree-HAT

Llama-3.1-70B-TFree-HAT-SFT

by Aleph Alpha · Available open-weight research release

Aleph Alpha’s Llama-3.1-70B-TFree-HAT-SFT is an open-weight, approximately 69.3-billion-parameter English-German instruction model. It combines a Llama 3.1 70B backbone with a byte-level Hierarchical Autoregressive Transformer, offering a research-focused alternative to conventional subword-tokenized models. Its main trade-offs are self-hosting complexity, an unfinished inference implementation, no verified multimodal or native tool support, and a non-commercial research and education license.

Text Reasoning Coding
Llama-3.1-70B-TFree-HAT-SFT is Aleph Alpha Research’s approximately 69.3-billion-parameter supervised fine-tuned language model for English and German text generation. Its defining feature is the T-Free, or tokenizer-free, design: instead of converting text into a fixed vocabulary of subword tokens, the model processes byte-level text through hierarchical encoder and decoder components around a Llama-style backbone. It is available as an open-weight research release under the Open Aleph License for non-commercial research and educational use.
Outputs

What Llama-3.1-70B-TFree-HAT-SFT can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

6/10 Reasoning
4/10 Coding
5/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Llama-3.1-TFree-HAT
Model type General Purpose
Context window 98K tokens
Release date 2025-01-22
Status Available open-weight research release
Knowledge cutoff notes

Aleph Alpha's model card warns that the model has fixed historical training data and may not know newer information, but it does not publish a precise knowledge-cutoff date for this exact model.

Model notes

The canonical Hugging Face identifier is Aleph-Alpha/llama-3_1-70b-tfree-hat-sft. The model has approximately 69.3 billion parameters and uses BF16 safetensors. It is based on the Llama 3.1 70B backbone but replaces conventional tokenization with a Hierarchical Autoregressive Transformer using byte-level encoding and decoding. The architecture reports a 12,288-position encoder sequence and a 98,304-position decoder sequence; these are HAT sequence positions and should not be interpreted as a conventional subword-token context window. The model was supervised fine-tuned for instruction following in English and German. Aleph Alpha states that it was not optimized for code generation or mathematics. Weights are available under the Open Aleph License for non-commercial research and educational use. A vLLM inference implementation is available but remains under active development, and the model is not deployed through a Hugging Face inference provider. Editorial scores are comparative estimates rather than vendor specifications.

Model guide

Llama-3.1-70B-TFree-HAT-SFT: Aleph Alpha’s Tokenizer-Free Bilingual Model

Aleph Alpha’s Llama-3.1-70B-TFree-HAT-SFT is an open-weight English and German instruction-tuned model that replaces conventional subword tokenization with a Hierarchical Autoregressive Transformer architecture. Built around a Llama 3.1 70B backbone, it uses byte-level encoding and decoding to investigate better text compression and language adaptability.

What Llama-3.1-70B-TFree-HAT-SFT is

Llama-3.1-70B-TFree-HAT-SFT is an open-weight language model released by Aleph Alpha Research. The model is designed for instruction following and text generation in English and German, with a particular focus on research into tokenizer-free language modeling.

The name describes several important properties. “Llama-3.1-70B” identifies the Llama 3.1 70B lineage used for the main transformer backbone. “TFree” refers to the tokenizer-free approach, and “HAT” stands for Hierarchical Autoregressive Transformer. “SFT” means supervised fine-tuning, indicating that the base model was further trained on examples designed to improve instruction following.

The model contains approximately 69.3 billion parameters, including a reported 68.45-billion-parameter backbone. It is distributed in BF16 safetensors format through Hugging Face, where the canonical repository is Aleph-Alpha/llama-3_1-70b-tfree-hat-sft.

Why the tokenizer-free architecture matters

Most modern language models first divide text into subword tokens. A token may represent a complete word, part of a word, punctuation, or a combination of characters. This approach is efficient, but the resulting vocabulary and segmentation can favor some languages, writing patterns, and domains over others.

Llama-3.1-70B-TFree-HAT-SFT takes a different route. Its HAT architecture processes raw UTF-8 byte sequences within words, converts them into word-level representations, passes those representations through the Llama-style backbone, and then decodes the predictions back into bytes. In practical terms, the model replaces the conventional token embedding and output-token machinery with hierarchical encoder and decoder components.

This design is intended to reduce sequence fragmentation and improve text compression, particularly for languages or specialized text that may be represented inefficiently by a fixed subword vocabulary. Aleph Alpha’s model information reports stronger compression than the compared Llama 3.3 70B Instruct baseline across several evaluations. That is a research result rather than a guarantee that every deployment will be faster: the tokenizer-free implementation adds its own processing stages, and actual throughput depends on the software, hardware, and optimization level.

Architecture and context details

The model’s three main components are a byte-level encoder, a Llama-style backbone, and a byte-level decoder. The encoder turns byte sequences into representations that the backbone can process, while the decoder converts generated representations back into text.

The published architecture reports a 12,288-position encoder sequence and a 98,304-position decoder sequence. The latter is commonly recorded as the model’s context length, but it should not be interpreted exactly like a conventional subword-token context window. These are HAT sequence positions, not ordinary tokenizer-generated token counts. The practical amount of source text that fits will depend on how the model’s byte- and word-level processing represents that text.

A maximum output-token limit has not been verified in the supplied documentation. The model is a self-hosted research release rather than a managed service with a single provider-enforced output quota, so usable output length will also depend on the inference implementation and available hardware.

Training, languages, and instruction following

Llama-3.1-70B-TFree-HAT-SFT was pretrained and post-trained on curated English and German data. Its instruction-tuning mixture includes single-turn and multi-turn conversations, reasoning prompts, programming-related prompts, safety data, multilingual conversations, tabular reasoning, and mathematics.

English and German are the clearest supported languages for normal use. The model can accept and generate text through its byte-level interface, but it should not be treated as a broadly multilingual model solely because some multilingual material appeared in the training mixture.

The model card states that it was not optimized specifically for code generation or mathematics and was not extensively evaluated in those areas. It may still respond to programming or mathematical questions, but users seeking dependable specialist performance should validate results carefully or select a model designed and tested for those workloads.

Capabilities and published evaluations

The primary capability is text generation following natural-language instructions. Suitable tasks include drafting or transforming English and German text, answering questions from supplied context, summarizing, translating between its principal languages, conducting conversational exchanges, and experimenting with long-context or tokenizer-free architectures.

Published evaluations include MTBench, MMLU, MMLU Pro, GPQA, BBH, ARC, Winogrande, HellaSwag, TruthfulQA, German GSM8K, WMT16, AlpacaEval, and long-context tasks. Against Llama 3.3 70B Instruct, Aleph Alpha reports MTBench win rates of 0.633 in English and 0.639 in German. Results across broader knowledge and reasoning benchmarks are mixed, while some German instruction-following and reasoning evaluations are reported as competitive or favorable.

These figures are provider-reported research results, not independent production guarantees. Benchmark outcomes can vary with prompting, evaluation setup, decoding settings, and implementation. They should be used to understand the model’s research positioning rather than as a promise of performance for a particular application.

Supported modalities, tools, and structured output

This is a text-only language model. Verified capabilities cover text input and text output; there is no supplied evidence that this model natively accepts images, audio, or video, or that it generates those media types.

Native tool or function calling is not documented for this release. Similarly, a provider-defined JSON mode or structured-output guarantee has not been verified. An application may be able to constrain or parse generated text through its own serving layer, but that should not be confused with a built-in structured-output feature.

The model is not described as having web search, real-time information access, or built-in external actions. It has fixed historical training data and may produce outdated information. Applications that need current facts, retrieval, database access, or controlled business actions would need to add and validate those systems separately.

Deployment and availability

The weights are downloadable from Hugging Face for users who can operate a compatible local or private inference environment. Aleph Alpha provides a work-in-progress vLLM-based inference implementation adapted to the HAT architecture. The implementation remains under active development, and the model is not currently deployed through a Hugging Face inference provider.

This availability model makes Llama-3.1-70B-TFree-HAT-SFT different from a typical hosted chatbot or API model. Users should plan to manage model files, compatible software, hardware capacity, performance tuning, monitoring, and application safeguards themselves. The approximately 69.3-billion-parameter size also makes this a substantial deployment for individuals or small teams, particularly when using BF16 weights.

A tokenizer-free architecture should not automatically be equated with lower serving cost. Aleph Alpha reports improved compression, but the public implementation is not fully optimized, and end-to-end speed depends on the encoder, backbone, decoder, hardware, and serving stack. No public recurring price or usage-based inference price was verified for this model. The principal cost is therefore the infrastructure and engineering required to run it, unless an organization obtains access through a separately negotiated deployment.

Main strengths and limitations

  • Tokenizer-free research value: The model provides a substantial open-weight example of byte- and word-level processing around a Llama 3.1 70B backbone.
  • English and German focus: Its training and evaluations are especially relevant to bilingual English-German instruction following.
  • Open-weight access: Researchers can download the weights and investigate the architecture rather than relying only on a closed hosted endpoint.
  • Large-model capability: The 70B-class backbone gives the model a substantial general language-model foundation for text tasks.
  • Compression research: Aleph Alpha reports stronger compression than the compared Llama 3.3 70B Instruct baseline in its evaluations.
  • Deployment complexity: The model is not a one-click hosted service, and the available vLLM implementation is still a work in progress.
  • Unclear production throughput: Compression measurements do not establish real-world latency or throughput.
  • Limited modality and tool support: The verified model is text-only, with no documented native image, audio, video, web-search, or function-calling capability.
  • Specialist gaps: It was not specifically optimized or extensively evaluated for code generation and mathematics.
  • License restrictions: The Open Aleph License supports non-commercial research and educational use, so commercial deployment requires careful license review.
  • Normal language-model risks: The model can produce inaccurate, outdated, biased, repetitive, or unsafe content and requires application-level validation for high-stakes uses.

Reasoning, coding, speed, and cost trade-offs

Llama-3.1-70B-TFree-HAT-SFT includes reasoning-oriented material in its instruction-tuning mixture and has been assessed on several reasoning benchmarks. However, it is not presented as a dedicated reasoning model with a separate deliberation mode. Its reasoning quality should therefore be judged through task-specific testing rather than assumed from the parameter count.

Coding is a secondary capability rather than the model’s defining use case. Programming-related prompts were included in training, but Aleph Alpha explicitly states that the model was not optimized for code generation. Teams choosing it for software development should compare it with a code-specialized option and test practical tasks such as repository navigation, debugging, code transformation, and test generation.

Editorial assessment places its reasoning at a moderate level, coding below general-purpose strengths, speed at a moderate level, and cost favorably relative to paid hosted models when the weights can be run efficiently. These are comparative editorial estimates, not Aleph Alpha specifications. In practice, a large self-hosted model can be expensive to operate even when there is no per-token provider charge.

When to choose Llama-3.1-70B-TFree-HAT-SFT

Choose this model when the project specifically benefits from an open-weight, English-German, tokenizer-free architecture and the team is prepared to manage research deployment. It is a strong candidate for studying byte-level language modeling, evaluating compression behavior, building bilingual text-generation experiments, and testing private inference with a Llama 3.1 70B-class backbone.

It may also suit organizations that want to investigate how language models behave when conventional subword tokenization is removed, particularly for German or mixed English-German workloads. The model is more appropriate for controlled research and engineering environments than for casual users looking for an immediately available assistant.

Another option may be more appropriate when the priority is a polished hosted API, predictable latency, native multimodal input, web-grounded answers, built-in tools, guaranteed JSON output, highly specialized coding, or a clearly documented commercial license. A smaller model may also be preferable when serving cost, memory use, or response speed matters more than experimenting with a 70B-class architecture.

License and intended use

The model is released under the Open Aleph License for non-commercial research and educational use. Because it incorporates the Llama 3.1 backbone, users should also review the applicable Llama licensing and usage requirements, as well as export-control, privacy, and other legal obligations relevant to their deployment.

Its intended uses include research into tokenizer-free architectures, bilingual English-German generation, instruction-following experiments, text-compression studies, and further adaptation by technically capable users. Before using it in production or a high-stakes workflow, teams should verify licensing, evaluate the exact deployment stack, test both languages on representative data, and add safeguards for factuality, bias, privacy, and unsafe output.


Answers to Frequently Asked Questions

Is Llama-3.1-70B-TFree-HAT-SFT suitable for commercial use?
The model is released under the Open Aleph License for non-commercial research and educational use. Commercial deployment requires careful license review, including the applicable Llama 3.1 licensing and usage requirements. Organizations should also assess privacy, export-control, safety, and other legal obligations before using the model.
How can Llama-3.1-70B-TFree-HAT-SFT be deployed?
The weights are available in BF16 safetensors format from Hugging Face in the repository Aleph-Alpha/llama-3_1-70b-tfree-hat-sft. Users must operate a compatible local or private inference environment; the model is not currently available through a Hugging Face inference provider. Aleph Alpha provides a work-in-progress vLLM-based implementation, but deployment requires substantial hardware, software, and engineering resources.
Which languages and tasks does Llama-3.1-70B-TFree-HAT-SFT support?
The model is primarily designed for English and German. It can handle instruction following, drafting, summarization, translation between English and German, question answering from supplied context, conversation, and long-context experiments. It may respond to programming and mathematics prompts, but it was not specifically optimized or extensively evaluated for those areas.
What does tokenizer-free mean in Llama-3.1-70B-TFree-HAT-SFT?
Tokenizer-free means the model processes raw UTF-8 byte sequences rather than conventional subword tokens. Its hierarchical encoder converts bytes into word-level representations for the Llama-style backbone, and a decoder converts generated representations back into bytes. This approach is intended to improve text compression and reduce language-dependent tokenization effects.
What is Llama-3.1-70B-TFree-HAT-SFT?
Llama-3.1-70B-TFree-HAT-SFT is an open-weight, text-only language model released by Aleph Alpha Research for English and German instruction following and text generation. It uses a tokenizer-free Hierarchical Autoregressive Transformer architecture with an approximately 69.3-billion-parameter Llama 3.1 70B-class backbone.


Sources 4
Provider

About Aleph Alpha