What Llama-3.1-70B-TFree-HAT-SFT is
Llama-3.1-70B-TFree-HAT-SFT is an open-weight language model released by Aleph Alpha Research. The model is designed for instruction following and text generation in English and German, with a particular focus on research into tokenizer-free language modeling.
The name describes several important properties. “Llama-3.1-70B” identifies the Llama 3.1 70B lineage used for the main transformer backbone. “TFree” refers to the tokenizer-free approach, and “HAT” stands for Hierarchical Autoregressive Transformer. “SFT” means supervised fine-tuning, indicating that the base model was further trained on examples designed to improve instruction following.
The model contains approximately 69.3 billion parameters, including a reported 68.45-billion-parameter backbone. It is distributed in BF16 safetensors format through Hugging Face, where the canonical repository is Aleph-Alpha/llama-3_1-70b-tfree-hat-sft.
Why the tokenizer-free architecture matters
Most modern language models first divide text into subword tokens. A token may represent a complete word, part of a word, punctuation, or a combination of characters. This approach is efficient, but the resulting vocabulary and segmentation can favor some languages, writing patterns, and domains over others.
Llama-3.1-70B-TFree-HAT-SFT takes a different route. Its HAT architecture processes raw UTF-8 byte sequences within words, converts them into word-level representations, passes those representations through the Llama-style backbone, and then decodes the predictions back into bytes. In practical terms, the model replaces the conventional token embedding and output-token machinery with hierarchical encoder and decoder components.
This design is intended to reduce sequence fragmentation and improve text compression, particularly for languages or specialized text that may be represented inefficiently by a fixed subword vocabulary. Aleph Alpha’s model information reports stronger compression than the compared Llama 3.3 70B Instruct baseline across several evaluations. That is a research result rather than a guarantee that every deployment will be faster: the tokenizer-free implementation adds its own processing stages, and actual throughput depends on the software, hardware, and optimization level.
Architecture and context details
The model’s three main components are a byte-level encoder, a Llama-style backbone, and a byte-level decoder. The encoder turns byte sequences into representations that the backbone can process, while the decoder converts generated representations back into text.
The published architecture reports a 12,288-position encoder sequence and a 98,304-position decoder sequence. The latter is commonly recorded as the model’s context length, but it should not be interpreted exactly like a conventional subword-token context window. These are HAT sequence positions, not ordinary tokenizer-generated token counts. The practical amount of source text that fits will depend on how the model’s byte- and word-level processing represents that text.
A maximum output-token limit has not been verified in the supplied documentation. The model is a self-hosted research release rather than a managed service with a single provider-enforced output quota, so usable output length will also depend on the inference implementation and available hardware.
Training, languages, and instruction following
Llama-3.1-70B-TFree-HAT-SFT was pretrained and post-trained on curated English and German data. Its instruction-tuning mixture includes single-turn and multi-turn conversations, reasoning prompts, programming-related prompts, safety data, multilingual conversations, tabular reasoning, and mathematics.
English and German are the clearest supported languages for normal use. The model can accept and generate text through its byte-level interface, but it should not be treated as a broadly multilingual model solely because some multilingual material appeared in the training mixture.
The model card states that it was not optimized specifically for code generation or mathematics and was not extensively evaluated in those areas. It may still respond to programming or mathematical questions, but users seeking dependable specialist performance should validate results carefully or select a model designed and tested for those workloads.
Capabilities and published evaluations
The primary capability is text generation following natural-language instructions. Suitable tasks include drafting or transforming English and German text, answering questions from supplied context, summarizing, translating between its principal languages, conducting conversational exchanges, and experimenting with long-context or tokenizer-free architectures.
Published evaluations include MTBench, MMLU, MMLU Pro, GPQA, BBH, ARC, Winogrande, HellaSwag, TruthfulQA, German GSM8K, WMT16, AlpacaEval, and long-context tasks. Against Llama 3.3 70B Instruct, Aleph Alpha reports MTBench win rates of 0.633 in English and 0.639 in German. Results across broader knowledge and reasoning benchmarks are mixed, while some German instruction-following and reasoning evaluations are reported as competitive or favorable.
These figures are provider-reported research results, not independent production guarantees. Benchmark outcomes can vary with prompting, evaluation setup, decoding settings, and implementation. They should be used to understand the model’s research positioning rather than as a promise of performance for a particular application.
Supported modalities, tools, and structured output
This is a text-only language model. Verified capabilities cover text input and text output; there is no supplied evidence that this model natively accepts images, audio, or video, or that it generates those media types.
Native tool or function calling is not documented for this release. Similarly, a provider-defined JSON mode or structured-output guarantee has not been verified. An application may be able to constrain or parse generated text through its own serving layer, but that should not be confused with a built-in structured-output feature.
The model is not described as having web search, real-time information access, or built-in external actions. It has fixed historical training data and may produce outdated information. Applications that need current facts, retrieval, database access, or controlled business actions would need to add and validate those systems separately.
Deployment and availability
The weights are downloadable from Hugging Face for users who can operate a compatible local or private inference environment. Aleph Alpha provides a work-in-progress vLLM-based inference implementation adapted to the HAT architecture. The implementation remains under active development, and the model is not currently deployed through a Hugging Face inference provider.
This availability model makes Llama-3.1-70B-TFree-HAT-SFT different from a typical hosted chatbot or API model. Users should plan to manage model files, compatible software, hardware capacity, performance tuning, monitoring, and application safeguards themselves. The approximately 69.3-billion-parameter size also makes this a substantial deployment for individuals or small teams, particularly when using BF16 weights.
A tokenizer-free architecture should not automatically be equated with lower serving cost. Aleph Alpha reports improved compression, but the public implementation is not fully optimized, and end-to-end speed depends on the encoder, backbone, decoder, hardware, and serving stack. No public recurring price or usage-based inference price was verified for this model. The principal cost is therefore the infrastructure and engineering required to run it, unless an organization obtains access through a separately negotiated deployment.
Main strengths and limitations
- Tokenizer-free research value: The model provides a substantial open-weight example of byte- and word-level processing around a Llama 3.1 70B backbone.
- English and German focus: Its training and evaluations are especially relevant to bilingual English-German instruction following.
- Open-weight access: Researchers can download the weights and investigate the architecture rather than relying only on a closed hosted endpoint.
- Large-model capability: The 70B-class backbone gives the model a substantial general language-model foundation for text tasks.
- Compression research: Aleph Alpha reports stronger compression than the compared Llama 3.3 70B Instruct baseline in its evaluations.
- Deployment complexity: The model is not a one-click hosted service, and the available vLLM implementation is still a work in progress.
- Unclear production throughput: Compression measurements do not establish real-world latency or throughput.
- Limited modality and tool support: The verified model is text-only, with no documented native image, audio, video, web-search, or function-calling capability.
- Specialist gaps: It was not specifically optimized or extensively evaluated for code generation and mathematics.
- License restrictions: The Open Aleph License supports non-commercial research and educational use, so commercial deployment requires careful license review.
- Normal language-model risks: The model can produce inaccurate, outdated, biased, repetitive, or unsafe content and requires application-level validation for high-stakes uses.
Reasoning, coding, speed, and cost trade-offs
Llama-3.1-70B-TFree-HAT-SFT includes reasoning-oriented material in its instruction-tuning mixture and has been assessed on several reasoning benchmarks. However, it is not presented as a dedicated reasoning model with a separate deliberation mode. Its reasoning quality should therefore be judged through task-specific testing rather than assumed from the parameter count.
Coding is a secondary capability rather than the model’s defining use case. Programming-related prompts were included in training, but Aleph Alpha explicitly states that the model was not optimized for code generation. Teams choosing it for software development should compare it with a code-specialized option and test practical tasks such as repository navigation, debugging, code transformation, and test generation.
Editorial assessment places its reasoning at a moderate level, coding below general-purpose strengths, speed at a moderate level, and cost favorably relative to paid hosted models when the weights can be run efficiently. These are comparative editorial estimates, not Aleph Alpha specifications. In practice, a large self-hosted model can be expensive to operate even when there is no per-token provider charge.
When to choose Llama-3.1-70B-TFree-HAT-SFT
Choose this model when the project specifically benefits from an open-weight, English-German, tokenizer-free architecture and the team is prepared to manage research deployment. It is a strong candidate for studying byte-level language modeling, evaluating compression behavior, building bilingual text-generation experiments, and testing private inference with a Llama 3.1 70B-class backbone.
It may also suit organizations that want to investigate how language models behave when conventional subword tokenization is removed, particularly for German or mixed English-German workloads. The model is more appropriate for controlled research and engineering environments than for casual users looking for an immediately available assistant.
Another option may be more appropriate when the priority is a polished hosted API, predictable latency, native multimodal input, web-grounded answers, built-in tools, guaranteed JSON output, highly specialized coding, or a clearly documented commercial license. A smaller model may also be preferable when serving cost, memory use, or response speed matters more than experimenting with a 70B-class architecture.
License and intended use
The model is released under the Open Aleph License for non-commercial research and educational use. Because it incorporates the Llama 3.1 backbone, users should also review the applicable Llama licensing and usage requirements, as well as export-control, privacy, and other legal obligations relevant to their deployment.
Its intended uses include research into tokenizer-free architectures, bilingual English-German generation, instruction-following experiments, text-compression studies, and further adaptation by technically capable users. Before using it in production or a high-stakes workflow, teams should verify licensing, evaluate the exact deployment stack, test both languages on representative data, and add safeguards for factuality, bias, privacy, and unsafe output.

