Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

Kolibri: Aleph Alpha's 78B Open Model, Benchmarked

At a glanceQuick answers
What is Kolibri?
Aleph Alpha’s open-weight mixture-of-experts language model for German and English, with 78.1B total and 3.46B active parameters, released October 3, 2026.
Can I use it commercially?
Yes. The full weights are on Hugging Face under the Apache 2.0 license.
Is it the best small open model?
Not across the board. On Aleph Alpha’s own table it leads math and banking-agent rows, and trails Qwen’s 35B-A3B models on agentic coding and function calling.
Data illustration on off-white paper: a teal hummingbird hovering over a grid of small grey blocks, three of them lit in amber and joined by a dotted line to its beak, with the large numbers 78B total and 3.46B active and a tag reading Apache 2.0
Fig 0A small bird with a big body: Kolibri holds 78B parameters and uses 3.46B per token. Made by CellCog's image agent, running GPT Image 2.5.

Kolibri is Aleph Alpha’s new open-weight model: a German-English mixture-of-experts with 78.1 billion parameters, 3.46 billion of them active per token, a context window validated to about one million tokens, and an Apache 2.0 license. Aleph Alpha released it on October 3, 2026, the Day of German Reunification. On its own benchmark table it leads the mixture-of-experts models it compares on math and a banking-agent test, and trails both Qwen 35B-A3B models on agentic coding. This page reads it from Aleph Alpha’s blog, model card and newsroom as of October 4, 2026.

On this page · 6 sectionsOpen
  1. What Kolibri is
  2. Not a lab from nowhere
  3. How it was trained
  4. Where it wins and where it doesn’t
  5. What we are watching
  6. Sources
Key points6 · 7 min full read
  1. Kolibri is Aleph Alpha’s open-weight German-English model: a mixture-of-experts with 78.1B total parameters and 3.46B active per token, released October 3, 2026 under Apache 2.0.
  2. Its native context is 262,144 tokens; Aleph Alpha validated it to 1,048,576 tokens and recommends staying at or under 262,144 for serving.
  3. The FP8 weights take about 78 GB, so it runs on one H200, B200 or B300, or on two H100 or A100 80 GB cards.
  4. Aleph Alpha trained it on 768 NVIDIA B200 GPUs: 20T tokens of pre-training in 21 days, then 3.44T of mid-training and 201B of long-context training.
  5. On Aleph Alpha’s own harness it leads the mixture-of-experts models it compares on AIME, GPQA and Tau3 banking, and trails both Qwen 35B-A3B models on SWE-Bench Verified, TerminalBench 2.1 and BFCL v4.
  6. Aleph Alpha is not a new lab: Kolibri follows Kolibri Origin, and the company signed a definitive agreement to combine with Cohere on September 16, 2026, still pending regulatory approval.

§ 01What Kolibri is

Question Answer Source
Who makes it Aleph Alpha; Aleph Alpha Research GmbH is the developer Model card
Released October 3, 2026; Hugging Face repository created October 2 Blog, Hugging Face API
Size 78.1B total parameters, 3.46B active per token Model card
Architecture 50-layer mixture-of-experts, 384 experts per layer (1 shared, 6 routed), sliding-window and global attention at 4:1 Model card
Context 262,144 tokens native, validated to 1,048,576 Model card
Languages German and English, with a tokenizer built for German word structure Model card
License Apache 2.0, full weights on Hugging Face Model card
Hardware About 78 GB of FP8 weights: one H200, B200 or B300, or two H100 or A100 80 GB cards Model card
Knowledge cutoff June 18, 2026 Model card
Reasoning and tools Explicit reasoning mode with adjustable effort, tool calling Model card
Adoption 388 likes and 1,135 downloads on Hugging Face Hugging Face API, Oct 4
Table 1Kolibri 1, read on October 4, 2026

Aleph Alpha’s launch post opens: “On the Day of German Reunification, we are releasing our new model: Kolibri.” Kolibri is German for hummingbird, and the company’s own announcement on X read “Small bird, fast wings, Kolibri is here.” The model card carries the full specification, and a tech report has the training detail.

The trade-off sits in the hardware row. Only 3.46B parameters do work on each token, which keeps serving cheap, but all 78B have to be in memory. That is why the floor is a single 141 GB or larger card, or two 80 GB cards, rather than a laptop.

§ 02Not a lab from nowhere

The post that carried Kolibri across X on October 4, from Charles Maddock, began: “Random German lab appears from nowhere and drops SOTA open-weight model”. Replies pushed back within hours, and the record backs them.

  • Kolibri has a predecessor. Aleph Alpha first built Kolibri Origin, a 30B model with 3B active and a 65k context window, to validate its training pipeline, then ran Kolibri through the same pipeline.
  • The Cohere deal is signed, not closed. On September 16, 2026, Cohere and Aleph Alpha signed a definitive business combination agreement, following a plan announced in April. The combined company will operate globally as Cohere, with headquarters in Berlin and Toronto. In the companies’ words: “The transaction remains subject to final regulatory approvals.”
  • It is aimed at a specific buyer. Aleph Alpha says Kolibri is built for “sovereign mission-critical work in regulated areas including public administration, industrials and aerospace”, run on a customer’s own hardware without sending data to outside inference services.

§ 03How it was trained

Stage Tokens Sequence length Compute
Pre-training 20T 16,384 21 days, 392k GPU hours
Mid-training 3.44T 65,536 5 days, 90k GPU hours
Long-context 201B 262,144 13 hours, 10k GPU hours
Table 2Kolibri’s three training stages, from the model card

All three stages ran on 768 NVIDIA B200 GPUs, nearly 24T tokens in total, which Aleph Alpha says is roughly three times what Kolibri Origin consumed. The pre-training mix was about 62.5% English, 23.9% German and 13.6% code, so German alone came to about 4.3T tokens. The card puts total training compute at 6.4e23 FLOPs and notes that Aleph Alpha has signed the EU general-purpose AI Code of Practice.

§ 04Where it wins and where it doesn’t

Aleph Alpha published one large comparison table. The card’s note says “All models use the same evaluation setup”, run on its open eval-framework and the Harbor harness for SWE-Bench and TerminalBench, with Kolibri at reasoning effort high. These are the lab’s own numbers; we have not seen independent runs yet. Here are nine rows against the four mixture-of-experts models people are comparing it with.

Benchmark Kolibri (3.46B active) Qwen3.6 35B-A3B Qwen3.5 35B-A3B Nemotron 3 Super 120B-A12B Mistral Small 4 119B-A6B
AIME 2025 96.9 84.6 88.1 91.7 79.8
AIME 2026 96.0 91.0 92.1 90.4 83.1
GPQA Diamond 84.3 83.4 83.8 78.0 74.7
LiveCodeBench v6 85.9 82.5 77.8 82.0 71.2
Tau3-Bench (Banking) 38.1 10.6 11.3 15.5 5.7
Tau2-Bench (Telecom) 94.7 99.1 97.7 68.1 41.5
BFCL v4 (overall) 61.4 67.2 70.5 61.0 58.0
SWE-Bench Verified 66.4 73.8 71.6 60.2 60.8
TerminalBench 2.1 27.7 – 39.7 39.7 21.0
Scroll to compare all columns
Table 3Selected rows from Aleph Alpha’s Kolibri table (higher is better; a dash means no score reported)
SWE-Bench Verified, Aleph Alpha's harness: Kolibri trails both Qwen 35B-A3B modelsBar chart of SWE-Bench Verified scores: Qwen3.6 35B-A3B 73.8, Qwen3.5 35B-A3B 71.6, Kolibri 66.4 highlighted, Mistral Small 4 60.8, Nemotron 3 Super 60.2Qwen3.6 35B-A3B73.8Qwen3.5 35B-A3B71.6Kolibri66.4Mistral Small 460.8Nemotron 3 Super60.2SWE-Bench Verified, Aleph Alpha's harness: Kolibri trails both Qwen 35B-A3B modelsBar chart of SWE-Bench Verified scores: Qwen3.6 35B-A3B 73.8, Qwen3.5 35B-A3B 71.6, Kolibri 66.4 highlighted, Mistral Small 4 60.8, Nemotron 3 Super 60.2Qwen3.6 35B-A3B73.8Qwen3.5 35B-A3B71.6Kolibri66.4Mistral Small 460.8Nemotron 3 Super60.2
Fig 1SWE-Bench Verified, Aleph Alpha's harness: Kolibri trails both Qwen 35B-A3B models

Where it wins. Kolibri tops every mixture-of-experts model in the table on AIME 2025, AIME 2026 and GPQA Diamond, including Nemotron 3 Super, which runs 12B parameters per token. The widest gap is Tau3-Bench banking, a multi-step customer-service agent test: 38.1 against 16.0 for the next mixture-of-experts model in the full table. That fits Aleph Alpha’s own claim that “Kolibri sits on the Pareto frontier for quality versus serving cost, for both English and German.”

Where it doesn’t. Agentic coding and function calling go the other way. Both Qwen 35B-A3B models beat it on SWE-Bench Verified, Qwen3.5 leads it on BFCL v4 by 9 points, and on TerminalBench 2.1 it scores 27.7 against 39.7 for Qwen3.5 and Nemotron 3 Super. Those are the rows the “loses to 35b a3b” replies on X pointed at. The full table also includes Qwen3.8 27B, a dense model that uses 27B parameters per token; it beats Kolibri on eight of the nine rows above, all but Tau2 telecom, at about eight times the compute per token.

So Charles Maddock’s line that it “beats Qwen, Mistral, Nemotron models with similar active parameter count” holds for math, knowledge and the banking agent, and not for coding agents. If your workload is German-language reasoning on your own hardware, it is the strongest small option in this table. If it is an agent that writes and runs code, the Qwen 35B-A3B models still lead.

§ 05What we are watching

  • Independent evaluations. Every number above is Aleph Alpha’s. We will add outside results when they publish.
  • Hosted access. Kolibri was not listed on OpenRouter when we checked on October 4; today you run it yourself with vLLM from the Hugging Face weights.
  • The Cohere close. Once regulators approve the combination, Kolibri becomes a Cohere model, and its roadmap may change with it.
  • The tech report. We will fold in any training or evaluation detail that changes how to read the table.

§ 06Sources

Frequently asked5 questions

Q1Who made Kolibri?

Aleph Alpha, the Heidelberg AI company; the model card names Aleph Alpha Research GmbH as the developer. It follows Kolibri Origin, a 30B model with 3B active that the same training pipeline produced first.

Q2What hardware does Kolibri need?

About 78 GB for the FP8 weights. The model card lists a minimum of two A100 80 GB, two H100 SXM5, one H200, one B200 or one B300. Only 3.46B parameters are active per token, but the whole model has to sit in memory.

Q3Does Kolibri really support 1M tokens?

Aleph Alpha says it validated quality and serving efficiency up to 1,048,576 tokens. The native training length is 262,144 tokens, and the model card recommends staying at or below that for latency-sensitive work and complex tasks.

Q4Did Aleph Alpha merge with Cohere?

Not yet. The two companies signed a definitive business combination agreement on September 16, 2026, after announcing the plan in April. The combined company will operate as Cohere, and the deal remains subject to regulatory approvals.

Q5Does CellCog run on Kolibri?

No. Every CellCog tier runs on Anthropic’s Claude Opus 5.5, and each AI employee works in its own secure VM with its own file system, browser identity and logins. We track open models like Kolibri because many of our readers build with them.

Published 04 October 2026 All Choosing a platform →