Kolibri: A Sovereign Open-Weight Model

bastitx 562 points 311 comments October 03, 2026
aleph-alpha.com · View on Hacker News

tech report: https://aleph-alpha.com/downloads/tech-report.pdf additional paper: https://tej.as/blog/aleph-alpha-kolibri

Discussion Highlights (20 comments)

sajithdilshan

> It knows less from memory, Multi-turn tool calling is weaker, It’s not the best coding agent Then what does it good at? Sending faxes?

martianvoid

I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8

cyanydeez

Interesting they recommended high end software without considering quant 4 or 8 and still used A3B which should give good throughput on cheap hardware. If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.

cbarrick

I got distracted by that scroll-wheel UI component on the page. Neat!

pythonic_hell

The benchmarks are impressive given the problem space they are working in.

9dev

Aleph Alpha is just a sad joke by now. The talent isn't there anymore, they never managed to catch up to the other labs, failed to deliver on several projects, and by now are just a cash grab for the investors.

d2kx

German here. We are cheering for Mistral, which is making some good moves before the year is over, and Black Forest Labs for non-coding. But that's about it.

Lucasoato

> 4. It thinks in German This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.

woadwarrior01

> A bigger dense model beats it. Qwen3.8 27B ... How is a 27B dense model bigger than a 78B MoE?

x1watt

Was expecting that a "sovereign" AI model would at least use their own sovereign language (German) on the website as one of the options. Anyways, all the best and happy reunification day.

mistyvales

The Sega 32X game??

orifito

At least Germany is moving smarter than UK government...

veryfancy

Nice to see public goods in this space.

petesergeant

I wish nothing but luck for an EU model, but: > intellectual-property safety My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.

hypfer

The ignorant, hostile, negative, and, frankly, kinda racist comments here really are just a sad showing for the currently online crowd. But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs. We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).

spijdar

The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago. I get that doesn't invalidate the real "point" of the model, but...

tosh

i wonder if the custom tokenizer is better in practice, the examples look interesting though

niemandhier

I think at the moment the main thing a sovereign AI model needs to be good at is auditing the results of other models. Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these. So having a sovereign controlled model audit the first one would basically act like a “trust adapter”. If the second model is cheap and fast enough, there is a business model. You don’t even need to audit all the intermediate steps, just tool calls and end results.

CorezIoOfficial

Im surprised by how well this works. What is the difference from this and union alpha (other than the fact that it is open weights)?

kkm

Thank you Aleph Alpha team for making it open. We as many other’s were curious to try and benchmark it. On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days. No GPU. No setup. Just try it. tesseracted.com/kolibri-1-chat/ https://x.com/konarkmodi/status/2106373678589960260?s=46

Semantic search powered by Rivestack pgvector
8,416 stories · 78,695 chunks indexed