KI
·
4 MIN
GLM-5.2, Sakana Fugu, Claude Fable 5: three frontier models, three answers on control
Three frontier models, one real question for a B2B buyer: where does your data go, and how much of the stack do you control?

LOCATION
Rheinland-Pfalz
SERIES
Sovereign & Offline AI
AUTHOR
Aashwin Shrivastava
PUBLISHED
When I line these three up for a client in Rheinland-Pfalz, the benchmark deltas are rarely what decides it. All three are frontier-grade in mid-2026. The decision is where the data goes and how much of the stack the client controls, and GLM-5.2, Sakana Fugu, and Claude Fable 5 give three genuinely different answers. Only one of them can run inside your own building. That is the comparison that survives a procurement review, so it is the one I lead with.
01.
Three releases, in plain terms
Each of these shipped within a fortnight of the others in June 2026, and each is a different kind of thing. It is worth being precise, because two of them are widely described wrong.
GLM-5.2, from Z.ai in Beijing, released on 17 June. It is an open-weight Mixture-of-Experts model, roughly 750 billion parameters with about 40 billion active, a one-million-token context, and crucially an MIT licence with the weights published on HuggingFace.
Sakana Fugu, from Sakana AI in Tokyo, released on 22 June. It is not a conventional model, and the common framing gets this wrong: Fugu is a trained orchestrator that calls a pool of other models and synthesises their work. It is API-only, offered as an OpenAI-compatible endpoint.
Claude Fable 5, from Anthropic in the United States, released on 9 June. The second common mistake is to file Fable as a fast or creative tier. It is the flagship, Anthropic's most capable widely released model, safety-gated, cloud-only, at 10 and 50 dollars per million tokens. Its benchmark numbers are strong; for this piece they are also beside the point.
02.
The only comparison that survives procurement
Strip out the leaderboard and line them up on the dimensions a German B2B buyer is actually accountable for. The picture is clear, and it is not about which one is smartest.
Dimension | GLM-5.2 | Sakana Fugu | Claude Fable 5 |
|---|---|---|---|
Openness | Open weights, MIT | Closed, API-only | Closed, API-only |
Jurisdiction | China (self-host removes it) | Japan | United States |
Run on-prem | Yes, about 744 GB GPU | No | No |
API cost per M tokens | About 1.40 and 4.40 | Not disclosed | 10 and 50 |
Evidence quality | Vendor and secondary | Vendor self-report only | Vendor, strong track record |
Two honest caveats belong on this table. GLM-5.2's default path, the Z.ai API, sits under China's data laws, with their mandatory access provisions, which is exactly why the open weights matter: self-hosting in the EU neutralises that risk. And Fugu's claim of frontier parity is entirely self-reported, weakened by the fact that Fable and the restricted Mythos model are not even in its pool. I would not stake a decision on either vendor's benchmark.
03.
What I actually tell a client
I do not recommend one of these in the abstract. I match the posture to what the client can fund and what they are accountable for.
🔸 GLM-5.2 is the on-prem play. The open MIT weights are the whole point: you can run it in your own data centre, and the prompts never leave your network. The catch is the GPU footprint, roughly 744 gigabytes in FP8, so it fits the client who can fund the hardware and needs data to stay in the building. It is the cleanest sovereignty story of the three.
🔸 Claude Fable 5 is the managed-assurance play. You are renting capability from a US vendor, at the highest price here, with real safety gating and regional data-routing on the major clouds. For a team that wants a top model without owning the stack, and can live with a cloud dependency, it is the strongest managed option.
🔸 Sakana Fugu is the convenience play, with the weakest control story. One API that routes across a pool of models is clever, and Japan is softer geopolitics than China. But you cannot run it on-prem, you do not pick which model sees the data, and the evidence is thin. I would treat it as interesting, not as a default for regulated work.
This is the same lesson Claude Fable 5 taught the hard way when an earlier model was switched off in 72 hours: rented capability is revocable, and control is a property of the stack, not the score. It is why I keep pointing clients toward open-weight models they can actually own.
04.
The question worth designing for
If you take one thing from this, let it be the question, not the ranking. The models will trade places on the leaderboard again within a quarter; that part is noise. The durable question is the one a procurement officer should be asking on day one: which of these can you still run, audit, and afford when the vendor changes the terms? For most of the regulated clients I work with, that question answers itself, and it does not point at the highest benchmark. So before you pick the smartest model, what would it cost you to lose it?

