GPT-6 Astra versus Claude Fable 5.1: an honest comparison
Most comparison tables for these two models do not survive scrutiny. This piece shows why, and names the small set of numbers that actually holds.
This translation was produced automatically using AI. The German version is the editorially reviewed original.
Since early September 2026 there have been two frontier models a German company can seriously weigh against one another: Claude Fable 5.1 from Anthropic, released on 1 September, and GPT-6 Astra from OpenAI, in limited preview since 3 September and generally available since 4 September. Much of what has circulated as a comparison table since then is unsound, not out of bad faith but because numbers from different measurement frames end up in the same column. This piece works the other way round: first discard the comparisons that do not hold, then name the few that do. What remains at the end is the column that actually decides a German architecture question, and it is not a benchmark row.
01. The table almost everyone is publishing is broken
The most common sentence in current coverage runs roughly like this: Anthropic's own benchmarks show Claude Fable 5.1 ahead of GPT-6 Astra. That sentence is false, and a calendar disproves it.
Anthropic published its benchmark table on 1 September 2026. The comparison column in it is labelled GPT-5.6 Sol. GPT-6 Astra went into limited preview on 3 September and into general availability on 4 September. Anthropic therefore did not merely fail to measure Astra, it could not have measured it: the model was not public at the time of publication.
Anyone reading the Terminal-Bench 4.0 row, where 55.8% for Fable 5.1 stands against 37.3%, and taking that 37.3% for Astra is comparing Anthropic's current model with the competitor's predecessor. That is not a rounding error, it is a different claim. And it is precisely the misreading now spreading fastest through aggregators and summaries.
The value of this piece therefore lies less in the numbers than in the sorting: which comparisons hold, which do not, and how to tell the difference.
02. What genuinely compares like for like
A small but clean core remains. Every row in the table below comes from the Anthropic model documentation, the Anthropic pricing page and the OpenAI developer documentation for gpt-6-astra, retrieved on 10 September 2026. These are each vendor's statements about its own product, in the same unit, with no measurement frame in between.
| Attribute | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|
| Released | 1 September 2026 | 3 September 2026 (preview), 4 September 2026 (general) |
| API id | claude-fable-5-1 | gpt-6-astra |
| Input / output per MTok | USD 10 / 50 | USD 10 / 50 |
| Cache read per MTok | USD 0.25 | USD 1.00 |
| Cache write per MTok | USD 12.50 (5 min) / 20 (1 hr) | USD 12.50 |
| Max output | 128K tokens | 128K tokens |
| Modality | text and image in, text out | text and image in, text out |
| Knowledge cutoff | June 2026 | 30 April 2026 |
One caveat belongs on the cache write row: Anthropic sells two hold durations, five minutes and one hour, while OpenAI publishes no equivalent tier. The two USD 12.50 figures sit close together but not perfectly parallel.
The economically interesting row is the cache read. A factor of four in Anthropic's favour sounds like a minor line item, but in agentic work it is not: there the same context, repository, system instruction and tool descriptions, is read again and again across hundreds of steps. Anthropic itself puts the resulting saving at roughly 25% against Fable 5 for typical workloads and up to roughly 45% for highly agentic work. That figure is a vendor model calculation, not a price cut: the headline input and output prices are unchanged, and a workload with few cache hits saves close to nothing.
Two further items belong in any calculation. Anthropic's Batch API halves both directions, to USD 5 input and USD 25 output per MTok. And pinning inference to the United States on Anthropic's side applies, according to Anthropic's data residency documentation, a 1.1x multiplier across input, output, cache writes and cache reads.
03. Terminal-Bench 4.0: the one figure with a real cross-check
One benchmark row in this comparison deserves particular confidence, for a reason that is rarely explained.
On 1 September 2026 Anthropic reports Claude Fable 5.1 on Terminal-Bench 4.0 at 55.8%. In the results OpenAI reports for GPT-6 Astra, relayed by DataCamp on 3 September 2026 and explicitly labelled there as vendor-reported, Astra stands at 57.7% and Fable 5.1 is carried at an unchanged 55.8%.
Both vendors therefore arrive independently at the same number for the other camp's model. That is the strongest single check available anywhere in this comparison. A vendor has little incentive to rate a competitor too highly, and when two parties with opposing interests report the same figure, it suggests the figure survives the setup rather than being a house measurement.
Two limits remain nonetheless. First, neither publication states which harness, which scaffold and which effort level were used. Second, the gap of 1.9 percentage points is small enough that it could come from exactly those factors. The defensible reading is therefore: on Terminal-Bench 4.0 the two models sit close together according to both vendors, with a slight lead for Astra on OpenAI's measurement.
04. What does not compare, and why
The larger part of the published numbers does not belong in a shared table. This is not formalism, it is the difference between a basis for a decision and a wall of figures.
- OSWorld 2.0. Anthropic reports two figures for Fable 5.1, 77.9% with partial credit and 41.7% under strict scoring. On the OpenAI side there is a single figure for Astra of 72.6%, with no stated scoring mode. Setting 72.6 against 77.9 is as unfounded as setting it against 41.7. Until the scoring mode matches, there is no comparison here, only a selection.
- Context window. Anthropic states 1,000,000 tokens, OpenAI 1,050,000 with a maximum input of 922,000 tokens. These are not capacity figures on the same scale, because a token means something different per tokenizer. Anthropic states itself that its current tokenizer produces roughly 30% more tokens for the same text than its own previous one. No cross-vendor tokens-per-word figure is publicly available, so the question of which model holds more text cannot be answered from open sources.
- AutomationBench. Anthropic publishes 31.4% for Fable 5.1. In its own announcement post OpenAI claims the top spot on the same benchmark but states no number there. There is simply nothing to compare.
- Rows where the Claude figure does not come from Anthropic. For ScreenSpot-Pro, FrontierMath Tier 4 v2, ExploitBench, ARC-AGI-3, GPQA Diamond and FrontierCode 1.1 only OpenAI-side figures exist. The Claude number in those rows is OpenAI's measurement of Claude, and in several cases a measurement of Claude Fable 5 or Claude Opus 5 rather than Fable 5.1. Attributing them to Anthropic cites the wrong source and, in part, the wrong model.
- Effort levels. Both vendors allow varying compute per request. At Anthropic even the default differs by surface:
highin Claude Code,mediumin claude.ai and Claude Cowork. Artificial Analysis measures at the "max" and "xhigh" settings. A benchmark number without a stated effort level is not comparable to one at a different setting. This is not a nicety, it is the most likely reason the same leaderboard shows three different results on three different websites.
One special case deserves a warning: the 99.9% circulating for Astra on ARC-AGI-3 carries an "adapter harness" qualifier in the source and sits beside 7.8% for the previous model. A swing of roughly 92 points across a harness change describes the harness first and the model second. That figure belongs in no headline.
05. The third-party leaderboards contradict each other outright
Anyone distrusting the vendor figures turns to third-party leaderboards. In this case that does not help, because the leaderboards disagree with one another, on one and the same index.
Artificial Analysis, in its own article of 3 September 2026, gives for the Intelligence Index: GPT-6 Astra 61, GPT-5.6 Sol 61, Claude Fable 5.1 66, the last at maximum effort with a fallback configuration that is not the API default. BenchLM reports the same index differently in September 2026: GPT-5.6 Sol leads at 58.9%, with Fable 5.1 at 53.7%. llm-stats, retrieved on 10 September 2026, reports Fable 5.1 at 56.8, Astra at 54.7 and Claude Opus 5 at 54.1.
Three sources, one index, three orderings. They cannot all be current. Plausible causes are different snapshot dates, different effort settings, and points being mixed with percentages. The practical consequence is simple: an Intelligence Index figure without the site, the date and the effort setting is not information. Artificial Analysis itself, incidentally, flags mixed findings, among them a drop of roughly 80 Elo points on GDPval-AA v2.
That leaves the arena. There is nothing to be had there either: neither Claude Fable 5.1 nor GPT-6 Astra holds a ranked LMArena position in the September 2026 snapshots reachable here, because arena Elo needs vote volume and lags frontier releases by weeks. Even if the figures existed, they would be the wrong instrument for this decision. Blind pairwise preference measures perceived answer quality on self-selected prompts, reacts strongly to formatting and verbosity, and says little about agentic work over long horizons. For a coding agent or a knowledge-work platform, that is not the yardstick a purchase should follow.
06. FrontierCode 1.1: the number that cuts against the vendor reporting it
There is one row in this comparison that carries more weight than the others, for a methodological reason. On FrontierCode 1.1, GPT-6 Astra stands according to OpenAI's own reported results at 53.3%, behind Claude Fable 5 at 53.5%.
Two clarifications are obligatory here. The comparison figure is for Claude Fable 5, not Fable 5.1, and Anthropic has published no figure of its own for Fable 5.1 on this benchmark. The gap is 0.2 percentage points, well inside any plausible measurement spread.
Even so, it is the most trustworthy kind of number a product announcement can contain. A vendor publishing a row in which its new flagship trails a rival model has no marketing incentive to do so. Anyone assessing how load-bearing a benchmark table is should look first for whether rows like this appear in it at all. A table in which the publishing vendor wins every single row is not a measurement, it is a selection.
07. The direction reverses: EU data residency and zero data retention
Up to this point Anthropic leads on price and cache economics and is roughly level on the benchmarks. On data governance the picture reverses completely, and for a German company that is the column that changes an architecture.
| Attribute | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|
| EU in-region inference | No, only us and global | Yes, via eu.api.openai.com for EEA and Switzerland |
| Data storage region | United States only | Europe selectable per Project |
| Zero data retention | Not available, Covered Model with mandatory 30-day retention | Documented for the main inference endpoints, approval-gated |
| Trains on customer API data | No, not without express permission | No, not unless opted in |
On the Anthropic side this is in the vendor's own documentation. The inference_geo parameter accepts exactly two values, global and us, and the only selectable workspace storage geo is us, fixed once the workspace is created. The retention side is more explicit still: Anthropic designates Fable 5.1 and Mythos 5.1 (as well as Fable 5 and Mythos 5) as Covered Models that require 30-day data retention and are not available under zero data retention unless expressly authorised by Anthropic. An organisation on ZDR has to switch 30-day retention on for a specific workspace, or the request is rejected.
On the OpenAI side, the developer documentation on data handling states that Europe (EEA and Switzerland) supports both regional storage and regional processing via eu.api.openai.com, configured per Project, with non-US regions requiring approval for abuse-monitoring controls and a modified retention amendment. The same page lists the ZDR-eligible endpoints, among them /v1/chat/completions and /v1/responses, and explicitly excludes Assistants, Threads, Vector Stores, fine-tuning and Batches.
Precision matters more here than a smooth statement: no source found states positively that GPT-6 Astra is ZDR-eligible. What is documented is that those endpoints support ZDR and that no model-specific exclusion is published for gpt-6-astra. The absence of a restriction, however, is not a permission. Anyone building an architecture on it should have it confirmed contractually rather than inferring it from a documentation gap. At Anthropic the position is unambiguous in the other direction, because the exclusion is written down.
One widely repeated line also needs correcting: "Claude has EU data residency because it runs on AWS Frankfurt." That is true precisely when you obtain Claude through a partner cloud such as Amazon Bedrock or Google Cloud, where the cloud provider sets the region and is the data processor. It does not hold for Anthropic's own API. For guaranteed regional routing Anthropic points to the partner clouds' regional endpoints itself, at a 10% premium over the global endpoints. For a GDPR review that distinction is the whole point, because it determines who the data processing agreement is signed with.
08. Article 50: a difference that can only be half evidenced
Since 2 August 2026 the transparency obligations of Article 50 of the AI Act have applied, requiring machine-readable marking of synthetic content. Anthropic states that Fable 5.1 and Mythos 5.1 carry an invisible text watermark from day one, and applies that marking worldwide rather than only in the EU. As far as the available sources reach, that is a genuine difference.
The sentence needs both halves, though. No source was found stating whether GPT-6 Astra's text output is marked. OpenAI's documented provenance approach, C2PA Content Credentials plus SynthID, covers images and audio; the reachable pages say nothing about text. That gap does not mean OpenAI fails to mark text. We therefore do not claim it. OpenAI separately publishes customer guidance on the AI Act.
And Anthropic's watermark carries less than the term suggests. Anthropic writes in its own help pages that a detected mark indicates content may have been processed by Claude, is not fully conclusive, and does not on its own confirm the provenance of the content. The corresponding detection API is moreover in private preview for selected organisations. As one building block of a compliance argument the watermark is usable; as evidence it is not.
09. What this means for a German decision
If you take one rule from this comparison, take this one: with two frontier models sitting close together on the defensible figures and priced identically at the headline, a benchmark row rarely decides anything. The gap on Terminal-Bench 4.0 is 1.9 percentage points with no stated harness. On FrontierCode 1.1 it is 0.2 points the other way, against an older Claude model. Gaps like that shift with the next release, and they change no architecture.
What does change an architecture sits in the governance column. An organisation bound to zero data retention cannot use Claude Fable 5.1 through Anthropic's own API without express authorisation, however good the model is. An organisation requiring inference inside the EU finds that route at Anthropic only through a partner cloud, with the corresponding premium and a different data processor in the contract. Conversely, Anthropic's cache-read advantage is real and grows with the share of agentic work in which the same context is read hundreds of times.
In practice: settle the obligations first, the costs second and the benchmarks last. In the reverse order you build a system that measures well and fails the data protection conversation. And for every number someone shows you, check three things: who reported it, when, and at which effort level.
What the release itself changed for a German company is covered in our piece on Claude Fable 5.1. How the model behaves in live project work is covered in capabilities in practice. And if the answer to the governance question is that the data must not leave the building at all, the route runs through the on-premise versus cloud trade-off rather than through a leaderboard.

