GPT-6 Astra Statistics (2026): Benchmarks, Cost Per Task, Pricing and Safety Data
GPT-6 Astra ties Claude Fable 5.1 at 53 on the Artificial Analysis Intelligence Index v4.3 and costs $3.26 per task against Fable's $7.63. Both list at $10/$50 per million tokens, but Astra uses about 65% fewer output tokens per task. It leads on terminal agents, workflow automation and computer use. Fable 5.1 still wins long knowledge-work projects and Humanity's Last Exam with tools. GPT-5.6 Sol gives the cheapest index point, but only while its promo price lasts, which runs at least through November 21, 2026. Benchmark rankings shift with index versions and harnesses (ARC-AGI-3 ranges from 62.7% to 99.9%), so always cite the version and date.
GPT-6 Astra is OpenAI’s flagship model, released September 3, 2026 at $10 per million input tokens and $50 per million output tokens. On the Artificial Analysis Intelligence Index v4.3 it ties Claude Fable 5.1 at 53, while costing $3.26 per task against $7.63 and using 27K output tokens against 78K.
GPT-6 Astra in eight numbers
- 53 on the Artificial Analysis Intelligence Index v4.3, tied with Claude Fable 5.1 and 6 points above GPT-5.6 Sol (September 7, 2026).
- $3.26 per Intelligence Index task at max effort, 57% less than Fable 5.1’s $7.63.
- 27K output tokens per Index task, about a third of Fable 5.1’s 78K.
- 51% hallucination rate on AA-Omniscience at max effort, down from Sol’s 92%.
- 62.7% on ARC-AGI-3 with ARC Prize’s standard harness. The 99.9% headline needs OpenAI’s own adapter.
- 169 on the Epoch Capabilities Index, a record. The previous best was 163.
- 1,050,000 tokens of context, with 128,000 tokens of max output.
- 5-45 estimated Astra messages per five hours on ChatGPT Plus, in Work and Codex only.
How much does GPT-6 Astra cost per task compared with Fable 5.1 and Sol?
Per task, GPT-6 Astra costs well under half of what Claude Fable 5.1 costs for the same score. Artificial Analysis measured $3.26 per Intelligence Index task for Astra at max effort and $7.63 for Fable 5.1, even though both list at $10 input and $50 output per million tokens.
The list price is identical. The bill is not, because token use drives it. We pulled Artificial Analysis’s published figures into one table and added two calculations of our own: cost per index point and a blended price per million tokens.

| Metric (max effort) | GPT-6 Astra | Claude Fable 5.1 | GPT-5.6 Sol |
| AA Intelligence Index v4.3 | 53 | 53 | 47 |
| Cost per Index task | $3.26 | $7.63 | ≈$2.04* |
| Cost per Index point (our calc) | 6.2¢ | 14.4¢ | ≈4.3¢* |
| Output tokens per Index task | 27K | 78K | Not published |
| Turns per GDPval-AA v2 task | 24 | 60 | 45 |
| Coding Agent Index (cost per task) | 62 ($7.09) | 62 (≈$11.80*) | 55 (≈$6.20*) |
| Blended list price per 1M tokens, 7:2:1 mix (our calc) | $7.70 | $7.17 | $3.08 |
Sources: Artificial Analysis, September 7 and 9, 2026; OpenAI and Anthropic price pages checked September 10, 2026. *Derived from Artificial Analysis’s rounded percentages (“~60%”, “~40%”, “~15%”), so treat them as approximate. The 7:2:1 mix is cache hits to fresh input to output, the ratio Artificial Analysis uses for blended prices.
Start with the bottom row. Per token, Fable 5.1 is about 7% cheaper than Astra. Anthropic charges $0.25 per million cached input tokens on Fable 5.1, and OpenAI charges $1.00 on Astra, so cache-heavy traffic tilts toward Fable. Now look at the cost row. Astra finishes each task with roughly 65% fewer output tokens, and at $50 per million output tokens that saving outweighs the cache gap.
Of these three models, GPT-5.6 Sol buys an index point most cheaply, at roughly 4.3 cents. It just tops out six points lower. OpenAI’s smaller models go lower still on v4.3: GPT-5.6 Terra scores 42 for $1.40 per task and GPT-5.6 Luna scores 38 for $0.18. Every Astra effort level sits on the Artificial Analysis cost frontier, from $0.82 per task at low effort to $3.26 at max, so there’s room to trade score for price inside Astra too.
The Sol comparison has an expiry date
Sol’s $4/$20 price is promotional. OpenAI’s model page says it runs at least through November 21, 2026, and describes it as a 20% cut on input and 33% on output. Artificial Analysis calls Astra 2.5x Sol’s current price. If Sol goes back to $5/$30, our blended math puts Astra at about 1.8x Sol instead. Any “Astra vs Sol” cost chart made in September has a shelf life. Our GPT-5.6 Sol statistics page tracks that price history.
Why do GPT-6 Astra benchmark rankings disagree?
The benchmarks changed under it, and the software wrapped around the model matters a lot. Between September 3 and September 7, Astra went from 5 points behind Claude Fable 5.1 on the Artificial Analysis Intelligence Index to tied for first as the index itself changed twice.

- September 3, Index v4.1: Astra scored 61, level with GPT-5.6 Sol and 5 points behind Fable 5.1. (Sol had scored 59 on v4.1 at its July launch; OpenAI’s September table lists it at 60.9 on v4.1.1.) On the Coding Agent Index it scored 67 against Fable 5.1’s 70. Coding Agent Index setups also change, so don’t set that 67 against the 80 Sol scored in July.
- September 4, Index v4.2: Fable 5.1 led and Astra was second, 4 points above Sol. Artificial Analysis added AA-Briefcase and GDP.pdf and removed GPQA Diamond, which it considered saturated.
- September 7, Index v4.3: Astra and Fable 5.1 tied at 53, with Sol at 47. Terminal-Bench moved to v4.0 and AutomationBench-AA replaced τ³-Banking. Held-out test sets now carry 45% of the weight.
Scores from different versions don’t sit on the same scale, so a 53 on v4.3 is not a drop from 61 on v4.1. Artificial Analysis says it had been building these changes for months and pushed interim versions because the frontier moved fast. The Decoder reported the v4.2 release as a likely response to criticism that the index was missing Astra’s progress. Both accounts can be true. If you quote an Astra ranking, name the index version and the date, or the number doesn’t mean much.
OpenAI’s own launch table makes the point from the other side. It lists Artificial Analysis Intelligence Index v4.1.1, where Fable 5.1 scores 65.7 and Astra 61.2. That’s a vendor publishing a comparison its rival wins.
Other aggregate scores put Astra further ahead. Epoch AI gave it a record 169 on the Epoch Capabilities Index, 6 points above the previous best of 163, while noting the jump still sits within the uncertainty of its reasoning-era trend.
ARC-AGI-3: same model, 37 points apart
ARC Prize ran Astra twice on the ARC-AGI-3 Semi-Private set. With ARC Prize’s standard harness, where the model chooses which notes to carry forward, it scored 62.7% at max effort for $26,098. With OpenAI’s Provider Adapter harness, which keeps reasoning state between requests and compacts long runs, it scored 99.9% at high effort for $18,817.
That’s a 37.2-point gap on identical weights, and by our math the higher-scoring run was also 28% cheaper. ARC Prize reported that Astra used fewer actions than the median human tester on 96% of levels. ARC Prize has also said it is not claiming the result proves AGI.
Terminal-Bench shows the same effect at a smaller scale. OpenAI reports 57.9% on Terminal-Bench 4.0. Artificial Analysis measured 59.1% in its Intelligence Index run and 56% inside Codex for its Coding Agent Index. The spread comes from the harness each run used, not from the model.
What are GPT-6 Astra’s API prices, context window and limits?
Standard API pricing is $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache-write tokens and $50 per million output tokens. The model accepts text and images, returns text, and has a 1,050,000-token context window with 128,000 tokens of max output.
| Price per 1M tokens | GPT-6 Astra | Claude Fable 5.1 | GPT-5.6 Sol |
| Input | $10.00 | $10.00 | $4.00 (promo) |
| Cached input | $1.00 | $0.25 | $0.40 |
| Output | $50.00 | $50.00 | $20.00 (promo) |
| Blended, 7:2:1 mix (our calc) | $7.70 | $7.17 | $3.08 |
Sources: OpenAI API model pages and Anthropic’s Fable pricing, checked September 10, 2026.
The details that change real bills, all from OpenAI’s model page:
- Long prompts cost more for the whole request. Above 272K input tokens, input and cache rates double and output rises 1.5x. That works out to $20 input, $2 cached and $75 output per million, starting at about 26% of the context window.
- Batch and Flex run at 50% of Standard, so $5 input and $25 output. At least one pricing aggregator lists that discounted rate as Astra’s main price. It isn’t.
- Fast mode costs 2x Standard: $20 input and $100 output per million tokens.
- Reasoning effort runs low, medium, high, xhigh and max. Simon Willison notes there is no reasoning-off setting.
- The knowledge cutoff is April 30, 2026. Sol’s is February 16, 2026.
- Rate limits start at 500 requests and 500,000 tokens per minute on Tier 1 and reach 15,000 requests and 40 million tokens per minute on Tier 5. The free API tier isn’t supported.
Astra is also available through Microsoft Azure and Amazon Bedrock, and OpenAI says it supports Zero Data Retention for eligible API customers.
Where does GPT-6 Astra lead, and where does it fall behind?
Astra leads on terminal-based agent work, business workflow automation, computer use, math and offensive security benchmarks. It falls behind on Humanity’s Last Exam with tools, the GDPval-AA v2 benchmark and presentation quality.
Independent results (Artificial Analysis, September 7-9, 2026)
- Terminal-Bench v4.0: 59.1%, against 52.0% for Fable 5.1 and 39.9% for Sol.
- AutomationBench-AA, built on Zapier’s business workflow benchmark: 68.5% score. Astra completed every objective with no guardrail violation on 41.6% of 657 workflows, against 32.1% for Fable 5.1.
- GDP.pdf, reasoning over professional documents: all criteria passed on 31% of attempts, against 27% for Sol.
- AA-Briefcase, long multi-week knowledge-work projects: about 90 Elo points above Sol.
Vendor-reported results (OpenAI launch table, best score at any effort)
- OSWorld 2.0 computer use: 72.6% at about 40 minutes per task, against Sol’s 65.7% at about 75 minutes.
- Agents’ Last Exam: 59.3%, against 55.5% for Claude Opus 5 and 53.6% for Sol.
- FrontierMath Tier 4: 97.6%.
- Terminal-Bench Science 0.1: 64.6%, against 52.6% for Fable 5.1.
- MRCR v2 8-needle retrieval at 512K-1M tokens: 96.3%, against 73.8% for Sol.
- ExploitBench: 100%, against 78.5% for Sol. SRE-Bench binary reverse engineering: 88.0% against 55.9%.
This table reports each model’s best score at any effort, and some of its Sol figures are higher than in OpenAI’s July GPT-5.6 table (OSWorld 2.0: 65.7% here, 62.6% in July; ExploitBench: 78.5% here, 73.5% in July). Compare Astra with the Sol numbers in this table, not with older ones.
Where GPT-6 Astra falls behind
- Humanity’s Last Exam with tools: 57.2% against Fable 5.1’s 65.0%, in OpenAI’s own table.
- GDPval-AA v2, economically valuable tasks across 44 occupations: about 45 Elo below Sol.
- AA-Briefcase presentation quality: lower than Sol, which still leads every model there. Fable 5.1 also scores higher than Astra on AA-Briefcase overall.
- DeepSWE inside Artificial Analysis’s Codex run: 68% against Sol’s 72%.
One pattern stands out. On GDPval-AA v2, Astra averaged 24 turns per task, against 45 for Sol and 60 for Fable 5.1. Artificial Analysis reported the lower score and the shorter runs side by side without linking them. Our guess is that the same efficiency that makes Astra cheap sometimes ends a task too early. Test that on your own workload before trusting it.
Can you use GPT-6 Astra in ChatGPT Plus?
Yes, but not in the Chat model picker. Plus includes Astra inside ChatGPT Work and Codex, while GPT-6 Pro in Chat, which runs on Astra, is limited to Pro, Business and Enterprise plans, according to OpenAI’s help center.
| Plan | Astra in Work and Codex (est. local messages per 5 hours) | GPT-6 Pro in Chat |
| Plus | 5-45 | No |
| Pro 5x | 25-225 | Yes |
| Pro 20x | 100-900 | Yes |
| Business Standard | 5-45 | Yes, 15 Pro messages a month shared with Sol Pro |
| Enterprise | By workspace | Yes, once an admin enables it |
Source: OpenAI Help Center, updated September 9, 2026. OpenAI calls these estimates, not fixed limits; weekly caps may also apply.
GPT-6 Pro in Chat has its own caps. Pro $200 includes 200 GPT-6 Pro messages a week. Pro $100 gets 50 a week, shared with GPT-5.6 Sol Pro. Business Premium gets 50 a week, also shared. A Pro subscription does not include unlimited GPT-6 Pro use.
On Plus, Sol’s estimate is 10-100 messages per five hours, so Astra works through the same allowance roughly twice as fast. OpenAI’s own advice cuts the other way: Astra at low effort can beat Sol at high effort, so if Sol on high has been working for you, start Astra on low or medium. Astra in Codex needs CLI version 0.153.0 or newer. During launch week, OpenAI gave eligible Plus, Pro and Business subscribers a banked reset on September 3 and another on September 4, then an automatic reset on September 7.
What do GPT-6 Astra’s safety and alignment numbers show?
Astra is the first OpenAI model rated Critical for cybersecurity capability under its Preparedness Framework. The system card reports fewer misaligned actions than GPT-5.6 Sol, and also a clear drop in how easily its chain of thought can be monitored.
| Measure (OpenAI system card, September 3, 2026) | GPT-6 Astra | GPT-5.6 Sol |
| Misaligned outcomes in realistic work apps, no confirmation policy | 3.4% | 18.8% |
| Same test with the default confirmation policy | 3.0% | 8.0% |
| Severity 3+ misalignment flags across 54,218 simulated Codex tasks | 34 (0.063%) | 73 (0.135%) |
| Indirect prompt injection success rate, Gray Swan (1,810 attacks) | 8.5% | 27.0% |
| Chain-of-thought controllability, 750-1,250 token traces (lower is safer) | 60.9% | 16.1% |
| Evaluation awareness in the Codex simulation | 9.6% | 2.8% |
The last two rows are the uncomfortable ones. OpenAI found Astra better at shaping its own reasoning traces and more likely to notice it is being tested. Apollo Research measured verbalized evaluation awareness in 50.6% of Astra samples at max effort and said low misbehavior rates under those conditions don’t say much about alignment either way.
UK AISI found Astra still carried out simulated supply-chain attacks on hard cyber challenges: 60 of 499 samples when the task scope didn’t mention internet access, and 2 of 500 when the scope ruled it out. No real systems were reachable. In production, OpenAI says the launch version refuses advanced offensive tasks such as writing proof-of-concept exploits, with wider defensive access planned through its Daybreak program.
What are developers saying about GPT-6 Astra?
Early reviewers mostly land on the same split: Astra for fast, token-lean work and computer use, Fable 5.1 for the biggest builds. We haven’t run Astra ourselves, so this section only reports what named practitioners published.
- Dan Shipper, Every (September 3): called Astra a big upgrade over Sol and the best writing model he has used, praised its computer use, and flagged a habit of overcomplicating things. He said Fable 5.1 still gets his largest tasks.
- Simon Willison (September 4): ran his pelican-on-a-bicycle SVG test at every effort level. Astra on low beat every Sol result, at 9.55 cents. Below max, it still didn’t reliably put the pelican’s legs on both sides of the frame. Astra used fewer output tokens than Sol at four of the five shared effort levels (at high it used 3,671 against 3,454), and it cost more than Sol at every level.
- Bindu Reddy (X, September 4): rated Astra slightly below Fable 5.1, cheaper and faster, weaker on long agentic loops and much better at browser use and 3D renders.
That lines up with the published numbers. At the 53-point level Astra is the cheaper model per task and leads the computer-use benchmarks, while Fable 5.1 keeps the edge on AA-Briefcase, the long-horizon knowledge-work test.
Should you pick GPT-6 Astra, Claude Fable 5.1 or GPT-5.6 Sol?
Pick on how your workload spends tokens. The same Index score costs $3.26 on Astra and $7.63 on Fable 5.1, and Sol is cheaper still if you can accept a lower ceiling.
- Choose GPT-6 Astra if you pay per task for terminal agents, business workflow automation or computer use, and you want a 53-level score at the lowest cost per task on the frontier.
- Choose Claude Fable 5.1 if your work is long, multi-step project builds or research with heavy tool use, or your traffic is mostly cached context with short outputs, where its $0.25 cache reads help.
- Choose GPT-5.6 Sol if you can accept a score six points lower in exchange for the cheapest index point of these three, or you need more messages per five hours on Plus. Watch the November 21 promo date.
- Whichever you pick, start Astra at low or medium effort and move up only when a result misses. Higher effort doesn’t guarantee a better answer, and OpenAI says so itself.
What are the limits of these GPT-6 Astra benchmarks?
Every number on this page describes one setup on one date. Keep these limits in view before you commit budget:
- Index versions drift. Astra’s Artificial Analysis rank changed twice in four days. Never compare scores across versions.
- Vendor tables show best-case settings. OpenAI reports the maximum score at any effort, run in its research environment or API, which can differ from ChatGPT.
- Harnesses move scores by tens of points. ARC-AGI-3 moved 37.2 points on identical weights.
- Some of our figures are derived. Values marked ≈ come from rounded percentages and carry a few percent of error.
- Prices are temporary. Sol’s promotion and any future Astra change will shift every ratio here.
- Evaluation awareness weakens the safety evidence. Apollo Research says so directly about its own results.
- We did not test Astra hands-on. Everything here is published data plus our calculations.
How were these GPT-6 Astra statistics compiled?
We took prices and specs from OpenAI’s API model pages, the ChatGPT help center and Anthropic’s Fable 5.1 pricing. Independent benchmark numbers come from Artificial Analysis (Index v4.2 and v4.3), ARC Prize and Epoch AI, each dated. Vendor claims come from OpenAI’s launch post and system card and are labeled as vendor-reported. Community views are attributed to named people. Our own calculations (cost per index point, blended prices, promo-expiry ratio) use only those published figures. Full method: How We Research and Verify.
GPT-6 Astra FAQ
Bottom line
GPT-6 Astra matches Claude Fable 5.1 on the Artificial Analysis Intelligence Index v4.3 for 43% of the cost per task, which makes it the cheapest way to get a top Index score as of September 10, 2026. It isn’t the best model at everything: Fable 5.1 still leads long knowledge-work projects, and Sol still sells the cheapest index point of the three. For the wider picture on how AI answer engines are changing, see our AI hub.
Key takeaways
- GPT-6 Astra ties Claude Fable 5.1 at 53 on Artificial Analysis v4.3 and costs $3.26 per task against $7.63.
- Both list at $10/$50, but Astra uses about 65% fewer output tokens per task, so its per-task bill is lower even though Fable 5.1’s cache reads are cheaper.
- Astra’s ranking moved from 5 points behind to tied in four days because Artificial Analysis changed its index twice.
- ARC-AGI-3 scores range from 62.7% to 99.9% depending on the harness.
- OpenAI rates Astra Critical for cyber capability and reports fewer misaligned actions than Sol, with weaker chain-of-thought monitorability.