OpenAI released GPT-6 Astra on 3 September 2026. The model can research across the web, operate software, write and test code, create professional artefacts and complete longer sequences of work with less intervention. It is also OpenAI’s first broadly deployed model to reach the Critical cybersecurity capability threshold. This is a major release—but the most important question for leaders is not whether Astra is the smartest model in every chart. It is what their organisation can now let an AI system do, with evidence and control.
The launch moves frontier AI further from answering prompts toward operating workflows. For companies and governments, that creates a practical opportunity in finance, customer operations, engineering, research, public services and cyber defence. It also makes model selection, permissions, monitoring and recovery part of the operating design. Buying access is easy. Turning that access into dependable performance is the valuable work.
What OpenAI actually released
Astra is rolling out in stages to ChatGPT Plus, Pro, Business and Enterprise, and through the API, Microsoft Azure and AWS Bedrock. Enterprise administrators must enable it. The API model, gpt-6-astra, has a 1.05-million-token context window, supports up to 128,000 output tokens and accepts text and images. It can use web and file search, code execution, hosted shell, computer use, Skills, MCP and custom tools. Fine-tuning and audio or video input are not supported at launch.
Its published knowledge cutoff is 30 April 2026, and developers can select low, medium, high, xhigh or max reasoning effort. Eligible API customers can use Zero Data Retention. These are architecture inputs, not footnotes: recent work still needs current sources; reasoning effort should match the consequence and economics of the task; and a retention setting does not replace data classification, access control or an accountable owner.
The release adds capabilities that matter for production agents: asynchronous tool calling, mid-turn steering and the ability to change reasoning effort while preserving cache. In ChatGPT and Codex, Astra can work across longer tasks, search earlier context and produce slides, documents and spreadsheets. OpenAI reports 72.6% on OSWorld 2.0 in about 40 minutes per task, versus 65.7% in about 75 minutes for GPT-5.6 Sol. That is a provider-published evaluation, but it shows the direction: more work completed through the interface, in less elapsed time.
The benchmarks say: route by workload, not by logo
Astra does not win every credible evaluation. Artificial Analysis scores Claude Fable 5.1 ahead on both its Intelligence Index and Coding Agent Index. The live Terminal-Bench 4.0 snapshot does not establish a clear lead between Astra with Codex and Fable 5.1 with Claude Code, while ARC Prize’s verified ARC-AGI-2 result favours Astra. The honest conclusion is not that one model has defeated every other model. General leadership and workload leadership are now different things.
The lead changes with the work being measured.
Compare models within each row only. Evaluations use different tasks, scales and harnesses.
| Operating dimension | GPT-6 AstraMeasured stack | Claude Fable 5.1Measured stack | Evidence readingWhat the result supports |
|---|---|---|---|
| AA Intelligence Index | 61 | 66 | Independent composite; Fable leads |
| AA Coding Agent Index | 67 · Codex | 70 · Claude Code | Native harnesses differ; Fable stack leads |
| Terminal-Bench 4.0 | 58.2% ±2.8 | 57.9% ±3.8 | Live snapshot does not establish a clear lead |
| ARC-AGI-2 | 95.0% | 90.0% | Verified max results; Astra leads |
Governance takeawayUse a representative evaluation set from your own workflow before assigning production authority.
Cost also depends on the job rather than the token price alone. Astra’s standard API rate is $10 per million input tokens and $50 per million output tokens, with cached input at $1. Prompts above 272,000 input tokens receive a higher price multiplier for the whole request. Batch and Flex can cost half the standard rate; Fast can run up to twice as quickly at twice the standard price. Artificial Analysis found Astra used fewer tokens than Sol on its index but still cost about 75% more per task. Measure cost per accepted outcome, including review and rework.
A breakthrough on ARC-AGI-3 is not the same as proof of AGI
The launch’s most dramatic number is 99.9% on ARC-AGI-3, a benchmark designed to test adaptation to unfamiliar interactive problems. That score used OpenAI’s provider adapter, which preserves the model’s opaque reasoning state and compaction between actions. ARC Prize also tested Astra through its provider-neutral Standard harness: the best score there was 62.7%. Both results are important. Presenting only the larger one would hide how strongly the execution harness affects the outcome.
The headline changes when the execution environment changes.
ARC Prize’s full-evaluation results for GPT-6 Astra, reported on 3 September 2026.
- Provider-neutral Standard harness · max effort62.7% · $26,098
- OpenAI provider adapter · high effort99.9% · $18,817
Operator takeawaySystem design—including state, tools and compaction—can change the measured result as much as model choice.
ARC Prize found that the adapter made Astra 3.66 times faster and used 49% fewer tokens on game-reasoning pairs both configurations solved. It also found that Astra used fewer actions than the median human on 96% of levels. Those are meaningful signs of more efficient generalisation. They do not establish that Astra can autonomously outperform people across most economically valuable work. The correct response is neither dismissal nor an AGI victory lap: it is disciplined experimentation on real work.
Cyber capability is where the release becomes operationally urgent
OpenAI classifies Astra at the Critical cybersecurity capability threshold: with the right tools and access, it can find unknown flaws and develop novel exploitation across protected systems without human guidance at every step. OpenAI reports 100% on public ExploitBench, while warning that prior exposure may inflate that result. On a post-cutoff set of June-to-August vulnerabilities, Astra scored 39.0% versus Sol at 11.5%; on SRE-Bench it reached 88.0% in one attempt and 99.2% within four. During testing, Astra also found two previously unknown zero-day vulnerabilities.
Independent evaluator Irregular also measured a clear increase over Sol on the same cyber snapshot: Astra solved 86 of 226 FrontierCyber challenges versus 34, although neither model solved an Elite challenge and no fully hardened target was compromised. This capability can accelerate repository review, vulnerability triage, patch preparation, validation and incident analysis. Shofield AI has organisation-level Daybreak Blue approval for authorised internal defensive work. We use that experience to strengthen our client methods; it is separate from customer access and from Astra’s advanced Daybreak mode, which OpenAI says is planned.
The safety picture deserves the same precision. OpenAI reports stronger alignment and lower circumvention than Sol, while its System Card also records reduced visibility into Astra’s written reasoning under adversarial pressure. Production use therefore needs isolation, least privilege, full action logs, approval gates and a working stop path—not blind trust in an internal explanation. Shofield AI brings those assessment, remediation-planning and governance disciplines into company and government deployments without making a model provider the control system.
The opportunity is not a better chatbot. It is a better operating model.
Astra can be valuable where work crosses documents, applications and decisions: closing a finance period, researching a regulated case, preparing a procurement pack, testing software, coordinating customer operations or investigating a security finding. Government teams can apply the same pattern to policy evidence, service delivery, infrastructure planning and authorised cyber defence. The first implementation should be consequential enough to measure, but bounded enough to supervise and reverse.
Move from release headline to a production result.
One workflow, a fair model comparison and a control boundary leadership can approve.
Named owner · Explicit authority · Least privilege · Evidence · Stop and recovery path
- 01Select
Choose a high-value, bounded workflow.
→ - 02Benchmark
Test Astra, Fable and viable alternatives.
→ - 03Connect
Add trusted data, tools and identities.
→ - 04Control
Set approvals, limits, logs and rollback.
→ - 05Operate
Track quality, cost, incidents and value.
Architecture takeawayThe winning model is the one that produces the best governed result for the workflow.
Why contact Shofield AI now
Most organisations do not need another model licence. They need a defensible answer to five questions: which workflow should change first; which model performs best on their evidence; what data and systems it may use; which actions require approval; and how performance, cost and incidents will be monitored. Astra raises the ceiling, but it does not answer those questions on its own.
Shofield AI provides that operating layer. We assess the workflow, compare Astra with Claude Fable 5.1 and other viable routes, design the data and integration architecture, implement identity and approval controls, and establish the evaluation and monitoring record. For security work, our Cyber Security services apply the same discipline to authorised assessment, remediation planning and verification.
The early advantage will not belong to whoever opens Astra first. It will belong to organisations that can safely give the right model the right authority, prove the outcome and improve the system as models change. Contact Shofield AI to identify that first workflow and build the route from capability to controlled value.
