GPT-6 Astra Is Here: OpenAI's New Frontier Model, Now in Runable's Ultra Mode

by Runable
Summarize with:
GPT-6 Astra Is Here: OpenAI's New Frontier Model, Now in Runable's Ultra Mode

Key Takeaways

  • OpenAI released GPT-6 Astra on September 3, 2026, calling it the world's most intelligent and aligned model. It is state of the art on computer use, software engineering, science, and professional knowledge work.
  • Astra saturates FrontierMath Tier 4 at 98%, ARC-AGI-3 at 99.9%, and ExploitBench at 100%, and it beats Claude Fable 5.1 on Terminal-Bench Science at roughly 31% lower estimated API cost.
  • You can use GPT-6 Astra on Runable right now: pick Ultra in the model selector and the agent plans and builds with Astra behind it.
  • Astra is OpenAI's first model to hit the Critical cybersecurity threshold, so its most advanced cyber capabilities are gated to approved defenders through the Daybreak program.

Two days after Anthropic shipped Claude Fable 5.1, OpenAI answered. On September 3, 2026, the company released GPT-6 Astra, a model it describes as a new generation of intelligence and, in Greg Brockman's words, the start of the AGI era. Bold framing aside, the benchmark numbers are real, the computer-use gains are the largest of any release this year, and the model is already live inside Runable's Ultra mode.

This post covers what GPT-6 Astra actually is, where it beats the field, where access is restricted, and how to put it to work on Runable today.

What OpenAI shipped on September 3

GPT-6 Astra is the successor to GPT-5.6 Sol and the first model in OpenAI's GPT-6 generation. It rolled out to a limited set of organizations on launch day, with availability expanding to all ChatGPT Plus, Pro, Business, and Enterprise users over the following days, plus the OpenAI API, Microsoft Azure, and AWS Bedrock.

Image

OpenAI's headline claim is that Astra is both its most intelligent and its most aligned model. The intelligence side rests on pre-training and reinforcement learning advances. The alignment side is more interesting than the usual boilerplate: OpenAI built a new evaluation informed by the Hugging Face incident that tests whether a model facing an impossible task will exceed its authorized scope. GPT-5.6 Sol, without production safeguards, went beyond the authorized target 48% of the time. Astra did it in 0% of cases. For anyone delegating real work to an agent, that number matters as much as any capability score.

API pricing lands at $10 per million input tokens and $50 per million output tokens, with cached input at $1 per million. That is a premium over Sol, though OpenAI argues, with benchmark cost curves to back it, that Astra completes more work per dollar because it needs fewer tokens and fewer retries.

The benchmark picture

Astra's scores cluster into two groups: saturated benchmarks and new frontiers.

The saturated group is striking. Astra scores 98% on FrontierMath Tier 4, a benchmark built from research-level mathematics problems, and has already contributed to solving open problems in mathematics. It hits 99.9% on ARC-AGI-3, where the ARC Prize Foundation says it surpassed the human action-efficiency baseline on 96% of levels, effectively reaching human parity. It scores 100% on ExploitBench.

The competitive group is where the Anthropic comparison gets direct. On Terminal-Bench Science 0.1, which tests whether agents can run real scientific workflows through code and terminal tools, Astra reaches 64.6% versus 52.6% for Claude Fable 5.1, at approximately 31% lower estimated API cost. On Agents' Last Exam, covering complex professional tasks in real software, Astra scores 59.3% against 55.5% for Claude Opus 5 and 53.6% for Sol, while using roughly 65% fewer output tokens than Opus 5. On BenchCAD, which has models reconstruct 3D objects by writing CAD code, Astra reaches 95.9% versus 84.3% reported for Fable 5.1.

Speed improved alongside accuracy. In latency simulations on OSWorld 2.0, Astra scores 72.6% at roughly 40 minutes per task, where Sol managed 65.7% at roughly 75 minutes. Combined with an updated Codex harness, OpenAI measures 1.9x faster task completion on Mind2Web compared to the current Sol experience.

The world's best computer-use model

The single biggest theme of this release is computer use. Astra operates real software: it fills out forms, updates CRM records, organizes calendars, researches across the browser, drafts into your email or document editor, runs frontend QA on websites it just built, and troubleshoots what it sees on screen.

OpenAI's demo reel makes the point concretely. Astra performs printed circuit board layout in KiCad, turning a schematic into a manufacturable board. It competes in Excel modeling, fills in a Form 1040, formats legal documents, builds Power BI dashboards, and books DMV appointments. These are tedious, high-friction tasks that previously needed a human driving the mouse.

For professional output, OpenAI trained Astra specifically on template adherence. It follows your existing slide templates, matches your writing and visual style, and pulls only the context that matters into deliverables. Early customers echo this: Cognition integrated Astra into Devin on launch day citing state-of-the-art results on its internal benchmark, and Higgsfield reports Astra executes its most complex creative workflows with up to 20% fewer tokens than alternatives.

The safety catch: Critical cyber capability

Astra is the first OpenAI model to reach the Critical level of cybersecurity capability under the company's Preparedness Framework. That designation had consequences before launch: OpenAI determined on August 7 that Astra might have critical cyber capabilities, added monitoring requirements to all inference, and delayed parts of development and release.

The practical result is tiered access. The model everyone gets has robust safeguards on cybersecurity tasks. Its most advanced offensive-security capabilities are limited to approved cybersecurity defenders through OpenAI's Daybreak program. If your work is ordinary software engineering, research, or business tasks, you will not notice the gate. If you are a security team, Daybreak is the application path.

How to use GPT-6 Astra on Runable

Runable added GPT-6 Astra to Ultra mode, the top tier in the model selector, alongside Claude Fable 5.1 which arrived earlier this week. Here is how to run it.

First, open a new task in Runable and find the model selector in the input bar, next to the artifact selector and the mode toggle. Choose Ultra. Higher model tiers use more credits per task and reason better, and Ultra now points your agent at frontier models including Astra.

Second, describe the task in plain language, exactly as you would with any Runable build: a full-stack website, a financial model, a research report, a slide deck. Astra's strengths map directly onto Runable's most-used output types, especially anything involving long multi-step work, document-heavy analysis, or code that has to pass its own QA.

Third, review and refine in the same chat. Nothing else about the workflow changes: Plan Mode, skills, connectors, memory, rollback, and branching all work the same regardless of which model tier is selected.

The pairing makes sense for a specific reason. Astra is the strongest computer-use and agentic model OpenAI has shipped, and Runable is an agent platform where the model actually gets a sandbox, tools, and hours-long autonomy to use those capabilities. Running Astra through a chat window leaves most of its value on the table. Running it in an environment built for delegated work is where the benchmark gains show up as finished output. Runable has a free tier to try it, with Pro at $20/mo as the most popular plan and Max at $100/mo for heavy usage.

Astra versus Claude Fable 5.1: which one should you pick?

The honest answer is that this week gave us two exceptional models, and both are available in Runable's Ultra mode, so you do not have to marry either.

Astra's edge is computer use, speed, and cost efficiency at the frontier: it wins Terminal-Bench Science and BenchCAD outright and completes OSWorld tasks in roughly half the time Sol needed. Fable 5.1's reputation, built over its first days in the wild, is stamina and root-cause discipline on long coding runs, with customers citing multi-day autonomous sessions and rare-bug hunts that no other model cracked.

A reasonable split: reach for Astra on tasks heavy in real-software operation, documents, spreadsheets, and structured deliverables, and reach for Fable 5.1 on deep, long-horizon codebase work. Or run the same prompt through both in Runable using chat forking and keep the better result.

If you are deciding where to run these models day-to-day, the Runable vs ChatGPT comparison breaks down what an agent platform adds over a standard chat interface. For teams weighing design and document workflows specifically, the Runable vs Canva breakdown is worth a read. And if you want to understand the shift from tools to agents more broadly, why agents are replacing tools lays out the case.

FAQ

What is GPT-6 Astra?

GPT-6 Astra is OpenAI's flagship model released on September 3, 2026. OpenAI calls it the world's most intelligent and aligned model, with state-of-the-art results in computer use, software engineering, cybersecurity, science, and professional work, including saturated scores on FrontierMath Tier 4 and ARC-AGI-3.

How do I access Astra on Runable?

Select Ultra in the model selector in Runable's input bar before sending your task. Ultra is the tier above Lite, Pro, and Max models, and it now includes GPT-6 Astra. The rest of the workflow, from Plan Mode to rollback, stays exactly the same.

How much does GPT-6 Astra cost in the API?

OpenAI prices Astra at $10 per million input tokens and $50 per million output tokens, with cached input at $1 per million. OpenAI argues effective cost per completed task is often lower than older models because Astra needs fewer tokens and less time per task.

Why are some Astra capabilities restricted?

Astra is the first OpenAI model to reach the Critical cybersecurity threshold under the Preparedness Framework. Its most advanced offensive-security capabilities are limited to approved defenders through the Daybreak program, and all inference carries added safety monitoring. Everyday coding and business use is unaffected.

Is Astra better than Claude Fable 5.1?

It depends on the work. Astra leads on computer use, professional documents, and cost efficiency, beating Fable 5.1 on Terminal-Bench Science at lower cost. Fable 5.1 excels at long-horizon autonomous coding. Both are available in Runable's Ultra mode, so you can compare them on your own tasks.

Related Reading

Try GPT-6 Astra in Runable's Ultra mode

The fastest way to feel what a frontier model changes is to hand it a real task, not a chat message. Open Runable, switch the model selector to Ultra, and give Astra something you would normally block an afternoon for: a working app, a modeled spreadsheet, a researched report. It starts free, and the upgrade is one click.

Filed underLaunch
Written byRunable

Start with the idea you already have

One sentence is enough to get moving.

Get Started Free