Illustration representing GPT-6 Astra, OpenAI's newest AI model powering ChatGPT

If you searched for “ChatGPT 6 Astra,” you landed close to the right place, but not quite on the right name. OpenAI’s newest model is officially called GPT-6 Astra. ChatGPT is simply the app people use to reach it, alongside the OpenAI API, Microsoft Azure, and AWS Bedrock. That naming mix-up is understandable — OpenAI has trained everyone to think in “ChatGPT” terms for three years — so this guide uses both phrasings where it helps, but sticks to the accurate one, GPT-6 Astra, everywhere it matters.

What follows is based on OpenAI’s own launch materials, its published system card, and reporting from outlets that covered the release. Where something is OpenAI’s claim rather than an independently verified fact, that distinction is called out directly instead of stated as settled truth.

What Is GPT-6 Astra? (And Why People Call It “ChatGPT 6 Astra”)

GPT-6 Astra is OpenAI’s flagship large language model, positioned as the successor to GPT-5.6 Sol. OpenAI began rolling it out to a limited set of organizations on September 3, 2026, with wider availability to ChatGPT Plus, Pro, Business, and Enterprise users following within days. It’s also available through the OpenAI API under the model name gpt-6-astra, as well as through Microsoft Azure and AWS Bedrock.

OpenAI describes it in blunt terms: the most intelligent and aligned model it has shipped so far. That’s a marketing framing, of course, but it’s backed by a fairly unusual step for OpenAI — publishing a long list of side-by-side benchmark comparisons against its own previous model and against competing systems, rather than just a highlight reel.

One naming wrinkle worth clearing up: some searches also surface “GPT-6 Pro,” which is a separate, higher-tier variant available to Pro, Business, and Enterprise subscribers rather than a different generation of model.

Why This Release Matters

Most model releases nudge benchmark scores up a few points. OpenAI’s own reporting on Astra claims something larger: near-saturation on several evaluations that were considered difficult just months earlier. On its FrontierMath Tier 4 evaluation, OpenAI reports a 97.6% score, and on ARC-AGI-3 — a benchmark designed around novel problem-solving rather than memorized patterns — it reports 99.9%, describing the result as reaching human parity on action efficiency.

The more consequential shift for everyday use, though, isn’t the raw intelligence score. It’s computer use — the model’s ability to operate a screen, click through interfaces, fill out forms, and complete multi-step tasks inside real software without a dedicated API integration. OpenAI reports that on its OSWorld 2.0 latency benchmark, Astra completed comparable tasks in roughly 47% less time than GPT-5.6 Sol.

That’s the practical story: a model that doesn’t just answer questions better, but can sit inside your actual workflow — spreadsheets, browsers, code editors, design tools — and get things done with less hand-holding.

Shivam Thakur logo
GPT-6 Astra is OpenAI’s newest model, released starting September 3, 2026. It replaces GPT-5.6 Sol as the company’s flagship system, with the biggest reported gains in computer use, coding, professional document work, and scientific reasoning.

Key Features and Capabilities

Based on OpenAI’s launch documentation, here’s what Astra is built to do:

Computer use that’s faster and more accurate

Astra can operate a mouse and keyboard inside a virtual environment — filling out forms, updating records in a CRM, running QA checks on a website it built, or troubleshooting software it’s installing. OpenAI’s benchmark figures show meaningful gains here over GPT-5.6 Sol, both in accuracy and in time per task.

Stronger coding and agentic software work

OpenAI reports state-of-the-art results for Astra on several coding benchmarks, including Terminal-Bench 4.0 and its internal database-migration tasks. Coding partners quoted in OpenAI’s release, including engineers at Jane Street and Cognition, described clearer communication and less iteration to reach production-ready code — though these are vendor-supplied endorsements, not independent audits.

Professional document and spreadsheet output

Astra is trained to follow existing templates closely — producing slide decks, spreadsheets, and documents that match a company’s existing visual style rather than generic defaults. OpenAI frames this as a fix for a common complaint: AI-generated documents that technically contain the right information but look nothing like what a team actually uses.

Scientific and mathematical reasoning

OpenAI highlights two research contributions attributed partly to Astra: a tightened bound on the gap between consecutive prime numbers, and progress on a decades-old bound related to large prime gaps. These are framed as model-assisted results verified by human mathematicians, not fully autonomous discoveries.

Better task continuity

OpenAI says Astra is less likely to lose track of an original request when a user adds new instructions mid-task — a known failure mode in earlier models, where a follow-up question could accidentally overwrite the original goal.

How Astra Differs From GPT-5.6 Sol

Rather than list features in isolation, it helps to see how OpenAI’s own reported numbers stack Astra against its direct predecessor, GPT-5.6 Sol.

AreaGPT-6 AstraGPT-5.6 Sol
FrontierMath Tier 4 (v2)97.6%83.0%
ARC-AGI-399.9%7.8%
Terminal-Bench 4.0 (coding)57.9%37.3%
OSWorld 2.0 (computer use, offline set)72.6%65.7%
GPQA Diamond (science reasoning)96.0%94.6%
ExploitBench (offensive cyber capability)100.0%78.5%

Figures as reported by OpenAI at launch.

That last row matters for a different reason than the others: OpenAI itself flags Astra’s jump in cybersecurity capability as crossing into what it calls a “Critical” risk threshold under its internal safety framework, which is why it ships with tighter restrictions on offensive security tasks than the rest of the model’s capabilities (more on that below).

OpenAI also reports that Astra makes fewer inaccurate claims about its own capabilities — about three times less often than GPT-5.6 Sol in its internal testing — and that it never attempted to bypass a deliberately evadable safety review in one internal evaluation, compared to a 48% bypass rate for its predecessor under similar conditions without production safeguards.

GPT-6 Astra GPT-5.6 Sol FrontierMath T4 97.6% 83.0% ARC-AGI-3 99.9% 7.8% Terminal-Bench 4.0 57.9% 37.3% OSWorld 2.0 72.6% 65.7% GPQA Diamond 96.0% 94.6% ExploitBench 100.0% 78.5%
Benchmark scores as reported by OpenAI at launch. Bars are scaled to a 100% axis.

Practical Use Cases

Here’s where the reported capabilities translate into things an actual user or team might do with Astra:

  • Research and drafting — pulling information from the web and turning it into a structured document or email draft.
  • Spreadsheet and financial modeling — OpenAI cites a demo where Astra completed Financial Modeling World Cup–style challenges via computer use roughly four times faster than the human competition winner.
  • Software engineering — working inside a codebase across multiple files, running tests, and iterating on fixes with less back-and-forth.
  • Website and app prototyping — building, hosting, and sharing a working website or small app directly from a prompt using OpenAI’s Sites feature in ChatGPT.
  • Data science and lab work — inspecting sequencing data, running simulations, and generating plots inside specialized scientific software.
  • Defensive cybersecurity — secure code review and patch verification, with more advanced offensive tasks intentionally restricted at launch.

Availability and Pricing

GPT-6 Astra usage is included within existing ChatGPT subscription allowances, with the option to buy additional usage credits. Enterprise workspace administrators must manually enable it — it’s off by default for Enterprise accounts at launch.

ItemPrice
Input tokens$10 per million tokens
Output tokens$50 per million tokens
Fast modeUp to 2x speed at roughly 2x the standard price

OpenAI API pricing for GPT-6 Astra (standard mode).

The model also supports Zero Data Retention for eligible API customers, and OpenAI says it’s testing a “Private Safety Processing” system intended to preserve safety monitoring while limiting how much customer data is retained for that purpose.

Limitations and Safety Considerations

A model this capable comes with tradeoffs OpenAI is fairly upfront about in its published system card:

  • Cybersecurity restrictions. Because Astra crossed OpenAI’s “Critical” threshold for cyber capability, it’s designed to refuse advanced offensive tasks like building proof-of-concept exploits, even though it can perform defensive tasks like patching and code review.
  • Harder-to-monitor reasoning on simple tasks. OpenAI’s own evaluations found Astra’s internal reasoning is somewhat harder to audit than its predecessor’s on simpler tasks, which the company describes as an ongoing research priority rather than a solved problem.
  • Safety checks can interrupt legitimate work. OpenAI acknowledges that its added monitoring layer can occasionally pause or stop tasks — including legitimate defensive security work — while it continues tuning the system to reduce false positives.
  • Access is tiered and evolving. Not every plan or region gets identical access at launch, and enterprise rollout depends on administrator settings and rate agreements.

It’s also worth treating OpenAI’s benchmark numbers the way you’d treat any vendor’s own performance claims: informative, but not a substitute for testing the model against your own tasks before relying on it for anything consequential.

Final Take

Strip away the launch-day numbers and GPT-6 Astra tells a fairly simple story: the interesting progress has moved away from “answers better questions” and toward “finishes more of the job.” The FrontierMath and ARC-AGI-3 jumps make headlines, but the change you’ll actually feel day to day is computer use — a model that can sit inside a browser, a spreadsheet, or a codebase and carry a multi-step task to the end without being walked through every click.

That shift cuts both ways. A model capable enough to operate your software unsupervised is also capable enough that OpenAI flagged its own cyber capability as “Critical” and shipped it with deliberate restrictions. The tighter guardrails and the occasional interrupted task aren’t rough edges to be patched away — they’re the cost of the capability itself.

So the practical advice is unchanged from every model release before it: every figure above is OpenAI’s own. Pick the two or three tasks that actually matter in your workflow, run them against Astra and against whatever you use today, and let that decide it. Benchmarks tell you what a model can do on someone else’s problems. Only your own tests tell you what it does on yours.

FAQ
Let’s build

Putting a model like Astra to work?
Let’s build it properly.