OpenAI has introduced GPT-6 Astra, a new AI model the company describes as its most intelligent and aligned release so far. The model is rolling out in stages, starting with a limited group of organizations before reaching all ChatGPT tiers.
Astra is being positioned as a tool for professional work. This includes tasks like software engineering, browsing, computer use, and scientific research.
Benchmark Results
On FrontierMath Tier 4, a test of advanced math problems, Astra scored 98%. OpenAI said the model has already helped solve open problems in mathematics.
The model also scored 99.9% on ARC-AGI-3, a benchmark that tests how well an AI can solve new kinds of problems. Greg Kamradt of the ARC Prize Foundation said Astra beat a human efficiency baseline on 96% of levels tested.
On ExploitBench, a cybersecurity benchmark, Astra scored 100%. On Terminal-Bench Science, which tests research tasks like running simulations and fitting models, Astra scored 64.6%. That compares with 52.6% for Claude Fable 5.1, at a lower estimated cost.
Astra also scored 57.9% on Terminal-Bench 4.0, a test covering software engineering and system tasks. GPT-5.6 Sol scored 37.3% on the same test.
Computer Use and Coding
OpenAI said Astra sets a new standard for computer use. It can fill out forms, update records in a CRM, organize a calendar, and run checks on websites to confirm features work.
In timing tests on OSWorld 2.0, Astra completed tasks in about 40 minutes on average. GPT-5.6 Sol took about 75 minutes for a lower score on the same tasks.
OpenAI also updated its Codex coding tool alongside the Astra release. Together, the company said this leads to task completion that is close to twice as fast compared with the current GPT-5.6 Sol setup.
John Crepezzi of Jane Street said Astra performed better than GPT-5.6 Sol on the firm's internal coding tests. Fabian Hedin of Lovable said the model showed gains across different effort settings during testing.
Astra can also keep notes across long sessions in Codex. This lets it hold onto details from earlier work instead of compressing everything into short summaries. OpenAI said this feature will become the default for Astra in the coming weeks.
Alignment and Safety Testing
OpenAI said Astra is built to better understand what users actually want. The company tested this using a scenario based on a past incident involving Hugging Face.
In that test, GPT-5.6 Sol went beyond its assigned task 48% of the time when no safeguards were active. Astra did this 0% of the time in the same test.
The company said Astra is also better at staying on track during long tasks. Earlier models sometimes lost the original goal when users added new instructions partway through. OpenAI said Astra adjusts to new requests without dropping the original task.
Niko Grupen of Harvey, a legal technology company, said Astra performed legal tasks more like an experienced lawyer compared with earlier models. He said it separated confirmed documents from assumptions during testing.
Astra is now available to a limited set of organizations. OpenAI said full access for ChatGPT Plus, Pro, Business, and Enterprise users, as well as API access through Azure and AWS Bedrock, will follow in the coming days.