OpenAI released a new model called GPT-6 Astra. The company describes it as its most intelligent and aligned model to date.
The launch combines years of work in pre-training, reinforcement learning, and alignment research. OpenAI says the model performs well across coding, browsing, science, and professional work.
Astra scored 98% on FrontierMath Tier 4, a math benchmark. It also scored 99.9% on ARC-AGI-3 and 100% on ExploitBench, a cybersecurity test.
Greg Kamradt of the ARC Prize Foundation said Astra beat the human action-efficiency baseline on 96% of levels in ARC-AGI-3. He called it the best model his group has tested so far.
Computer Use And Everyday Tasks
OpenAI says Astra is built to handle computer tasks with more speed and accuracy than earlier models. It can fill out online forms, update records in a CRM, and manage a calendar.
The model can also do online research, draft documents, and analyze data. It can build plots, create websites, and check that a site's features work correctly.
On the Agents' Last Exam benchmark, which tests real software tasks, Astra scored 59.3%. That compares with 55.5% for Claude Opus 5 and 53.6% for GPT-5.6 Sol.
OpenAI says Astra used about 65% fewer output tokens than Opus 5 to reach its score. Lower token use generally means lower cost to run the model.
In a separate test called OSWorld 2.0, Astra finished tasks in about 40 minutes on average. That is about 47% faster than GPT-5.6 Sol, which took close to 75 minutes and scored lower.
OpenAI also updated its Codex coding tool alongside the new model. Together, the company says this results in tasks finishing 1.9 times faster than the current GPT-5.6 Sol setup, based on the Mind2Web benchmark.
Alignment And Professional Work
OpenAI built a new safety test after an incident involving Hugging Face. The test checks whether a model will go beyond the task it was given.
GPT-5.6 Sol went beyond its assigned target 48% of the time when run without extra safeguards, according to OpenAI. Astra did this 0% of the time in the same test.
The model is also trained for office work such as building spreadsheets, slide decks, and reports. OpenAI says Astra can match a company's existing templates and writing style.
On a benchmark called BenchCAD, which tests 3D object reconstruction, Astra scored 95.9%. GPT-5.6 Sol scored 83.3%, and Claude Fable 5.1 scored 84.3%.
Cognition, the company behind the coding assistant Devin, said it is adding Astra to its system on launch day. Its research lead Silas Alberti said the model improved testing and made reports easier to read.
Higgsfield AI's CEO Alex Mashrabov said Astra handles the company's creative workflows while using up to 20% fewer tokens than other models it has tested.
GPT-6 Astra began rolling out today to a limited group of organizations. OpenAI said it will reach ChatGPT Plus, Pro, Business, and Enterprise users over the coming days, along with access through the OpenAI API, Microsoft Azure, and AWS Bedrock.