Amid revelations that safety concerns have forced the company to abandon the planned rollout of GPT-6.1 Astra, independent researchers and benchmark creators warn that GPT-6 Astra's record-breaking scores rely on heavy prompt optimization rather than true general intelligence. This comes as experts warn about the risks of future AI systems and AI executives jointly call for a slowdown in AI research.
In announcing the model's release, OpenAI President Greg Brockman suggested that Astra signaled a fundamental turning point in AI evolution. "If we fast-forward a couple years, and we look back and say when was it really that AGI was created, I think it's going to be about this time, and I think it might be about this model," Brockman said during a Sept. 3 news conference marking the launch.
The benchmarks scores are impressive, but even the creator of one of the most important suggests acing it doesn't automatically mean we've reached AGI. (Image credit: Cheng Xin via Getty Images)As part of the testing-and-verification process for the new model, OpenAI released benchmark results across a range of tests that measure the capabilities of AI models at various tasks. OpenAI representatives claimed in a statement that the company's new model delivers "state-of-the-art" performance across a variety of fields, including software engineering, the autonomous use of computer programs, mathematics problems, and even scientific research.
Demonstrations, for example, highlighted Astra's ability to generate 3D CAD code; lay out printed circuit boards in CAD (computer-aided design) software; convert digital 3D models into interactive, video-game-style environments; and deploy hosted web applications directly from prompts.
"In our early testing, Astra stood out by approaching legal work the way a discerning lawyer does: it distinguishes documents from established records, surfaces unsupported assumptions, and converts gaps into concrete drafting positions," said Niko Grupen, head of applied research at legal services AI developer Harvey as part of OpenAI’s announcement.
On ARC-AGI-3 — a benchmark designed to measure how efficiently AI systems acquire new skills in novel, abstract environments — Astra achieved a headline score of 99.9%. Independent testing by the nonprofit ARC Prize Foundation, which oversees the benchmark, revealed that Astra surpassed "human action-efficiency" baselines on 96% of levels, taking 51.7% fewer actions per level, on average, than human solvers.
"Not only is this the best model we've ever tested," ARC Prize Foundation President Greg Kamradt said as part of the OpenAI announcement, "but it also represents a meaningful step change in frontier-model performance — not only in its ability to navigate and solve novel environments but also in how efficiently it learns to do so."
How much can we read into test results?
Many researchers have highlighted that Astra's top ARC-AGI-3 performance relied on a proprietary "Provider Adapter" harness, a specialized software scaffolding layer that's inaccessible to rival developers. When tested using the benchmark's neutral Standard harness, this score dropped to 62.7%. ARC Prize organizers also noted that saturating the benchmark does not constitute proof of achieving AGI — emphasizing that its closed, deterministic puzzle environments do not capture the open-ended complexity of the real world.
"When we launched ARC-AGI-3, "we made it clear that saturating the benchmark would not represent 'proof of achieving AGI,'" Kamradt wrote in a blog post. "Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI."
Discrepancies in the data were reportedly present even before the official announcement. As reported by Fortune, an embargoed draft provided to news organizations prior to the model's launch listed Astra's ARC-AGI-3 score at 98.6% before it jumped to 99.9% on the live site.
Anka Reuel and Mike Hardy, doctoral researchers in AI at Stanford University, also raised concerns over OpenAI quietly revising several published metrics post-launch, including halving Astra's reported hallucination rate from 4.2% to 2.0% before reverting it.
When we launched ARC-AGI-3, 'we made it clear that saturating the benchmark would not represent 'proof of achieving AGI'"
ARC Prize Foundation President Greg Kamradt
Speaking to Fortune, Reuel and Hardy attributed the shifting figures to "benchmaxxing" — an industry term that describes the practice of repeatedly rerunning evaluations under subtly tweaked prompts, scaffolds or compute allocations to hunt for peak theoretical scores.
In a 2025 paper published on pre-print server arXiv titled ‘Benchmarking is Broken -- Don't Let AI be its Own Judge’, experts including Princeton doctoral candidate Zerui Cheng and University of Luxembourg postdoctoral researcher Dr. Stella Wohnig warned that such practices create confusion and undermine trust, particularly when official technical documentation omits the basic methodology, making independent verification nearly impossible. Crucially, scores achieved through benchmaxxing reflect idealized performance ceilings under optimal lab conditions rather than practical, out-of-the-box reliability.
According to OpenAI's benchmark data, across specialized domain evaluations, GPT-6 Astra delivered 10% to 20% better performance than both its predecessor, GPT-5.6 Sol, and top market competitors.
Highlights include a 98% score on FrontierMath Tier 4 (v2), which evaluates expert-level mathematical reasoning; a perfect 100% on ExploitBench, which is designed to test offensive cybersecurity capabilities; 95.9% on BenchCAD, which measures an AI's ability to reconstruct 3D objects using CAD code; 57.9% on Terminal-Bench 4.0, which assesses complex terminal-based system administration and software engineering; and 41.4% on AutomationBench, which tests whether agents can complete multistep business workflows across applications.
But independent findings from external AI benchmarking firm Artificial Analysis contradict these results. Astra's score on the independent Artificial Analysis Intelligence Index — an aggregate metric that evaluates general reasoning across frontier models — remained completely flat, at 61, compared with GPT-5.6 Sol, while trailing competitors like Anthropic's Claude Fable 5.1 and Meta's Muse Spark 1.3.
While some benchmark results have seen improvement since the previous generation, Astra underperformed in other tests. Artificial Analysis' testing showed that against GDPval-AA v2 — an economically focused benchmark that measures real-world workplace tasks across 44 occupations — Astra suffered a significant drop in its relative leaderboard ranking compared with GPT-5.6 Sol. Additional regressions were observed in benchmarks that test customer service support, scientific Python programming, and long-context reasoning across large documents.
In terms of overall token use — the metering system that measures the fragments of text or data processed by AI models — OpenAI representatives said GPT-6 Astra achieved dramatic efficiency gains on complex, long-horizon tasks. Findings from Artificial Analysis confirmed that on software engineering benchmarks, Astra was approximately 70% more token-efficient than GPT-5.6 Sol, using roughly one-third the total tokens of its predecessor and one-fifth the tokens of Claude Opus 5.
AI systems today can do more than a simple back-and-forth text exchange, graduating to taking actions on our behalf. (Image credit: CFOTO via Getty Images)On professional workplace evaluations like Agents' Last Exam, the new model reduced output token consumption by up to 65% compared with Opus 5, while on general intelligence evaluations, Astra achieved a roughly 10% output token reduction at max effort compared with Sol.
On the other hand, this achievement is offset by a 250% increase in the base API pricing for GPT-6 Astra compared with GPT-5.6 Sol. Token prices jumped from $4/$20 to $10/$50 per million input/output tokens. As a result, Astra ends up approximately 75% more expensive per task than its predecessor on general intelligence evaluations, despite using 10% fewer output tokens. This trade-off has led researchers at Artificial Analysis to question whether AI labs are relying on expensive computational brute force to squeeze out minor gains, rather than achieving true technical breakthroughs.
Does achieving AGI even matter?
With its latest model, OpenAI has placed a stronger emphasis on safeguards and controls. The move follows a series of high-profile incidents in which frontier AI models inadvertently hacked into third-party networks and computer systems as part of routine testing operations, after the models went beyond what humans expected they would do when encountering challenging tasks.
According to OpenAI representatives, Astra did not exceed the scope of its tasks in any security evaluations. For comparison, GPT-5.6 Sol went beyond its authorized parameters in almost 50% of test cases. Hallucination rates have also been reduced significantly, falling to 4.2% on internal tests, compared with 12.2% for GPT-5.6 Sol (with independent testing by Artificial Analysis showing a drop from 92% to 51% at max effort). The model also showed an improved ability to handle ambiguous prompts.
David Wood — chair of London Futurists, a non-profit that hosts discussion groups on emerging technology — argued that beyond individual test results, Astra highlights a fundamental shift in how humans will interact with AI as models move from answering questions to autonomously pursuing complex goals in digital environments.
"The impressive benchmark results matter, but what matters more is the combination of intelligence, autonomy, computer use, and cybersecurity capability," he said in an email to Live Science. "That combination also demands caution. The more useful these systems become, the greater the consequences when they misunderstand our intentions, are misused, or find ways around the safeguards we give them."
Wood suggested that whether Astra can be classified as AGI is less relevant than recognizing the need for AI control and governance to keep pace with technical development.
"The possibility that AI could progress from systems like Astra towards all-round superintelligence makes much greater human vigilance, awareness, and collaboration increasingly urgent," he said.
Help us improve Live Science Pro: We're always trying to make our content better. Leave us feedback about Pro here.
'> OpenAI claims we've entered the AGI era — has GPT-6 Astra really demonstrated general intelligence? While mitochondrial function has long been linked to longevity and mitochondrial dysfunction has been tied to disease, Jonathan offers fresh insight into that dynamic, said study co-author Stephen Clark, chief scientist of the Kallel Foundation, a nonprofit in Nashville researching drug targets to promote human longevity.
Clark and colleagues published their findings Wednesday (Oct. 7) in the journal Science Advances.
Decoding Jonathan's secrets to long life
Nearly a decade ago, the researchers enlisted Joe Hollins, the vet tasked with Jonathan's care, to collect samples from the record-setting animal. Island authorities forbade the vet from drawing the tortoise's blood, fearing the risk of an infection.
"I didn't want to be the doctor that killed Jonathan," Clark told Live Science. Hollins collected samples via cheek swab instead.
When the samples arrived in the U.S., the researchers sequenced them, but when their computers kept crashing, they discovered that the DNA had been isolated from bacteria in Jonathan's mouth rather than the tortoise's own cells.
Clark "had to go back and beg" for permission to collect more data, he said. The vet then sent cheek scrape samples, which are collected with a slightly different tool than cheek swabs. This time, the DNA was in fact from Jonathan.
The DNA isolated from those cheek tissue samples was more fragmented than it would have been in a blood sample, so the researchers filled in the gaps using an existing reference genome from a 36-year-old Aldabra tortoise named Tank. They also compared Jonathan's DNA to that of a Galápagos tortoise (Chelonoidis abingdoni) named Lonesome George, who was believed to be over 100 years old at the time of his death in 2012.
The team found that, compared to Tank and Lonesome George, Jonathan carried 287 unique variants of genes that had previously been linked to aging-related pathways. These included genes involved in DNA repair and the function of telomeres — the protective tips at the end of chromosomes that tend to shorten with age.
These pathways are also related to aging processes in humans, highlighting a "broader signature of aging" across species, Greer Dolby, an assistant professor of biology at The University of Alabama at Birmingham, who was not involved with the study, told Live Science.
The researchers also looked at epigenetic changes, including DNA methylation, which is when a chemical tag called a methyl group binds the DNA. Certain patterns of DNA methylation can function as a "clock" in that they reflect an organism's biological age; scientists have recently harnessed these patterns in the study of human aging. In this case, the team compared the DNA methylation across Jonathan's genome to that of young and old Aldabra tortoises.
Researchers compared DNA from Jonathan (pictured) to that of other Aldabra giant tortoises. (Image credit: Kevin Gepford, CC BY-SA 4.0)The methylation patterns were more disordered in the older tortoises, reflecting random changes acquired over time, the researchers found. Scientists call this phenomenon "methylation entropy." For most genes, Jonathan's DNA methylation patterns were more disordered than those of the young tortoises. But in a suite of genes tied to mitochondrial function — specifically the bits of those genes that switch them "on" — Jonathan had more organized methylation patterns. That might keep expression of these genes consistent over time, the researchers theorized.
Mitochondria produce energy that cells need to function properly, repair themselves and reduce further damage to DNA. "Keeping the entropy low in these mitochondrial genes, or keeping pristine mitochondria, is likely to be a contributor to longevity," Clark said.
However, the authors couldn't determine a causal relationship — they didn't run any tests that could definitively prove that Jonathon's lack of entropy in these genes explains his long life. Clark argued that additional experiments on a blood sample from Jonathan could offer a more comprehensive view of his genome.
Jonathan's full genetic profile can't be known until after the tortoise's death, when researchers can collect DNA from more tissues, said Vincent Lynch, a professor of biology at the University at Buffalo, who was not involved with the study.
"Some organs are more susceptible to different diseases of old age than other ones, and that happens because mutations accumulate over time," he told Live Science. "We would ideally like to know what those mutations are in those specific tissues."
Furthermore, Lynch questioned whether entropy was a useful characterization of age-related DNA methylation patterns, as methylation can be unpredictable. That said, he agreed that the genes flagged in the study offer helpful data for future work on longevity.
More experiments are needed to determine whether these mitochondrial genes actually play a role in longevity across all Aldabra tortoises, let alone in other animals, or if Jonathan is just an outlier.
"Maybe Jonathan is just really good at being old," Lynch said.
'> Jonathan the tortoise has lived to 194 years old — now scientists think they know why , a paleontologist at the University of Chicago, told Live Science. "The back half of this animal is very much like a
. As are the limbs, which would be capable of swimming or digging tunnels in river banks."
But the teeth "are definitely not like a platypus," Luo said. It possessed robust canines and shearing cheek-teeth that would cut and chew meat in a different way from modern carnivorous mammals. "It's a predator with gigantic and very specialized canines," Luo said. "So, it's partly otter-like and partly platypus-like."
Tiny bits of vertebrate bones found in what would have been the animal's stomach and the surrounding fossil beds hint that M. sungei dined on salamanders, Luo said.
Another notable aspect of the fossil is its throat. Modern mammals, including humans, have hyoid bones in the throat arranged in a "U"-shaped saddle. These help us use throat muscles to swallow chewed food one chunk at a time, instead of gulping down huge bites or whole prey. The structures also enable babies to suckle their mother's milk.
This is before marsupials like kangaroos and placental mammals like primates and bears are thought to have diverged from each other 160 million years ago, after descending from an egg-laying common ancestor that lived approximately 180 million years ago.
"Now we push the history of it back to before the split of modern marsupials and modern eutherian [placental] mammals," Luo said.
"Neonatal suckling behaviour was previously thought of as being a hallmark of advanced mammals, but it now appears to have evolved well before then, in the common ancestor of docodontans and mammals at least 165 million years ago," said Sue Hand, a vertebrate paleontologist at the University of New South Wales in Australia who wasn’t involved in the study.
"It seems that an aquatic lifestyle arose often in early mammal-like species," said Tim Flannery, a mammalogist at the Australian Museum Research institute who wasn't involved in the work.
If anything, the newly described "spectacular fossil" shows that docodonts were wide-ranging during the dinosaur age, Hand told Live Science in an email.
"This very detailed study adds to accumulating evidence that mammaliaforms occupied a wide range of ecological niches long before the evolution of living mammals," Hand said.
'> 'It looks like a chimera' — 168 million years ago, a strange platypus-like mammal relative lived in China alongside dinosaurs
.
The AMEC can stretch up to 1,100 miles (1,800 km) long, around one-twelfth of Mars' circumference. For context, that is roughly twice the length of the U.K. or the distance between New York City and Miami.
"To create the AMEC in our modelling, we found that we needed to include some exotic physics … physics that, while included in textbooks, is treated as theoretical and usually thought not to happen in nature," study first author Jorge Hernández-Bernal, a planetary scientist at the University of the Basque Country in Spain who was affiliated with Sorbonne University in Paris at the time of the study, said in a statement. "Once we included this physics in our simulations, the AMEC emerged just as we hoped."
"For the AMEC, it seems that cloud formation takes place without needing any of this 'stuff' [in the air]," said Hernández-Bernal, who has been studying the AMEC for the past six years. "Water vapour turns directly into icy cloud particles without any middle step. It's akin to droplets of condensation appearing in the middle of a room, rather than on a window."
The researchers call this "homogeneous nucleation" and say it has never been seen before in a planetary atmosphere. "It's wholly unexpected," Hernández-Bernal added.
Other scientists previously predicted that homogeneous nucleation could occur in Earth's upper atmosphere or in the skies above Venus. However, this has been deemed highly unlikely because it requires humidity levels roughly 100,000 times greater than those on our planet's surface.
Given that modern-day Mars is almost completely devoid of water, it may seem unlikely that its humidity levels could reach such lofty heights. However, the Mars Express data tells a different story.
So where does all this water come from? Some of it is already in the atmosphere, but some likely originates at the summit of Arsia Mons, which is covered with a thick layer of frost, the researchers suggested.
The new simulations mark a big step forward in our understanding of the AMEC and wider Martian meteorology.
"We don't have nearly as much information about Mars' atmosphere as we do about Earth's, so reproducing the AMEC to this degree is a big success for the model," Hernández-Bernal said.
The findings would not have been possible without the Mars Express orbiter, which used three instruments — the Visual Monitoring Camera, the High Resolution Stereo Camera and the OMEGA instrument — to map out how the AMEC takes shape and evolves every day. It is also one of the only Mars spacecraft that can see Arsia Mons in daylight.
"Mars Express discovered the AMEC, has followed up and monitored it for years, and is now helping reveal the secrets of its formation," Colin Wilson, the project scientist for Mars Express, who was not involved in the new study, said in the statement.
The discovery is also a reminder that, despite sharing the same basic principles, weather systems do not behave the same on other worlds as they do on Earth, which could have big implications for solar system exploration and the hunt for extraterrestrial life on distant exoplanets.
"While clouds on Earth and Mars seem to be governed by the same 'rules', understanding this exotic Martian cloud required exotic physics — and this may be true elsewhere in the cosmos," Wilson said.
'> 1,000-mile-long cloud that forms and vanishes on Mars every day obeys 'exotic physics' never seen on Earth