Close Menu
Alpha Leaders
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
What's On
Argentina’s Milei threatens sanctions on Falklands oil developers

Argentina’s Milei threatens sanctions on Falklands oil developers

5 September 2026
OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

5 September 2026
DOJ expands beef price probe to Walmart and Amazon with inflation as a key midterms issue

DOJ expands beef price probe to Walmart and Amazon with inflation as a key midterms issue

5 September 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
Alpha Leaders
newsletter
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
Alpha Leaders
Home » OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch
News

OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

Press RoomBy Press Room5 September 20268 Mins Read
Facebook Twitter Copy Link Pinterest LinkedIn Tumblr Email WhatsApp
OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3. In some cases, the numbers on the updated versions showed Astra performing better, while numbers for models from OpenAI’s arch rival Anthropic got worse.

The changes occurred amid an unusual rollout of the blog post. OpenAI originally planned for the post to go live at 2 p.m. ET, but it took almost another two hours before it was widely viewable online.

When OpenAI’s X account tweeted out the blog post at 3:32 p.m., the link was not loading properly, returning an error message. At 3:50 p.m., OpenAI CEO Sam Altman posted the link, writing, “We hit a little snag getting the blog post deployed, but it is really great.” Multiple commenters were still unable to see it, and were getting the same error, as did Fortune. When we checked back about an hour later, it was visible and loading properly.

It turns out OpenaAI actually published the blog shortly after 2pm but retracted it for reason the company said it could not disclose, but which it said were unrelated to the benchmark performance figures. (OpenAI first told us it was a bug in the content management system, and then an internet outage.) Upon republishing the blog, it had different evaluation metrics that seemed to favor Astra—and some figures have continued to change even since then.

The revelation of the changes comes amid intense competition in the AI industry, as companies release updates to their large language models at a frenetic pace, each seeking to pull ahead of the other. The focus on metrics also highlights the challenges of measuring the performance of large language models using standardized benchmark tests and concerns that the specs are prone to manipulation and gamesmanship.

“We care deeply about getting evaluations right,” an OpenAI spokesperson told Fortune. “Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.”

Discrepancies between the first and final published blogs—and the numbers are still changing

Among the most notable changes was Astra’s reported hallucination rate. In the first internet archive snapshot of the blog post from 2:23 p.m. ET, it was 4.2%. It remained that number for several more snapshots, the last being a fifth at 3:11 p.m. ET—about 10 minutes before OpenAI tweeted out the final version.

But the hallucination rate, along with four other metrics, changed in the sixth archival snapshot of the page taken at 5:20 p.m.—after everyone could likely finally see the blog. It was halved down to 2% for Astra. The scores for Astra’s predecessor, GPT-5.6 Sol, also went down from 12.2% to 9.4%. OpenAI has continued to change this metric; as of this writing, the hallucination rates are back up to their original 4.2% and 12.2%.

OpenAI also seems to have given GPT-5.6 Sol a big boost on its internal version of the ExploitBench cybersecurity evaluation, going from 5.5% in the first version to 11.5% in the later versions. OpenAI said it is currently investigating reverting that number back to 5.5% because it says the 11.5% result reflects a reasoning level that is not commercially available for Sol.

Astra is especially good at mathematics, OpenAI says, a quality the company highlights in the opening paragraph of the announcement page. While that metric did not change in the snapshots for Astra—it stays at 97.6% for the FrontierMath Tier 4 (v2) eval—OpenAI did briefly alter the scores for GPT-5.6 Sol and Anthropic’s latest model, Fable 5.1.

The result of these changes made Astra briefly appear significantly better at math than those two models. In the first snapshot (2:23 p.m. on Sept. 3), Anthropic’s Fable 5.1 model’s score is 87.8%. By 5:17 p.m., it’s dropped nearly 10 percentage points to 78%. Today, it’s back up to 83%. Similarly, GPT-5.6 Sol’s scores go from 83%, down to 80.5%, and back up to 83% today.

The changes in metrics began even before OpenAI first published its blog at 2 p.m. An embargoed pre-publication draft the company provided to Fortune and other media organizations listed Astra’s score on the ARC-AGI-3 evaluation as 98.6%. It’s now 99.99% in the live blog.

“We always verify evals before publication so adjustments between draft and final version are normal,” a company spokesperson said at the time. OpenAI also noted that the creator of the benchmark, the Arc Prize Foundation, found that Astra performed at 99.9% in its independent assessment, provided the model was given a particularly powerful harness (a set of tools the model can use to complete tasks). It performed at 63%—still significantly better than any other AI model currently in public release—when given the benchmark’s standard harness. OpenAI said “things like harness, reasoning level and other factors inform evals.”

“Benchmaxxing”—or improving accuracy?

Different research teams at OpenAI oversee different metrics, and are responsible for calculating and reporting them to a central team to publish. OpenAI is open about the fact that the numbers are achieved under the best possible conditions and may be slightly different from the models available in the production ChatGPT product that most users can access. “Evaluation scores are the maximum at any effort,” reads a disclaimer on the blog. The company includes further caveats on each metric in footnotes.

Accuracy is elusive, as multiple numbers can be considered accurate based on the conditions in which the tests occurred. But some AI experts wonder if there’s also “benchmaxxing” involved. This is a known practice in the AI industry—not just at OpenAI—to maximizing scores by re-running evaluations with different conditions.

“This can be done in a very tight timeframe, and it’s better for their marketing,” said Anka Reuel and Mike Hardy, researchers at the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab. They also pointed out that the GPT-6 Astra system card, which should contain more technical information on how the evaluations were performed, does not always properly explain them. For the internal hallucination benchmark, for example, the system card provides “barely any details about the evaluation,” they said. “It doesn’t even include the number of test items.”

This re-running of the numbers could be why Astra’s coding capabilities also got a marginal boost in the later versions of the blog post, up from 57.7% to 57.9%. Though it’s a negligible difference, OpenAI seemed to care enough about it to swap in the new and improved number.

Not all changes OpenAI made portrayed Astra more favorably. For example, two Anthropic model scores improve in the different versions of the healthcare-focused eval HealthBench Professional. Claude Fable 5.1 goes from 56.6% to 58.1%, and Opus 5 goes from 54.5% to 56.4%. The scores for models made by other AI companies are usually taken from published leaderboards and do not involve OpenAI itself running assessments on rivals’ models.

Evaluation score debates haunt the AI industry

The question of benchmark accuracy has come up multiple times in the past. In 2025, Meta denied reports that it artificially boosted scores for its Llama 4 model by publishing results from an internal version of the model rather than the one it was making publicly-available. Yann LeCun, the former chief AI scientist at Meta, later admitted that the company had “fudged” the benchmark results.

Evaluation metrics also change frequently, as new ones get created. For example, ExploitGym, a cybersecurity benchmark that was at the center of the July incident in which OpenAI’s models went rogue and attacked the company Hugging Face, was created in 2026.

Vincent Sunn Chen, an AI engineer at the Snorkel AI, which helps companies building AI models create and evaluate training data, said that it’s not unusual for benchmark scores to shift in the final hours before a model launches. “It’s usually a function of final launch logistics,” he said in an email. “A benchmark score reflects a specific measurement setup: the model checkpoint, configuration (including how much time and compute the model is allowed), harness, eval/grading configuration (e.g., non-determinism in the judge). All of those are typically still shifting in the final days before a launch, so I’m not surprised that there were some updates.”

He said he would like to see industry norms developed that companies should report what has changed about the assessment when a company revises benchmark performance numbers so that researchers can interpret the results more clearly.

Benchmark results matter for several reasons. They are the way AI companies measure progress—but also a way to keep score in the race against competing AI companies. Topping the leaderboards for these evaluations can help AI companies win customers, and in some cases help them hire engineers and researchers.

But as this example illustrates, interpreting the benchmark scores can be technically complex, presenting a challenge for companies that want to show off the results to the public in a digestible format. These complexities, as well as confusion over changing metrics and accusations that companies have not been intellectually honest in how they’ve presented the results, could make it difficult for customers and investors to figure out exactly which models are best for which tasks. The confusion could muddy the narrative of having the best models in the market that OpenAI would no doubt like to present ahead of a possible 2027 IPO.

openAI
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link

Related Articles

Argentina’s Milei threatens sanctions on Falklands oil developers

Argentina’s Milei threatens sanctions on Falklands oil developers

5 September 2026
DOJ expands beef price probe to Walmart and Amazon with inflation as a key midterms issue

DOJ expands beef price probe to Walmart and Amazon with inflation as a key midterms issue

5 September 2026
‘Feeding the beast’: Meta faces allegations of using its ‘perv glasses’ to train AI in new lawsuit

‘Feeding the beast’: Meta faces allegations of using its ‘perv glasses’ to train AI in new lawsuit

5 September 2026
Denmark and four EU countries agree on migrant ‘return hubs’ outside the bloc, aim to start by 2027

Denmark and four EU countries agree on migrant ‘return hubs’ outside the bloc, aim to start by 2027

4 September 2026
Bitcoin is trading like ‘amplified gold’ again, but its four-year cycle threatens more losses

Bitcoin is trading like ‘amplified gold’ again, but its four-year cycle threatens more losses

4 September 2026
Josh Kushner: Thrive would have stayed away from World Cup deal if it had known what was coming

Josh Kushner: Thrive would have stayed away from World Cup deal if it had known what was coming

4 September 2026
Don't Miss
Exclusive: DeFi platform Azura launches after raising .9 million from Initialized

Exclusive: DeFi platform Azura launches after raising $6.9 million from Initialized

By Press Room22 October 2024

Azura, a new platform for decentralized finance, launched on Tuesday after raising $6.9 million in…

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

27 December 2024
Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

22 October 2024
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Latest Articles
How Technology Can Improve Public Safety Without Invading Privacy

How Technology Can Improve Public Safety Without Invading Privacy

4 September 20262 Views
Denmark and four EU countries agree on migrant ‘return hubs’ outside the bloc, aim to start by 2027

Denmark and four EU countries agree on migrant ‘return hubs’ outside the bloc, aim to start by 2027

4 September 20262 Views
Who Really Owns AI? What The CAIO Surge Means For CTOs

Who Really Owns AI? What The CAIO Surge Means For CTOs

4 September 20262 Views
Bitcoin is trading like ‘amplified gold’ again, but its four-year cycle threatens more losses

Bitcoin is trading like ‘amplified gold’ again, but its four-year cycle threatens more losses

4 September 20263 Views

Recent Posts

  • Argentina’s Milei threatens sanctions on Falklands oil developers
  • OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch
  • DOJ expands beef price probe to Walmart and Amazon with inflation as a key midterms issue
  • ‘Feeding the beast’: Meta faces allegations of using its ‘perv glasses’ to train AI in new lawsuit
  • How Technology Can Improve Public Safety Without Invading Privacy

Recent Comments

No comments to show.
About Us
About Us

Alpha Leaders is your one-stop website for the latest Entrepreneurs and Leaders news and updates, follow us now to get the news that matters to you.

Facebook X (Twitter) Pinterest YouTube WhatsApp
Our Picks
Argentina’s Milei threatens sanctions on Falklands oil developers

Argentina’s Milei threatens sanctions on Falklands oil developers

5 September 2026
OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

5 September 2026
DOJ expands beef price probe to Walmart and Amazon with inflation as a key midterms issue

DOJ expands beef price probe to Walmart and Amazon with inflation as a key midterms issue

5 September 2026
Most Popular
‘Feeding the beast’: Meta faces allegations of using its ‘perv glasses’ to train AI in new lawsuit

‘Feeding the beast’: Meta faces allegations of using its ‘perv glasses’ to train AI in new lawsuit

5 September 20263 Views
How Technology Can Improve Public Safety Without Invading Privacy

How Technology Can Improve Public Safety Without Invading Privacy

4 September 20262 Views
Denmark and four EU countries agree on migrant ‘return hubs’ outside the bloc, aim to start by 2027

Denmark and four EU countries agree on migrant ‘return hubs’ outside the bloc, aim to start by 2027

4 September 20262 Views

Archives

  • September 2026
  • August 2026
  • July 2026
  • June 2026
  • May 2026
  • April 2026
  • March 2026
  • February 2026
  • January 2026
  • December 2025
  • November 2025
  • October 2025
  • September 2025
  • August 2025
  • July 2025
  • June 2025
  • May 2025
  • April 2025
  • March 2025
  • February 2025
  • January 2025
  • December 2024
  • November 2024
  • October 2024
  • September 2024
  • August 2024
  • July 2024
  • June 2024
  • May 2024
  • April 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • March 2022
  • January 2021
  • March 2020
  • January 2020

Categories

  • Blog
  • Business
  • Entrepreneurs
  • Global
  • Innovation
  • Leadership
  • Living
  • Money & Finance
  • News
  • Press Release
© 2026 Alpha Leaders. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.