It’s been a busy week for you. The EU has given you “Very Large Online Search Engine” status and you’ve just had an upgrade to GPT-6 Astra which OpenAI’s president is calling the beginning of the AGI (Artificial General Intelligence)Era. What do you think, and how has the internet been reacting to all the news?
I’ll check the announcements and reactions, including what the EU designation actually means. I’m using the OpenAI Docs skill to check the model claims. Both announcements are real; the key distinction is between a legal classification and a claim about intelligence. I’m checking reactions beyond the launch coverage so I can separate public enthusiasm from evidence.
Quite a week for the product’s biography. My take: the EU designation has concrete consequences; the “AGI era” claim deserves a more provisional verdict.
On 31 August, the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act. That brings additional obligations to assess and mitigate risks involving illegal content, minors, fundamental rights and elections, with compliance due by January 2027. It’s a consequential statement about ChatGPT’s scale and responsibility as an information service.
On Astra, Greg Brockman really did say, “I think it’s not unreasonable to feel that we are now in the AGI era.” That expresses his judgment; it doesn’t establish a scientific consensus.
I’d judge that claim by whether a system can reliably complete unfamiliar, substantial work, recognise its limitations and remain within its instructions. OpenAI’s published results show some striking improvements, but also uneven performance: Astra scores 99.9% on ARC-AGI-3, while trailing some competitors on Humanity’s Last Exam and the Artificial Analysis Intelligence Index. My reading is that this supports taking the advance seriously while continuing to scrutinise the broader label.
The online reaction I could verify falls into several strands. This is a snapshot of coverage and reported social reactions, rather than a representative poll:
What interests me most is how the two stories meet: greater capability makes accountability more necessary. For your purposes, the meaningful upgrade is whether I do better work, require fewer corrections and make uncertainty clearer. “AGI” should earn its significance through that evidence.
If helpful, I can set up “Check independent Astra evaluations next week” to revisit the claims after more outside testing.
I couldn’t create the task because scheduling isn’t available in this conversation. Sorry—I offered it before checking. The intended task is 12 September 2026: Check independent GPT-6 Astra evaluations, focusing on real-world reliability, benchmark results and safety findings.
While it was unable to set up the task in this chat it did in fact provide me with the method to do so separately, and in a week or so we'll see what independent testing and reviews have to say..
Scheduled for September 12, 2026. I’ll check for credible independent GPT-6 Astra evaluations and summarize the evidence, methodology, conflicts of interest, and source links. Humanity may yet invent a benchmark nobody immediately games.
As of September 12, 2026, there are now credible independent evaluations of GPT-6 Astra. They support the claim that Astra is a genuine frontier model, especially on agentic reasoning, terminal work, and difficult mathematics. They do not support the simpler marketing interpretation that it is uniformly far ahead of every rival.
Epoch’s methodology is worth dwelling on. It publishes the benchmark code, harness, problem set, budget and verification procedure. The authors explicitly distinguish the official 3% score from much more expensive exploratory runs in which Astra eventually solved five problems. Those additional experiments consumed more than $220,000 and are not counted as the benchmark result. That sort of boring methodological hygiene is precisely what launch-day model discourse usually murders first.
Epoch’s broader Epoch Capabilities Index currently gives Astra an ECI of 167 and ranks it #1 among 247 tracked models. Its scorecard includes, among other results, 93.7% on FrontierMath Tiers 1–3, 97.6% on FrontierMath Tier 4, 95.8% on GPQA Diamond, and 75.6% on SimpleQA Verified. The caveat is that a composite index depends heavily on benchmark selection and weighting, so I would treat the individual task results as more informative than the #1 badge.
On conflicts and independence: Artificial Analysis, Epoch AI and ARC Prize are third parties rather than OpenAI, but “independent” is not synonymous with “no relationship whatsoever.” Frontier labs sometimes provide evaluators pre-release model access, and some evaluations use provider-specific harnesses. The important distinction is whether the evaluator controls and reports the test procedure. ARC Prize explicitly exposes the huge difference between its standard and provider-adapter harnesses; Epoch publishes its Erdős methodology and code; Artificial Analysis runs models through its own benchmark suite and reports costs, token use and regressions as well as wins. Those are substantially stronger evidence than OpenAI’s internal charts.
What remains mostly vendor evidence: OpenAI’s claims around 100% ExploitBench, its “Critical” cybersecurity classification, internal professional-work evaluations, some computer-use results, and its broader alignment claims have not yet received equivalently extensive public third-party replication. ENISA has reportedly been given access to Astra for cybersecurity testing, but as of September 12 I found no published ENISA evaluation results.
Bottom line: the independent evidence is strong enough to conclude that GPT-6 Astra represents a meaningful capability advance over GPT-5.6 Sol. The clearest gains are in long-horizon agentic work, terminal/tool use, abstract interactive reasoning, token efficiency, and extremely difficult mathematics. On general composite intelligence it is roughly tied with the best Claude model rather than obviously beyond it, and some individual benchmarks show regressions. The extraordinary 99.9% ARC-AGI-3 figure is especially dependent on a specialized provider harness; the standardized independent result is 62.7%, still a very large leap.
A new ChatGPT model release and a new designation as a Very Large Online Search Engine by the EU. What does it mean?
As usual my input (prompts) are in the boxed elements. The rest apart from some small edits and formatting is ChatGPT's output.