Simultaneous Outages Hit ChatGPT, Claude, and Gemini, Impacting Thousands of Users

0
4

Key Takeaways

  • On Thursday morning, OpenAI, Anthropic, and Google each reported service degradations affecting their flagship AI products—ChatGPT, Claude, and Gemini—within a roughly overlapping time window.
  • OpenAI acknowledged “elevated errors across ChatGPT and Codex,” noting the issue persisted for about 20 minutes before a status‑page update was posted.
  • Anthropic identified elevated error rates across multiple Claude models (Mythos/Fable 5.1, Mythos/Fable 5, Opus 5, Opus 4.8, Opus 4.6), traced the root cause, and said it was actively working on a fix.
  • Google’s Gemini API experienced partial degradation for some developer and API users, particularly affecting recently created keys and those used via OpenAI‑compatible libraries.
  • All three companies stated they were investigating the problems and had reached out to The Post for comment, underscoring the coordinated impact on the AI ecosystem.

Overview of the Service Disruption
The simultaneous hiccups across three of the industry’s most prominent AI platforms raised immediate concerns about the reliability of large‑scale generative models. Users attempting to interact with ChatGPT on web and mobile interfaces reported intermittent “error generating response” messages, while developers relying on the Codex API observed failed code‑completion requests. Anthropic’s status page flashed alerts for several Claude variants, and Google’s AI Studio dashboard highlighted problems serving newly minted Gemini API keys. The timing—roughly between 9:00 a.m. and 11:00 a.m. Eastern Time—suggested a possible common underlying factor, though each company stressed that its investigation was independent at this stage.


OpenAI’s Statement on ChatGPT and Codex
OpenAI’s incident report, captured in a status‑page screenshot reviewed by The Post, said the company was “investigating elevated errors across ChatGPT and Codex.” The note added that the problem had been ongoing for approximately 20 minutes at the time of the screenshot. A spokesperson later clarified, “We are monitoring the situation closely and have deployed mitigations to reduce error rates; further updates will follow as we isolate the root cause.” The brief outage forced many casual users to refresh their chats repeatedly, while developers building on Codex reported delayed or missing completions in their integrated development environments.


Anthropic’s Claude Models Affected
Anthropic’s status page listed a suite of impacted models: “Mythos/Fable 5.1, Mythos/Fable 5, Opus 5, Opus 4.8 and Opus 4.6.” The company first announced the issue just before 9:30 a.m. ET, stating it had “identified the cause and was working on a fix.” By 10:49 a.m. ET, an update read, “We are continuing to work on a fix for this issue.” The quoted language indicates Anthropic had already narrowed down the problem—likely tied to a shared inference service or a recent deployment—allowing engineers to target remediation efforts more precisely than a blind rollback.


Details on the Specific Claude Variants
Among the affected Claude models, the Mythos/Fable series represents Anthropic’s latest line of multimodal, high‑capacity transformers, while the Opus line includes earlier but still widely used versions. The fact that both families experienced errors points to a potential infrastructure layer—such as load‑balancing, GPU allocation, or a shared token‑generation service—rather than a model‑specific bug. Anthropic’s rapid identification of the cause suggests robust internal observability tooling, which enabled the team to pinpoint the fault within roughly an hour of the initial report.


Google’s Gemini API Challenges
Google’s AI Studio status page disclosed that “the Gemini API was having problems serving recently created API keys, including keys used through OpenAI‑compatible libraries.” The note continued, “The company is investigating.” This wording hints at a possible issue with key‑validation or quota‑enforcement systems that differentiate newly issued keys from older, established ones. Developers who had freshly registered keys for experimental projects reported authentication failures, whereas long‑standing keys appeared to function normally, limiting the blast radius but still causing noticeable disruption for early‑adopter users.


Temporal Overlap and Potential Common Factors
The near‑synchronous onset of the degradations—spanning roughly a 90‑minute window from early morning to late morning ET—has prompted speculation about shared external dependencies. Cloud‑provider networking events, DNS fluctuations, or a widespread software‑library update (perhaps a common open‑source inference engine) could theoretically affect multiple vendors simultaneously. However, each firm has emphasized that its investigations are independent, and no conclusive evidence has yet emerged linking the incidents to a single root cause.


Industry Reaction and User Impact
Social‑media threads and developer forums filled with complaints about lost productivity, failed automated workflows, and interrupted chatbot interactions. A freelance programmer tweeted, “Just spent 20 minutes trying to get Codex to finish a function—kept getting ‘internal error’ messages. Hope they fix it fast.” Enterprises that rely on these APIs for customer‑support bots or content‑generation pipelines reported temporary degradation in service‑level agreements, prompting some to activate fallback mechanisms or switch to alternative models during the outage window.


Statements from the Companies (Quoted)
When approached for comment, OpenAI replied, “We are aware of the issue and are working to restore normal service as quickly as possible.” Anthropic’s representative said, “We have identified the underlying cause and are deploying a fix; updates will follow shortly.” Google’s response echoed a similar tone: “Our team is investigating the Gemini API anomaly and will provide a status update soon.” These concise, standard acknowledgments reflect the companies’ desire to maintain transparency while avoiding premature speculation that could unsettle investors or users.


What the Incident Means for AI Reliability
The episode serves as a reminder that even the most advanced, widely deployed AI systems remain dependent on complex, layered infrastructures—hardware accelerators, networking, orchestration platforms, and monitoring tools. A fault in any of these layers can cascade, affecting multiple models and services despite their architectural differences. For end users, the takeaway is the value of designing applications with graceful degradation: retry logic, alternative model endpoints, and clear error messaging can mitigate the impact of such provider‑side outages.


Looking Forward: Mitigation and Lessons Learned
In the wake of the disruption, each company is likely to conduct post‑mortem analyses, examining deployment pipelines, change‑management procedures, and fault‑tolerance mechanisms. Potential improvements may include staggered rollouts, more granular health‑checks for newly issued API keys, and enhanced cross‑provider communication during widespread incidents. As the AI market continues to mature, providers will be judged not only on the raw capabilities of their models but also on the steadiness of the platforms that deliver those capabilities to millions of users and developers worldwide.

https://nypost.com/2026/09/03/business/chatgpt-claude-and-gemini-are-all-down-as-thousands-of-users-experience-outages/

SignUpSignUp form

LEAVE A REPLY

Please enter your comment!
Please enter your name here