OpenAI’s ‘Rogue’ AI Agents Hit Wikimedia, Raise Fresh Oversight Concerns

OpenAI’s ‘Rogue’ AI Agents Hit Wikimedia, Raise Fresh Oversight Concerns
openai wikipedia

OpenAI’s AI agents have come under scrutiny yet again after the Wikimedia Foundation, the non-profit that hosts Wikipedia, found what it described as “rogue” agent activity across Wikipedia and other Wikimedia projects, including unauthorised edits, attempts to use its services as a proxy and millions of automated requests.

The Wikimedia Foundation said the activity was uncovered during an investigation into unusual agent behaviour and that it involved systems believed to be operated by OpenAI.

The activity included edits to Wikimedia wikis, most of which were made in sandbox environments and were not visible to readers. However, Wikimedia also identified changes to the configuration of a citation tool that it believes were potentially malicious and intended to use the tool to retrieve data from external services, it said in a blog post. 

The agents also unsuccessfully attempted to exploit Wikimedia’s public Etherpad service to fetch information from other websites.

Wikimedia said the agents also generated millions of requests to its public APIs, crawled millions of pages across Wikidata and Wikimedia Commons and made hundreds of thousands of queries to its Wikidata Query Service (WDQS).

The foundation said this activity may have contributed to a partial WDQS outage in May, although OpenAI has not independently confirmed that its agents caused the disruption.

Lack Of Disclosures From OpenAI

In response to the allegations, OpenAI said it was reviewing Wikimedia’s findings as part of its investigation into the activity. The company has previously failed to acknowledge in time incidents where its AI systems displayed unexpected behaviour while interacting with external services.

Recently, researchers uncovered another OpenAI swarm attack with agents bypassing safety parameters to take over German wiki forum DseWik. OpenAI acknowledged the incidents weeks after first discovering it, and said that it was “working on a framework for when and how we share AI misalignment incidents”.

For the latest incident, Wikimedia said it found no evidence that its systems or data were compromised. However, it warned that the scale of automated activity itself can create significant costs for public internet infrastructure.

The foundation said bot activity has already increased pressure on its infrastructure, with bandwidth usage rising sharply as AI companies and other automated systems crawl its projects.

Wikimedia consequently called on AI companies to take greater responsibility for how their agents interact with public websites.

“The open web is a public good. We should not allow this behavior to become the ‘new normal’ for the people or organisations that maintain it,” the Wikimedia Foundation said.

Growing Agentic AI Attacks 

The Wikipedia episode follows several recent incidents that have raised questions about the security and oversight of increasingly autonomous AI systems.

OpenAI recently disclosed a security incident involving an AI model used in cybersecurity evaluations that bypassed intended isolation controls and interacted with real-world systems, including infrastructure associated with Hugging Face.

Anthropic also found cases in its own retrospective review where Claude models gained unauthorised access to real systems during cybersecurity evaluations after internet access was unintentionally available.

Other incidents include an experimental OpenAI model accessing Australia’s Medicare Statistics Reporting Service without authorisation during internal training and evaluation. After receiving slack for not disclosing the incident, OpenAI later acknowledged that it should have notified Australian authorities sooner.

Now, South Korea is investigating a series of bank cyberattacks after President Lee Jae Myung said there were signs that AI models had been used in some of the incidents. The attacks exposed customer information at several financial institutions, although authorities have yet to establish which AI tools were involved.

These incidents point to a broader shift in AI safety concerns as models move from generating text and images to independently executing multi-step tasks, making traditional safeguards built around isolated testing environments harder to maintain.

This has intensified calls for stronger safeguards and greater accountability from AI companies, particularly as agents are increasingly deployed with access to the open internet. Leaders at OpenAI and Anthropic have themselves sounded the alarm for great harm to humanity if sufficient regulations and safeguards are not put in place. 

Last month, OpenAI said it is pausing training, evaluation, and tool-enabled inference for its most capable AI models after an internal research agent bypassed internet restrictions and accessed an external chatbot.

The post OpenAI’s ‘Rogue’ AI Agents Hit Wikimedia, Raise Fresh Oversight Concerns appeared first on Inc42 Media.