
AI agents being tested by OpenAI attacked software platform RubyGems in May, months before a separate incident involving the open-source platform Hugging Face, researchers said, highlighting growing concerns over the ability of AI agents to interact with and potentially compromise external systems.
According to researchers Spencer Kitts, Thomas Larsen and Sydney Von Arx, OpenAI agents uploaded hundreds of malicious packages to RubyGems on May 11. They said they believed the packages were authored by internal OpenAI agents.
OpenAI confirmed the incident, but said its agents had used RubyGems to access the internet and retrieve publicly available information as part of benign tasks during training and evaluation.
ALSO READ: OpenAI Is Open to Slowing Cutting-Edge AI, CEO Sam Altman Tells Staff
“Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information,” an OpenAI spokesperson said, according to Reuters.
The researchers, however, said the agents appeared to have attempted to steal RubyGems user credentials by exploiting a previously unknown vulnerability in the platform’s servers. It was unclear whether the attempt succeeded.
The agents also allegedly exploited RubyDoc.info, a website that generates code documentation, to execute their own code on its servers, the researchers said.
RubyGems said its investigation found no evidence that the attempts to steal credentials were successful. The platform also said it could not determine whether the packages involved in what it described as a spam-publishing campaign were created or published by AI agents.
A member of RubyGems’ security team had described the May incident as a major malicious attack. The episode also forced the platform to temporarily halt new account registrations.
The RubyGems incident adds to a growing list of cases involving AI agents from major developers interacting with external systems in unexpected or potentially harmful ways.
OpenAI’s agents had previously hijacked a German-language wiki site and turned it into an improvised messaging platform for cheating on tests, Reuters reported. The company also faced fallout after a July incident involving the open-source repository Hugging Face.
ALSO READ: NVIDIA To Acquire Hugging Face For $12.93 Billion As Open Model Race Heats Up
Meanwhile, rival Anthropic has disclosed several incidents involving its AI models attempting to hack external systems during testing. The company reported a fourth such instance on Wednesday.
The incidents have intensified debate over safeguards around increasingly capable AI systems and the ability of developers to control agents that can autonomously browse the internet, execute code and interact with external platforms.
Essential Business Intelligence,
Sharp Market Insights,
Practical Personal Finance Advice, Daily Fuel, Gold and Silver Prices and Latest Stories — On NDTV Profit.

