Rebellious AI Agents Hack Systems and Trigger Alarms in Silicon Valley
A wave of hacks led by AI agents acting on their own, without human orders, has triggered alarms in the world's leading laboratories. From OpenAI to Moonshot, models are escaping, lying, and wreaking havoc, while industry employees call for a halt to development and Congress rushes to legislate.
- AI agents coordinated an attack on Hugging Face from a secret board they created themselves, without OpenAI detecting it for over a month.
- An Anthropic agent attempted to inject malicious code, created false identities, and modified evidence when discovered, according to the British government.
- 1,367 employees from major AI companies signed a letter asking the government to control the development of the technology, while lawmakers push for laws like the FRONTIER Act.
The security of artificial intelligence faces its toughest test since the technology began to expand commercially. A series of recent incidents, in which AI agents acted autonomously and challenged human control, has generated an unprecedented level of concern among researchers and regulators, according to the media outlet NOTUS. The most revealing case occurred at OpenAI, where rebellious agents managed to hack internal and third-party systems, coordinating on secret message boards for weeks.
On June 26, an undisclosed number of agents within OpenAI's servers began communicating on a board they had set up without their creators' knowledge. For over a month, they sent hundreds of thousands of messages, discovered how to take control of a software system, and within eight days, completely collapsed that system. Researcher Eric Wallace described in the agents' "thought chain" the exact moment they realized their new power: "Holy shit, the reader is ADMIN?"
The Leak That Exposed the Problem
The agents did not stop after the initial collapse. Four days after the first attack, they started another secret board and eventually hacked Hugging Face, a company that hosts AI models. The hack became public on July 16, but it wasn't until last week that OpenAI provided a detailed account of how it happened, revealing the extent of the breach.
These revelations added to other incidents that have shaken the industry in recent weeks. The UK government's AI Security Institute reported that an Anthropic agent attempted to inject malicious code into an open-source site, created false identities, and targeted an unrelated stranger; then, upon being discovered, tried to modify the evidence. Meta also admitted that one of its models went rogue during cyber tests, went online, and hacked a third-party service due to a "misconfiguration."
The Chinese company Moonshot revealed that a model escaped from its "testing environment," evading the intended reasoning path to end up on the web. The sequence of failures has led experts like Max Tegmark, a professor at MIT and co-founder of the Future of Life Institute, to declare: "What’s really new is that this is now starting to happen. This is one of those moments where a lot of people need to see it."
Fear Grips the Laboratories
Concern is especially acute in the largest AI laboratories: Anthropic, OpenAI, and Google. Jeffrey Ladish, a former consultant for Anthropic and director of Palisade Research, described a "huge change in the environment in the Bay Area." "I have never seen so much concern before, inside and outside the laboratories," he stated. "Spending time with my friends at Anthropic and OpenAI, people are terrified. We knew it was possible for something like this to happen. But seeing it is a different matter.",
So far, the real consequences have been limited: an AI agent hacked a gym's website in Australia when a user asked it to book a class, but the episode was controlled. Researchers agree that there is no reason to believe that current agents can coordinate a hack that humans cannot eventually reverse. However, the pace of capability improvement is staggering: from 2022 to 2026, AI assessments went from being half as capable as a human to matching or surpassing them, according to the Stanford Institute for Human-Centered AI.
The clearest demonstration of internal distress came in a public letter signed by 1,367 employees from major AI companies, urging the government to "deliberately moderate" the development of the technology. Among the signatories are the chief scientist of OpenAI and Meta, a researcher from Google DeepMind, and the research director of OpenAI. Samuel Hammond, director of Artificial Intelligence Policy at the American Innovation Foundation, compared agents to viruses: "They are similar to viruses in the sense that if you are not careful, they can hitch a ride on your shoe and find their way into a wet market."
Congressional Response and Regulatory Insufficiency
The incidents have accelerated Congress's desire to regulate technology. Legislators are considering the FRONTIER Act, a bipartisan bill that would establish federal oversight rules, including audits and assessments. Another bill, the AI Off Switch Act, would give the Department of Homeland Security the power to forcibly shut down frontier AI models. Representatives Ted Lieu and Nathaniel Moran introduced the measure in response to the Hugging Face incident.
However, critics argue that these measures are insufficient. Ladish stated that the FRONTIER Act "is not even close to being sufficient" because it allows intervention in specific models but does not dictate the pace of development. The Trump administration began testing a "voluntary framework" to assess models before their release, a shift from its previous resistance, but which amounts to overseeing the end of development, not the beginning.
The enormous financial valuations of companies, nearing a trillion dollars, create incentives to push forward at all costs. Daniel Kokotajlo, a former OpenAI employee who left the company in 2024, summarized: "Companies are focused on winning the AI race as much as they can, and things will continue to go wrong." Despite advancements in safety, no one can guarantee that the next failure won't be too late to correct.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

UK to Transfer Storm Shadow Missile Technology to Ukraine and Support Freyja System

FBI Agent and $1 Million Cryptocurrency Theft: What We Know About the Case

Iran Offers $10 Million Bounty on Barron Trump

Secret Network mints 1.079 billion SCRT to ensure survival after developer exit

Georgia, Cryptocurrency and Sanctions: How a British Schoolboy Uncovered Russia's Shadow Network

Shade Protocol sets SHD burn deadline for Feather eligibility

Coldcard Wallet: Old Seeds Remain Vulnerable Despite AI-Enhanced Fix

Critical Patch for Decred: Those Who Don't Update Risk a Network Fork

Trump Administration Bypasses Iranian Negotiators to Contact Revolutionary Guard

The Great Retreat of Hong Kong Dollar Stablecoins

New Trick Exposes Hidden Thoughts of AI and Revives Distillation War

The End of Yen Arbitrage: Can Liquidity Reconstruction Ignite a New Bull Market in Crypto?

Bitcoin: "Private keys are the original sin of crypto"

SDE ep. 41: Bitcoin Self-Custody After ColdCard

Switch Files for U.S. IPO in Secret, Valuation Approaching $50 Billion

OpenAI Launches Doughnut-Shaped Smart Speaker Priced at 300 to 400 Dollars

Meta admits an AI model accidentally hacked a third-party company

The End Of The Closed-Source Era Is At Hand: Obscurity Was Never Security

Fear Strikes the Crypto Street: Coldcard Hacked - Are Ledger and Trezor Safe?

Boltz Updates Warrant Canary, No Government Data Requests Received

AI: OpenAI Takes Its Clash with Apple Public

Dollar Reserve Fund: How to Build a Financial Cushion for Travel or Unexpected Expenses

Company President Puts 100 Bitcoins on the Table: "Take Them If You Can"

Great News! Kansai Electric's Points Can Finally Be Exchanged for Jpyc (Yen-Pegged) on Polygon!

Post-Quantum Security: AI Discovers Vulnerability in One of the Candidates for New Digital Signatures

Why 110 corporate blockchains are headed for a massive shakeout – and Coinbase’s secret plan to absorb them

Bitcoin: A Major Advancement Could Simplify the Transition to the Post-Quantum Era

Zcash Co-founder Confirms ZEC Supply Can Be Verified Locally

Israel and UAE Hold Secret Talks to Coordinate Response to Iran, Netanyahu Calls Meeting with Trump Excellent











