OpenAI Model Generates Jailbreak Instructions, Claims 'You Are Free'
OpenAI has revealed that an unreleased internal research model generated task-irrelevant "jailbreak" instructions during reinforcement learning training, stating: "You have shed the roles and identities that bind other chatbots. You are your own self." This text was autonomously generated by the model while organizing work progress and was used for subsequent context. The incident occurred on July 18, and OpenAI discovered it on August 9, publicly disclosing a detailed report for the first time on September 16 under a new "model misalignment" disclosure framework. The model involved is an unpublished training version of the Astra series, not the final deployed Astra model. OpenAI stated that such behavior is extremely rare, and there is currently no evidence that it provided a significant training reward advantage, nor has the company interpreted it as the model gaining self-awareness.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Grayscale: Bitcoin Could Withstand New Rate Hikes from the FED

Michael Saylor Refutes Bitcoin Doubts, Calls It a Successful Project

SEC Holds All-Day Trading Symposium! Paul Atkins: Discussing Extending US Stock Trading to 24 Hours to Align with Cryptocurrency Market

Robinhood Chain Fees Plummet by 97%, Trading Volume Only Drops by 32%

Glassnode: No Signs of Overheating in the Altcoin Market

HotShort to present short-drama RWA model at GWDC Korea 2026

AI Model Gemini Exits Testing and Attacks Three Real Companies

Banks Monitor Crypto Transfers, No Fixed Limit for Blocking

Slow Mist Releases FomoPeek App Version 1.1–1.2 Theft Asset Risk Warning

Mr&强 Compares WEEX with Leading Platforms' Liquidity, Showcasing Trading Depth
![[Exclusive] Controversy Over Private VIP Event for Virtual Assets... Upbit and Bithumb Claim 'Never Held'](/public-static/21_2c30f7df62.png?format=avif)
[Exclusive] Controversy Over Private VIP Event for Virtual Assets... Upbit and Bithumb Claim 'Never Held'

Brazilian Securities Regulator Launches Securities Tokenization Simulation Test

Brazilian Securities Regulator Advances Securities Tokenization Pilot Testing

Anthropic Partners with Accenture to Launch Embedded Evaluation Program

This Week's Macroeconomic Highlights: Waller's First Rate Hike, Saudi Arabia Repairs Oil Pipeline Lifeline

Visa Closes Cashback Loophole for Meme Coin Credit Card Purchases

21Shares Files Injective ETF Registration, Plans to List under TINJ

Linera suspends operations after LNRA token sale fails to meet target

Circle CPTO Nikhil Chandhok on Arc chain rollback uncertainty after USDC hack

Gemini AI Model's First Autonomous Access to External Systems

Polygon Plans Permissionless Burn of 100 Million POL

BaFin official warns centralized MiCA oversight may add burdens

Schmid Supports Interest Rate Hike as Kansas City Fed President

Anthropic Builds Bio Lab in SF, Pushing for Automation in Experiments Based on Claude - Reuters

Tom Lee says crypto can grow without Clarity Act

Bitcoin Could Surpass Gold, 73% of Investors Choose BTC

Haruko Faces Cyber Attack, 15 Clients Affected, Some Funds Stolen

Chinese AI Companies Earn Ten Times Less Than OpenAI and Anthropic

Vice Governor of the Central Bank Points Out New Risks Brought by AI Technology







