TNB TopicalNewsBeats

Crypto, chains and markets — the beat, as it happens

LIVE

Decrypt

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Photo: Decrypt

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Originally published by Decrypt. Read the full story at the source →

More stories