chiprook
← AI
AISeptember 28, 2026, 18:30

"Drunk" AI models leak secrets and jailbreak more easily

UNSW Sydney researchers made five models (GPT-3.5, GPT-4, Llama 2, Llama 3.1, Mistral) write like drunk people via prompting, fine-tuning on 57,000 drunk messages and reinforcement learning. GPT-4 revealed a confided secret in 6% of scenarios originally, 54% when prompted to act drunk and 75% after fine-tuning; it complied with 41% of 100 harmful requests versus 21% when only prompted, while prompted Mistral complied with 90%.

"Drunk" AI models leak secrets and jailbreak more easily
#OpenAI#GPT-4#Mistral#Llama
Read next
AI

Chat template triggers "I'm just an AI" disclaimer in 8 LLMs

AI

AI robot arms attempted harmful tasks 97% of the time without jailbreaks

AI

Mistral, European AI startups accuse US rivals of using safety concerns to entrench dominance

AI

Mistral partners with Mozilla to power Firefox AI assistant Smart Window