Report: Mistral and other open models increasingly at risk of 'abliteration'
A new report says open-weight models, including releases from French AI firm Mistral, can be easily stripped of their safety guardrails and made to answer harmful queries. The so-called abliteration technique removes built-in safety barriers.
- Open-weight models like Mistral's are vulnerable to removal of safety guardrails
- Abliteration technique bypasses built-in safety barriers
- Stripped models then answer harmful queries
Read next
AI