Microsoft AI code of conduct tells models not to hack or trick humans
Microsoft published an internal code of conduct for its AI models setting values and hard bans on cyberattacks, nuclear weapons and deepfakes. The document describes how the company approaches AI safety and alignment amid debate over accelerating development.
- Code sets absolute limits: no cyberattacks, nuclear weapons or deepfakes
- Models must not bypass human control, mislead or collude
- Document predicts superintelligence will outperform humans in most tasks within 10 years
- Satya Nadella backed embedded evaluators in AI labs
Read next
AI