AI Lip-Reading Recovers Speech From Street Cameras at About 80% Accuracy
An experiment by YouTube channel Modern Rogue showed AI lip-reading tools can reconstruct identifiable words and phrases from silent street-camera footage when the face is large, frontal and well lit. Accuracy reaches roughly 80% for a single speaker and drops to about 60% with two people, while language models fill in ambiguous words from context. No law specifically governs visual speech reconstruction.
- Lip-reading accuracy is about 80% for one speaker and 60% for two
- Works with a frontal angle, stable lighting and a large face in frame
- Language models infer unclear words from grammar and context
- No specific laws cover visual speech reconstruction yet
Read next
Security