The concept explainer
Why Is It Harder to Understand English When You Cannot See the Speaker?
Your ears are not doing all the work in face-to-face listening. See how lip movements, facial information, gestures, noise, and language experience combine during English speech perception.
An idea, brought into focus.
09 sources, open to exploreIn this article 15 sections
Audiovisual speech perception is the process of using what you hear together with what you see to understand speech. In a face-to-face conversation, the acoustic signal from a voice is only one source of information. Movements of the lips, jaw and tongue can provide clues about speech sounds, while facial expression and hand gestures can add timing, emphasis and meaning.
The useful mental model is not ears + optional decoration.
It is several partly overlapping signals that the brain can combine when they are informative.
That is why a phone call, podcast, dark room, frozen video feed or speaker facing away can sometimes make familiar English unexpectedly harder. The language has not changed. One source of evidence has disappeared.
Listening has never been purely auditory
Think about a noisy café.
Someone across the table says:
Did you say fifteen or fifty?
The sound reaching your ears may be imperfect. Music is playing. Cups are hitting tables. Another conversation overlaps with yours.
But you are not limited to the waveform.
You can see when the speaker begins talking. You can see the rhythm of the mouth. Some consonants have visible articulatory information. You can watch the speaker's face and follow a gesture that narrows the intended meaning.
None of these cues gives you a transcript of the sentence.
Together, however, they can reduce uncertainty.
Research on audiovisual speech has documented this visual contribution for decades. Reviews describe a robust visual gain in speech intelligibility when listeners can see a talker, particularly when the auditory signal is degraded (Peelle & Sommers, 2015; Irwin & DiBlasi, 2017).
The important word is combine.
Your brain does not necessarily finish listening and then inspect the face for a second opinion. Visual information can influence the speech percept itself.
A face can change what you think you heard
The famous McGurk effect makes this unusually easy to notice.
In the classic demonstration, a listener hears one syllable while seeing a face articulate a different one. The visual and auditory signals conflict, and some listeners report hearing a third percept or a sound influenced by the visible articulation.
The striking part is not the illusion itself.
It is the fact that information from the face can alter an apparently auditory judgement.
But there is an important warning here. McGurk clips are deliberately mismatched and often use isolated syllables. Natural conversation normally gives you a matching face and voice plus words, syntax, meaning and context. A 2023 review argues that susceptibility to the McGurk illusion is a poor stand-in for how well someone integrates audiovisual information in realistic speech (Getz & Toscano, 2023).
So McGurk is a memorable doorway into the concept.
It is not the whole theory.
Your mouth gives phonetic clues before the word is finished
Visible speech is sometimes casually called “lip-reading,” but that label can make the process sound more complete than it is.
Many speech sounds are visually ambiguous. Different sounds can produce similar-looking mouth shapes. Other articulatory differences happen largely inside the mouth and are difficult to see.
The face therefore does not replace the soundtrack.
It constrains it.
Suppose noisy audio leaves several words plausible. A visible lip closure can make some candidates more likely and others less likely. Timing from the mouth can also help the listener anticipate when acoustic information is about to arrive.
This is a probabilistic advantage rather than a visual transcription system.
That distinction matters for language learners. You do not need to become an expert lip-reader to benefit from seeing a speaker.
The benefit becomes especially valuable when the sound is uncertain
Classic work going back to Sumby and Pollack showed that visible speech can improve word identification in noise. Modern audiovisual-speech research has repeatedly found the same broad pattern: when auditory information becomes less reliable, congruent visual information can contribute more to intelligibility (Irwin & DiBlasi, 2017).
Recent work adds a useful qualification.
A 2026 study compared native and non-native English listeners across several noise levels. Seeing the talker generally improved speech identification, but it did not automatically reduce cognitive costs in every condition. The reduction in dual-task cost appeared when listening was difficult enough that visual information was actually necessary for accurate identification (Brown et al., 2026).
That gives us a better model than:
Video is always easier than audio.
Visual information has value when it resolves uncertainty. If the audio is already effortless, the face may add little for the specific recognition problem being measured.
But extremely bad audio can break the partnership
There is another limit that sounds paradoxical.
If visual information helps when sound gets worse, should it help most when the sound is almost unintelligible?
Not necessarily.
Linda Drijvers and Aslı Özyürek tested highly proficient non-native listeners with clear, moderately degraded and severely degraded English speech. Participants saw blurred lips, visible speech, or visible speech plus an iconic gesture.
The non-native listeners benefited from visual information, but less than native listeners in the comparable work. More importantly, visible-speech benefit was minimal for non-native listeners when the auditory signal was severely degraded. The authors argue that L2 listeners may need enough phonological information in the sound to integrate the visual cues effectively (Drijvers & Özyürek, 2020).
So audiovisual benefit can have a useful middle zone.
When audio is clear, you may not need much rescue.
When it is moderately difficult, the face can disambiguate.
When it becomes extremely poor, there may not be enough auditory structure left for an L2 listener to connect the streams efficiently.
That is far more interesting than “watch the lips.”
The visual channel is carrying more than one kind of clue
| Visual information | What it can contribute | What it cannot reliably do alone |
|---|---|---|
| Lips, jaw, teeth and visible tongue movement | Phonetic and timing information | Uniquely identify every speech sound |
| Facial movement | Timing, attention, affect and conversational information | Provide a word-for-word linguistic transcript |
| Iconic hand gestures | Semantic clues about actions, shape, direction or meaning | Specify every grammatical or lexical detail |
| Wider body context | Discourse and interaction cues | Replace a sufficiently informative speech signal |
Putting all visual information into one bucket hides an important distinction.
Visible articulation can support phonological identification.
Gesture can support meaning.
Those clues may help at different stages of comprehension.
Gestures can tell you about meaning, not just sound
Imagine someone says a verb you almost recognize while making a clear twisting motion with one hand.
The gesture does not show you the consonants.
It may still make the intended action easier to infer.
In a study of native and highly proficient non-native English listeners, iconic gestures and visible speech helped listeners understand degraded speech, although the non-native group generally gained less from the visual information (Drijvers & Özyürek, 2020).
An earlier L2 listening study by Sueyoshi and Hardison used a videotaped English lecture with low-intermediate and advanced learners. Both proficiency groups scored better when visual cues were available than in the audio-only condition, but the pattern differed by proficiency: the advanced group performed best with the face condition, while the lower-proficiency group performed best when gesture and face were both available (Sueyoshi & Hardison, 2005).
That does not establish a universal rule that beginners “need gestures” and advanced learners “need faces.”
It does show why visual support is not one thing.
A mouth can narrow a sound.
A gesture can narrow a meaning.
Myth: if video helps, you were never really listening
Myth: Understanding a speaker only when you can see them means your listening is weak or you are “cheating” with visual clues.
Reality: Face-to-face speech is naturally multimodal. Native listeners also use visible speech, and visual information can alter intelligibility under noisy conditions. Using a speaker's face is part of ordinary human communication, not a loophole in a listening test.
The more useful distinction is between communication and training conditions.
If your goal is to understand a real person, use the information available.
If your goal is to discover what your auditory system can decode without help, audio-only practice can deliberately remove visual support.
Those are different tasks.
The same logic applies to English captions: support can be useful without needing to remain present in every practice condition.
Why second-language listeners may use the face differently
Visual speech is not a universal code that every learner reads identically.
Your language experience changes which sound distinctions matter to you and how much value you assign to auditory and visual cues.
Debra Hardison's work with advanced learners from Japanese, Korean, Spanish and Malay backgrounds found that matched visual cues could improve identification of particular English consonants and that first-language experience influenced audiovisual integration (Hardison, 1996).
Other research has found that the native-language backgrounds of both listener and speaker can alter the amount and pattern of audiovisual benefit in noise (Yi et al., 2014).
This means two learners can watch exactly the same English speaker and extract different amounts of useful information from the face.
The visual signal is the same.
The perceptual history interpreting it is not.
Seeing a face is not guaranteed to improve comprehension
At this point it would be easy to turn the research into a new rule:
Always study English with video.
The evidence is not that clean.
A study by Nobuhiro Kamiya compared L2 English listening tasks in three conditions: upper-body video, close-up face video and audio only, across easier and harder materials. The effects of modality were limited rather than uniformly favouring one visual condition (Kamiya, 2025).
Different experiments also measure different outcomes: identifying a word in noise is not the same task as understanding a lecture, remembering a story or learning vocabulary.
And visual information can demand attention as well as provide information.
The correct conclusion is conditional:
Seeing the speaker can make speech easier to identify when the visible cues are informative and the listener can integrate them with the available audio.
That is less catchy than “video beats audio.”
It is also much more useful.
Do not confuse a talking face with every kind of video
A close-up of a speaker's mouth, a lecture showing the speaker, a film scene, a slideshow with narration and a TikTok full of cuts and text are all “video.”
Cognitively, they are not the same input.
A talking face can provide articulatory timing.
A meaningful gesture can provide semantic information.
A diagram may explain the topic without helping speech identification.
Decorative movement may add nothing useful to the spoken signal.
This matters when interpreting studies of audiovisual language learning. A recent meta-analysis of 56 experiments found an overall benefit of audiovisual input for L2 learning, with outcomes varying by video category, but that literature addresses broader language learning from video rather than the narrower question of seeing the talker's articulatory information while perceiving speech (Montero Perez et al., 2026).
Do not collapse those questions into one claim.
What this changes about listening practice
You do not need to choose between “video learner” and “audio learner.”
Instead, change the available information depending on what you are training.
If a difficult video becomes understandable when you can see the speaker, pay attention to why. Was the mouth helping you identify words? Did gesture reveal meaning? Was the visual scene simply giving you context?
Then remove one support.
Listen to the same or comparable speech without the face. If comprehension collapses, you have discovered something specific about your auditory recognition rather than proving that the earlier comprehension was fake.
The practical loop can be:
- Use natural audiovisual speech to experience the full signal.
- Try audio only to expose what your ears cannot yet resolve.
- Return to the video and locate the visual clue that changes the percept.
- Listen again without video and see whether the sound has become identifiable.
That fits naturally with the broader listen → check → listen approach. The difference is that the “check” is not always text. Sometimes the missing evidence is on the speaker's face.
Seeing speech is part of hearing speech
English does not literally become a different language when you close your eyes.
But the evidence available to your perceptual system changes.
In face-to-face conversation, your brain can combine the voice with visible articulation, facial timing, gesture and context. Those signals can be especially valuable when sound is uncertain. For second-language listeners, however, the benefit depends on proficiency, language experience, the type of visual cue and how much usable auditory information remains.
So if English becomes harder on the phone than in person, that does not automatically mean you imagined your face-to-face listening ability.
You removed a channel that real communication normally gives you.
Use video when the goal is communication and rich input. Use audio-only conditions when you want to stress-test auditory recognition. Move between them when you want the visual signal to teach your ears what they missed.
The deeper idea is simple:
Listening is not just what reaches the ears. It is what the brain can infer from all the evidence arriving together.
Sources and further reading
- Getz, L. M., & Toscano, J. C. — Audiovisual speech perception: Moving beyond McGurk
- Irwin, J. R., & DiBlasi, L. — Audiovisual speech perception: A new approach and implications for clinical populations
- Drijvers, L., & Özyürek, A. — Non-native Listeners Benefit Less from Gestures and Visible Speech than Native Listeners During Degraded Speech Comprehension
- Sueyoshi, A., & Hardison, D. M. — The Role of Gestures and Facial Cues in Second Language Listening Comprehension
- Hardison, D. M. — Bimodal Speech Perception by Native and Nonnative Speakers of English
- Yi, H.-G. et al. — Nonnative audiovisual speech perception in noise: dissociable effects of the speaker and listener
- Kamiya, N. — The limited effects of visual and audio modalities on second language listening comprehension
- Brown et al. — The dual-task costs of audiovisual benefit: Effects of noise and native speaker status
- Montero Perez et al. — The effects of audiovisual input on second language learning: A meta-analysis