If you learn a language mostly by watching things in it, there's a particular moment that tends to catch people off guard. It isn't while you're watching. That part usually goes well, often better than you expected. It comes later, the first time someone asks you something directly and you realise that all those hours of understanding haven't left you able to say much back.
Many learners find that understanding a familiar video feels easier than producing an answer under conversational time pressure. That observation can be a useful prompt to practise output; it is not a measurement of what immersion will do for every learner.
The short version: comprehension and speaking overlap, but they place different demands on a learner. Input can build understanding; a little low-stakes output can reveal whether you can retrieve and assemble language under a new demand. This article offers ways to practise that gap without claiming one routine fits everyone.
That gap isn't a sign you've been doing immersion wrong. Our guide to a stalled routine explains how to examine one learning problem at a time. Here, the narrower question is whether understanding a video has prepared you to retrieve and arrange language for yourself.
Understanding and speaking pull on different muscles
Research on immersion classrooms helped motivate the idea that learners may need opportunities to notice and attempt language, alongside input. That does not mean input is ineffective or that every learner needs the same amount of speaking practice. It is a reason to add low-stakes production when a learner is ready for it.
Listening allows context and partial recognition to carry part of the meaning. Building a sentence requires selecting a word, putting it in the right form, and arranging it yourself. Those demands overlap, but they are not identical.
Why watching more doesn't close the gap by itself
The natural response is to immerse harder. More hours, more shows. That keeps growing your comprehension, which is genuine progress. There's just a catch that's easy to miss: you mostly get better at the exact thing you practise. Hours of listening grow your listening. They don't quietly turn into speaking, because the practice that builds understanding and the practice that builds production aren't the same work.
Vocabulary shows the same split. Across a lot of studies, the number of words a learner can recognise is much larger than the number they can produce on demand. Recognising an expression when it slides across the screen is a different act from reaching for it, unprompted, in a sentence you're trying to build right now.
まだ · yet
Don’t give up yet.
まだ adds ‘yet’; 諦めないで is a request not to give up.
Retrieval practice can improve delayed recall in some settings. The often-cited Roediger and Karpicke (2006) experiment compared study and test conditions with prose material. It supports including attempts to recall in a routine; it does not establish a percentage gain for Japanese video learning or for Glosaire.
Speed is its own skill
There's one more layer: knowing a word and reaching it in time are not the same thing. Speech rate is affected by task, familiarity, first language, and many other factors. The study on second-language speech rate is relevant context, but it does not provide a universal native-speaker comparison or show that retrieval alone explains a gap.
This part is trainable, and teachers have known how for decades. One classic drill has you give the same short talk three times over, with less time on the clock each round. You're not adding anything new. You're getting faster at what you already half-know, until reaching for it stops feeling like a search.
None of this is a knock on immersion
It's worth being clear about what this argument is. It isn't that immersion is overrated, or that input doesn't matter. Input is the foundation, and you can't practise your way into a language you haven't heard enough of. Even the most input-first communities take this seriously. Approaches like Refold build a dedicated output stage into their roadmaps. They just place it later, once you've taken in enough of the language for production to have something to pull from. The argument in that world is about when to start speaking, not whether speaking needs its own practice.
And the caution about timing is fair. Start output with language you have met before and keep the stakes low. That is different from treating every beginner as ready for a live conversation on day one.
If anything, immersion gives you a head start here. All those months of taking the language in mean that when you do turn to speaking, you're working with something rather than starting from a blank. That's a very different position from a beginner trying to produce and absorb at the same time.
How we built practice around the gap
This is the thinking behind the practice system in Glosaire, and the reason it doesn't stop at review.
It starts from what you actually watched. When you finish a video, Glosaire can build a session out of the words you just met in it, so the language you're producing is language you've already seen in context rather than a list from nowhere. A new word turns up first in gentle, recognition-style questions. As you keep getting it right and it settles in, the same word begins to ask more of you. You move from picking it out of a lineup, to typing it into a sentence, to writing your own sentence with it from scratch.
Choose the completed-action form.
傘を( )。
Writing a sentence yourself adds a different demand from recognising one. Feedback can help you inspect an attempt and choose what to adjust next, but it does not guarantee speaking progress or replace listening and interaction.
Choose the polite café response.
コーヒーをお願いします。
Writing a sentence yourself, rather than only understanding one, does more for your speaking than it might seem. The hard part of saying something is getting from your own meaning to the right word in the right form, and that runs much the same whether the sentence ends up typed or spoken. It's the step immersion never asks for, and practising it in writing has been shown to carry over to speaking, particularly earlier on. What writing doesn't give you is the sound of the word and the speed of reaching it, and that's what the next two pieces are for.
Then there's speed. Speed Round is a 60-second rapid-fire practice mode that sends words at you one after another and you answer as fast as you can. It's less about whether you know them and more about how fast the meaning comes, so the words you've learned start surfacing the moment you need them instead of a beat too late. That sits alongside the rest of the system: ten exercise types in all, interleaved within a session, from flip flashcards and cloze questions through to pair matching, listening (both single words and full sentences with a blank), kanji recognition, sentence construction, and speaking practice where you say a line aloud and see, word by word, which parts landed. Four of the ten make you produce the language rather than just recognise it, and new words climb five stages, New through Mastered, on that schedule.
None of this is meant to replace watching. Think of it as the other half of the same habit. You take the language in by watching things you enjoy, and then you spend a few minutes turning some of what you understood into something you can use.
Where to start
If you're already immersing and enjoying it, keep going. Add a small output experiment when it helps you check retrieval: take one expression you understand and try to use it in a low-stakes sentence. Notice whether that makes the next attempt easier, and adjust the amount of support you use.
To try this with your own watched material, create a Glosaire account. The research below describes general learning findings; it is not a study of Glosaire.
The research behind this post
- Robert DeKeyser and Robert Sokalski on why practice tends to build the specific skill you practise: DeKeyser and Sokalski (1996)
- The gap between words you recognise and words you can produce: Laufer and Paribakht (1998)
- Why recalling beats rereading: Roediger and Karpicke, "Test-Enhanced Learning" (2006)
- The speaking-speed gap between learners and native speakers: study on second-language speech rate
- Building speaking speed with timed repetition (the 4/3/2 technique): Nation's fluency activity
- Whether practising in writing carries over to speaking: Sletova (2023), on writing as a scaffold for speaking accuracy
Common questions
Does immersion alone teach you to speak Japanese?
Input can build understanding, but it may not fully prepare you to assemble a sentence under conversational time pressure. Comprehension and production overlap while placing different demands on a learner, so low-stakes speaking or writing can reveal what still needs practice.
What is the output gap in language learning?
It describes a difference a learner may notice between language they can understand and language they can retrieve and arrange themselves. Recognising a word and recalling it without a prompt are related but different tasks.
How does Glosaire train speaking and writing?
Glosaire includes several practice formats, including speaking and written-response activities. Check the current signed-in product flow for which formats are available to your account; the article's research links describe general learning evidence, not a Glosaire efficacy study.
Learn Japanese from real YouTube videos.
Tap any word. Understand it in context. Never forget it.
Try Glosaire free →Keep learning

Learning Method
傘を持ってくるのを忘れました。
I forgot to bring my umbrella.
Japanese Sentence Mining: What to Save and What to Skip
Build useful Japanese vocabulary cards from videos without saving every unknown word. Includes worked examples and a manageable review routine.

Learning Method
もう帰る?
Going home already?
Can Read Japanese but Can't Understand It Spoken? Try This
Find out whether vocabulary, sound recognition, or processing speed is limiting your Japanese listening, then choose a focused exercise.