Three stages sit between a thought and a spoken sentence, and if you understand English but can’t speak it, the stall is almost always in the middle one. Recognising a word asks you to match it against context, tone and the rest of the sentence, and much of that work is done for you. Producing the same word asks you to pull it out of nothing, build the grammar around it and move your mouth, live, with no support. Your knowledge is real. The retrieval path was never built, and more input will not build it.
Why understanding English asks less of you than speaking it
The two abilities have names. Your receptive vocabulary is the set of words you understand when you hear, read or see them; your productive vocabulary is the set you can produce in the right context with the meaning you intended, and the receptive one is normally the larger of the two. That gap is the ordinary shape of word knowledge. It is not a defect you developed.
Comprehension is also scaffolded from outside your head. The speaker’s tone, their gestures, the topic and the social situation can hand you the meaning of a word you do not actually know. None of that scaffolding exists when you are the one speaking. Merrill Swain made the same point in mechanism terms: learners can deduce meaning from context without ever processing the grammar, and that shortcut is much harder to use when you produce.
Here is a concrete test, and it is ours rather than a cited finding. You read “I’d rather you didn’t” and understand it instantly. Now imagine assembling it inside a live reply, under a two-second pause, with someone watching your face. Same item, two entirely different operations. One you have performed thousands of times. The other, possibly never.
What happens in the half second before a sentence comes out
“My mind goes blank” is a description, not a diagnosis. The standard model of speech production breaks that half second into named stages. In conceptualization you determine what to say/09%3A_Speaking/9.02%3A_The_Standard_Model_of_Speech_Production). In formulation, concepts are connected to words and built into a syntactic, morphological and phonological structure: word forms are selected in lexicalization and put together in syntactic planning. In articulation, that structure is phonetically encoded and comes out as sound.

Notice where you actually stall. You knew what you wanted to say. Your mouth works. The break is in formulation, the stage that has to find the words and assemble the grammar fast enough to keep a conversation moving. Explaining a rule and executing it at speed are not the same act.
That is where the strongest evidence on this page lands. Robert DeKeyser taught 61 subjects the same morphosyntactic rules over eight weeks, gave everyone the same amount of comprehension and production practice, and varied only which rules were practised in which mode. The finding: rule learning is “highly skill-specific” and these skills “develop very gradually over time, following the same power function learning curve as the acquisition of other cognitive skills.” The study used an artificial language and a small group, so hold it lightly. The shape is still hard to argue with, and our reading of it is this: you practised comprehension for years and got, precisely, better comprehension.
Is your English not ready or are you just being watched
Two different problems produce the same silence, and the fix for one does nothing for the other. Foreign language anxiety is situation-specific and affects people who are not otherwise anxious. It occupies working memory, leaving fewer resources for complex grammar and vocabulary, and listening and speaking are regularly cited as the most anxiety-provoking language activities. Because it is situation-specific, you can test for it by removing the situation. The test below is our own reasoning, not a published instrument.
- Alone in a room, unobserved, with no time limit, say the sentence you could not say. Does it still collapse?
- Now give yourself thirty seconds to write the same sentence. Does it collapse there too?
Read the result plainly. Collapse in both means a retrieval gap, and the rest of this page is written for you. Fine alone and fine in writing, but gone the moment a person is in front of you, means the cause is the room rather than your English: read why your sentence stops halfway and English speaking anxiety and what lowers it, because this page will not help you much. Both at once is very common, and retrieval is still the half you can train on your own.
Does this mean your real English level is lower than you think
No. Your reading level is real and it was not a trick. A receptive vocabulary being larger than a productive one is the normal case for anyone who has a vocabulary at all, in a first language as well as a second. What you have is an asymmetry, which is the usual shape of language knowledge, not a fake level.
The two are also measured separately, which is the useful part. Stuart Webb tested five aspects of word knowledge with both a receptive and a productive test for each target word. If researchers need two instruments, a reading score was never going to tell you what happens on a call.
Note
We will not estimate your level here, and Two-Minute Speaking does not sell a test. No reliable figure exists for how much smaller a typical productive vocabulary is, so you will not find a ratio or a percentage on this page.
Why more films and more podcasts do not close the gap
Input built the understanding you already have. It is very good at that and it stays useful. What it does not do is make you meet the gap. The comprehensible output argument is that learning happens when learners run into a gap in their linguistic knowledge and notice the distance between what they want to say and what they can say. You cannot run into that gap while listening. Listening is where the gap stays invisible.
The opposing case deserves stating properly, because it is not a fringe position. Stephen Krashen’s input hypothesis holds that language is acquired by receiving comprehensible input slightly above your current level, and his objection to output accounts is that output is rare and comprehensible output is rarer still. That is a real argument, and the question is not settled. A 2025 critique in Frontiers in Psychology answers it this way: understanding is necessary, but understanding alone without use is not sufficient to rewire linguistic capacity, and active use is a different and arguably deeper stimulus than passive input.
The closest thing to a natural experiment is immersion. Swain observed that French immersion students in Canada very rarely said anything longer than a clause, and that many graduates still had grammatical inaccuracies in their speech. Those were children in Canadian school programmes, not adults with your reading level, so do not take it as a finding about you. Take it as evidence that years of comprehensible input alone did not produce fluent production even under generous conditions.
Be careful with the evidence pointing the other way too. Webb’s first experiment, with time held equal, favoured the receptive reading task; his second, when tasks were given the time they genuinely need, found the productive writing task more effective. That is one study contradicting itself by design, and reporting both halves is the honest version. Shadowing sits in a similar place. It trains your mouth and your ear, and in our view it does not train retrieval, because the words are handed to you; that is our editorial position rather than a research finding, and we set out what shadowing trains and what it does not separately.
How to turn words you understand into words you say
The conversion is a loop, not a course. It runs on material you already have, which is everything you understood today and would not have produced. One item at a time.

- Notice one item from today’s reading, film or meeting that you understood instantly and would never have said.
- Build your own sentence with it, about your own life, before you check anything.
- Say it out loud. Out loud matters, because articulation is a stage you have to run.
- Get one correction on that sentence, and one point only.
- Say the corrected version once, immediately.
- Come back days later and use the item again in a new sentence about something else.
Tip
Build the sentence before you check it. The gap only shows itself if you attempt production first, and looking the answer up beforehand skips the exact step that does the work.
Make the sentence about your own life rather than a textbook topic. That is Two-Minute Speaking’s design choice rather than a research finding, and the reason is plain: your own life is what you will actually have to talk about. Spacing is better supported. A summary of Cepeda and colleagues’ 2006 review reports that spaced practice beat massed practice in 259 of 271 comparisons, a second-hand figure but a consistent one. Six short separated attempts beat one long evening, which is the whole case for two minutes of speaking a day. Step four is where how AI English speaking practice works becomes relevant, because the correction has to come from somewhere.
How to know a word has actually crossed over
“When it feels natural” is not a stopping rule. Two-Minute Speaking uses a countable one instead, and these are the product’s own design rules rather than findings about all learners. An item closes after three correct uses in three separate sessions at spaced intervals, three being the default and adjustable between one and ten. At most three items stay open at once, on the principle that no new territory opens while old problems are open. One point is taught per session. You can copy the same rule into a notebook today.
| Item you are activating | First correct use | Second | Third | Status |
|---|---|---|---|---|
| I’d rather you didn’t | Mon, own sentence | Thu, own sentence | not yet | open |
| by the end of the week | Tue, own sentence | not yet | not yet | open |
| I’m not sure that follows | not yet | not yet | not yet | open |
| a fourth item | blocked | blocked | blocked | waiting for one to close |
The cap is the part people skip, and it is why most self-study lists fail. Twenty open items means twenty items rehearsed once and none of them retrievable. Three means three that leave your mouth without a search. You can see the same rules running inside the app in how Two-Minute Speaking works.
What stays difficult after you can speak
Activation gives you retrieval. It does not give you idiom, register, humour or the cultural scripts that decide when a joke lands and when a direct request sounds rude. Accent remains its own long project. The curve is gradual by nature as well: these skills develop very gradually on a power function, so expect improvements to get smaller rather than to arrive as a breakthrough. Immersion graduates kept grammatical inaccuracies for years, and that is the ordinary ceiling of any single method.
Warning
What Two-Minute Speaking does not do: it is not a long free-form conversation partner, it does not score your accent, and it promises no timeline and no test result. No reliable figure exists for how long activation takes, so we do not give one.
Questions people ask about speaking the English they already know
Why can I read English but not speak it? Reading is recognition and speaking is recall, and only recognition gets help from tone, gesture, topic and context. When you read, the page does half the work; when you speak, nothing does. The practice you put in built the skill you practised, which was comprehension.
Is it anxiety or is my English not good enough? Test it by removing the audience. Say the sentence alone, unobserved and untimed, then try writing it in thirty seconds. Collapse in both means a retrieval gap; collapse only in the room means foreign language anxiety, which is situation-specific by definition. Many people have both at once.
Does understanding English without speaking it mean my level is fake? No. A receptive vocabulary is normally larger than a productive one for everyone, in every language they know. Researchers measure the two with separate tests precisely because they are separate abilities, so your reading level is a real measurement of reading and not a verdict on your speaking.
How long does it take to start speaking what I already understand? No citable figure exists and we will not invent one. What is documented is the shape: automatization develops very gradually, on a power function curve, so early gains feel larger than later ones. Anyone giving you a number of days is guessing.
Can I fix this without a speaking partner? Mostly, yes. The mechanism that matters is attempting production and noticing the gap between what you meant and what came out, and that needs an attempt plus a correction rather than a conversation. Our guide to how to practise English speaking alone covers the methods that actually produce speech.
The drill above needs one thing you cannot supply yourself: someone to tell you exactly what was wrong with the sentence you just built. Two-Minute Speaking does that in sessions under two minutes, one point at a time. Start a two-minute session and find out which of the English you already understand is ready to be said.

Leave a Reply