How to practice English speaking alone at home

Most solo “speaking practice” is listening in disguise. Here is which methods actually produce speech, how to correct yourself without a teacher, and how to tell when a mistake is finally closed.

·

·

9–13 minutes
Solo English practice methods arranged from zero spoken words on the left to self-built speech on the right

You can practice English speaking alone, but only with the methods that make you build a sentence yourself. Talking to yourself, answering a real question out loud into a recorder, and retelling your day produce speech. Shadowing, podcasts and series produce input wearing a speaking costume. The second thing you need is a correction source, because on your own you cannot reliably hear what went wrong. This page sorts the popular solo methods by how much speech they actually produce, gives you four ways to correct yourself without a teacher, and a countable rule for deciding when a mistake is closed.

Why speaking alone usually changes nothing

Reading, listening and watching train comprehension. Speaking is a different job: you have to assemble a sentence under time pressure and say it. Merrill Swain’s comprehensible output hypothesis names what happens when you do that. Learners “encounter gaps between what they want to say and what they are able to say, so they notice what they do not know or only know partially” (Comprehensible output). The same source summarises her argument that output “facilitates second language learning in ways that differ from and enhance input due to the mental processes connected with the production of language”.

This is not settled science, and you deserve the other side. Stephen Krashen answers that producing language is rare in everyday use, and comprehensible output is rarer still, so he treats more input as the safer bet. There is also evidence that building beats copying: in a pronunciation study summarised by The Learning Scientists, “retrieval practice outperformed imitation on both immediate and delayed tests“. That study tested vocabulary pronunciation, not free conversation, so read it as a signal rather than a proof.

Hold on to one distinction for the rest of this page. If you built the sentence, it is output. If you repeated or absorbed someone else’s, it is input. Both have value. Only one of them is speaking practice. If the sentence starts and then dies in your mouth, why your sentence stops halfway is the companion piece to this one.

Which solo methods actually make you speak

The table below tags each popular solo activity as INPUT or OUTPUT and estimates how many words you construct yourself in a ten-minute sitting. The word counts are our own arithmetic, not a research finding. We assume roughly 120 words a minute while you are actually talking. That sits at the low end of a published conversational range that runs up to about 150 words per minute, a figure the same page attributes to the National Center for Voice and Speech. Then we multiply by the minutes you really spend speaking rather than listening or reading. No verified figure exists for second-language speakers specifically, so treat every number as an order of magnitude.

MethodWho builds the sentenceTagMinutes talking in 10Words you construct
Podcast or series with subtitlesnobody, you listenINPUT00
Shadowing a recordingthe recordingINPUTabout 80
Reading a text aloudthe authorINPUTabout 90
Repeating app sentencesthe appINPUTabout 30
Narrating what you are doingyouOUTPUTabout 4about 480
Retelling your day out loudyouOUTPUTabout 5about 600
Recording an answer to a real questionyouOUTPUTabout 4about 480
Role-play with an AI coachyouOUTPUTabout 4about 480

The INPUT tag is a description, not an insult. Shadowing in particular does measured work: by “synchronising with the speaker’s voice in real time, learners refine their intonation, pitch contours, and rhythm” (The Language Gym). That is worth doing. It is simply not the thing that fails you when a colleague asks an unexpected question.

How do you correct yourself when nobody is listening

Here is the failure that quietly ruins solo practice. You rehearse your own mistakes for months because nothing interrupts them. Willem Levelt’s perceptual loop theory describes speakers as monitoring their own output by comparing it against an internal standard (self-monitoring in speech production). Without an explicit standard to compare against, the monitor “has only statistical information about the occurrence of an error to go on”. In plain words: alone, you often feel that something was wrong without knowing what or where.

Correction sourceFeedback arrivesCatchesBlind spot
AI coach naming the error in the sessionsecondsgrammar, word choice, the missing wordnot a human ear, no social judgement
Your own recording played backa minute laterpace, hesitation, obvious slipsonly if you already know the correct form
Automatic transcript you readminuteswrong words, dropped endingstranscription errors on accented speech
Grammar checker on that transcriptminuteswritten-level grammarnothing about pronunciation or rhythm

Automatic speech recognition sits behind three of those four rows, and its accuracy matters. One study of English learners used software whose developers reported “an accuracy rate exceeding 90%” (ASR and EFL speaking skills). The same paper notes that advanced systems “have the ability to offer feedback on the sentence, word, or textual level”. That figure is developer-reported, and accented speech is measurably harder for these systems, so read a transcript as a strong hint rather than a verdict. Using an AI coach as your correction source shortens the delay to seconds, which is the only reason it ranks first.

Tip

Record every session on your phone, even the bad ones. A month of audio is the only record that shows whether an error you named in March still appears in April.

What do you talk about when you are alone

Most people run out after forty seconds because the prompt was generic. “Describe your city” belongs to nobody. Swain’s noticing function suggests the useful gaps show up when you reach for something you actually need to say, so build prompts from the week in front of you. This is reasoning from the output hypothesis, not a separate research finding.

  • The hardest part of today, in five sentences
  • A message you sent in Persian today, said again in English
  • What you will ask in your next meeting, out loud
  • A decision you made this week and why
  • The last thing that annoyed you, described without swearing

Here is one worked example. Prompt: what was the hardest part of your workday? A recorded answer might come out as “I am agree with my manager, but I have problem with the new deadline.” Two named things, not a bad sentence. “I am agree” is a straight translation of من موافقم (man movafegh-am), where Persian builds agreement as an adjective, so English wants “I agree”. And Persian has no article system like English, so “problem” feels complete while English asks for “a problem”. Fix one of them tonight, not both, and name the mistake you keep making before you start.

Build a two-minute solo session from these methods

  1. Pick one prompt from your own week, ten seconds.
  2. Answer it out loud with no script, sixty seconds.
  3. Play the recording back once, thirty seconds.
  4. Name exactly one thing that was wrong, ten seconds.
  5. Say the corrected sentence twice, ten seconds.

One target per session is the point, and it is not modesty. Nelson Cowan’s review of working memory “proposed that the limited focus-of-attention capacity averages about four chunks in normal adults” (modelling working memory capacity). Speaking already occupies most of that. Add a grammar checklist on top and something gets dropped, usually the fluency. Two-Minute Speaking runs this same loop and teaches exactly one point per session; it prompts you in your own language and takes either voice or text. If you want the weekly shape rather than the single session, see what one week of two-minute sessions looks like and how a two-minute session works.

Four-stage loop showing prompt, spoken answer, playback with one part highlighted, and a corrected repeat

How do you know a mistake is actually fixed

“Practise more” has no finish line, so here is ours, published openly. Every error you meet becomes a named item in an error ledger across six categories. The item closes after three correct uses in three separate sessions, spaced on an FSRS schedule rather than crammed into one evening. You can set that number anywhere from one to ten. At most three items stay open at a time, which is why the session above names one thing and walks away from the rest.

One error card with three ticks on three separate days next to a tray holding only three open items

The spacing part rests on real evidence. Cepeda, Pashler, Vul, Wixted and Rohrer reviewed a large body of distributed-practice research across hundreds of experiments and found the optimal gap between study sessions grows as the retention interval grows (distributed practice in verbal recall tasks). Separate days beat one long night.

Warning

Three is a design choice, not a measured threshold. No study we found says an English error closes after exactly three correct uses, and Two-Minute Speaking has no published outcome data. Treat the number as a countable habit, not a guarantee.

The paper version costs nothing. Write the error at the top of a card, write the corrected sentence under it, and add one tick each time you use it correctly. Ticks only count from separate sessions on separate days. Three ticks, card goes in the drawer.

Why shadowing and podcasts still deserve twenty minutes a week

The table demoted shadowing, and the table is not the whole truth. Shadowing trains learners to process input and output at the same time, which builds mental agility and fluency under pressure. It is also one of the few solo activities that touches rhythm and intonation at all. If your English sounds correct on paper and flat in the air, that is a prosody problem, and no amount of self-talk fixes it. Read how shadowing is actually done before you judge it by two minutes of mumbling along.

Podcasts and series earn their place differently: they are where new words and phrasings arrive. Output practice with a thin vocabulary base stalls quickly, because you can only build sentences from material you have met. We suggest roughly two thirds of your weekly minutes on output and one third on input, which for most readers means twenty to thirty input minutes a week. That split is our recommendation, not a finding.

Pick tonight’s method in three questions

  • Did you build the sentence yourself, or repeat someone else’s?
  • Will something tell you whether it was right?
  • Can you name the one thing you are working on?

Two failures out of three means the activity is study, not speaking practice. Study is fine, just do not count it twice. If the answer to the second question is no, record the session so a correction source exists later tonight. If you cannot name a target, spend the first session naming one instead of chasing fluency. And if you genuinely cannot speak out loud right now, type the answer at full speed without editing, then say it aloud tomorrow before you read the correction.

Questions people ask about practising alone

Can I really improve my English speaking if I have nobody to talk to? Yes, within limits. Methods where you construct the sentence yourself train the thing that fails you, and you can run all of them alone. What solo practice cannot train is turn-taking, interruption, unpredictable questions and listening under pressure, so treat it as preparation for conversation rather than a replacement.

Is talking to yourself in English useful, or does it just build bad habits? It is useful for fluency and genuinely risky for accuracy. Talking to yourself produces self-built sentences, which is exactly what you want, but nothing in the room corrects them. Pair it with a correction source in the same session and the risk mostly disappears.

How do I correct my own mistakes without a teacher? Use one of the four sources in the table above: an AI coach in the session, your own recording, an automatic transcript, or a grammar checker run on that transcript. Rank them by how fast the feedback reaches you. Your unaided ear is the slowest option, because it only registers that something felt wrong.

Is shadowing speaking practice or listening practice? Functionally it is closer to listening and pronunciation practice. You reproduce someone else’s sentence rather than building your own, so it trains prosody, intonation and processing speed, not sentence construction. Keep it, but do not let it stand in for the sessions where you answer a question nobody prepared you for.

How long should I practise speaking alone each day? Short and daily beats long and occasional, because the spacing evidence favours separate sessions over one massed block. Two focused minutes with one named target is a real session. The weekly shape is a separate question, and the daily practice article answers it properly.

You now have the method; what you do not have alone is something that names the mistake while you are still in the sentence. A short placement session will name your first one, so start a two-minute session tonight.

Keep reading

More in this section, or start from the full list of articles.

Written by

Meet the team · How we write and check articles


Leave a Reply

Your email address will not be published. Required fields are marked *