Filler words — 'um,' 'uh,' 'like,' 'you know,' 'so,' 'actually' — are the verbal placeholders most of us insert while our brains catch up with our mouths. They feel harmless, but they add up. A speaker who says 'um' once every ten seconds litters a five-minute talk with thirty fillers, and listeners notice. Forbes has reported on the neuroscience behind this: fillers erode perceived credibility because the audience's brain interprets hesitation as uncertainty about your own material. The good news is that filler words are one of the most fixable speaking problems. Most people can cut their filler rate by half or more within four to eight weeks using deliberate practice. This guide covers why you say them, how to stop, what tools help (including AI transcription software that turns your speech into searchable text so you can count and analyze your own habits), and which common approaches backfire.

Why You Say 'Um': The Mechanics Behind Filler Words

Also worth reading: how to practice public speaking with AI feedback? · How can I fix whisper hallucination in AI transcriptions and what are the best tips to reduce errors in medical or clinical audio? · How can businesses effectively reduce speech to text API expenses in 2026 without sacrificing accuracy?

Filler words exist because silence feels dangerous to speakers. When you pause mid-sentence, you fear two things: that someone will jump in and steal your turn, and that the gap exposes you as unprepared or slow. So the brain deploys a low-cost sound — 'um,' 'uh,' 'like' — to hold the conversational floor while it retrieves the next word. Cognitive scientists describe this as managing working memory under load: when you're composing complex sentences in real time, especially on unfamiliar topics, retrieval slows down and the filler fills the gap.

The problem compounds under stress. In high-stakes settings — interviews, presentations, first dates — anxiety narrows your attention, which degrades working memory further, which produces more hesitation, which produces more fillers. It's a feedback loop. Research on public speaking anxiety consistently shows that filler frequency rises with perceived stakes. This is why people who speak fluently in casual conversation suddenly sound like they're stalling in a boardroom. Understanding this loop matters because it tells you the fix isn't 'try harder to be confident' — it's reducing cognitive load and rehearsing transitions so your brain never hits an empty buffer mid-sentence.

There's also a social dimension. The New York Times interviewed teenagers in 2024 about saying 'like,' and many described it as a social lubricant rather than a hesitation marker — a way to soften statements, signal relatability, and avoid sounding overly assertive. That means not every filler is a flaw; some are doing pragmatic work. The goal isn't zero fillers (broadcasters themselves average roughly one per minute), it's eliminating the ones that read as nervousness rather than style.

Measure First: Use Transcription to Count Your Fillers

You cannot reduce what you haven't measured. The single highest-leverage step is recording yourself and getting an accurate transcript, because memory lies — most speakers estimate they use far fewer fillers than they actually do. Modern AI transcription tools make this trivially easy. Services like transcribeall.io convert recorded audio into text automatically, and because AI transcripts capture every word verbatim — including every 'um,' 'uh,' and repeated phrase — you get an objective, timestamped record of your verbal habits. Speech-to-text engines from Google and others now recognize punctuation and sentence structure directly in transcription output, so you can see exactly where your sentences break down.

The workflow is simple: record a three-minute impromptu talk on any topic (your job, a hobby, a recent news story), run it through a transcription tool, then search the text for your filler words. Count them. Divide by minutes spoken to get your filler rate. Anything above four or five per minute is distracting; below two per minute, most listeners stop noticing entirely. Repeat weekly. Because the transcript gives you timestamps, you can also identify trigger moments — do fillers cluster at the start of answers? During topic changes? After questions? That pattern tells you precisely where to focus your rehearsal effort, which is far more efficient than vague advice to 'speak more confidently.'

A practical note on tooling: raw ASR output can be messy, and the industry has responded. Superwhisper released S1-mini in 2025, a 462 MB open-weights text normalizer specifically designed to clean up raw ASR transcripts, and dictation apps like Wispr Flow and ShoutFlow (which launched as a pay-once, on-device Mac app) compete on producing clean written text from natural speech. For filler analysis you actually want the opposite of aggressive cleanup — verbatim accuracy matters more than polish — so choose a transcription mode labeled 'verbatim' if available, or review the raw output before any auto-formatting strips your fillers away.

The Pause Replacement Technique

The core skill replacing fillers is the deliberate pause. Radio operators figured this out decades ago: aviation radiotelephony procedure documents like CAP R100-3 and military radio operator handbooks train communicators to use structured silence instead of noise. On a two-way radio, keying the mic with dead air is preferable to transmitting garbage, and disciplined operators learn to think silently, then transmit complete thoughts. You can borrow this discipline directly. When you feel the urge to say 'um,' close your mouth instead. Hold eye contact. Breathe through your nose. Then continue.

The reason this works is perceptual: research consistently shows audiences judge a silent pause as shorter than it actually is. A two-second pause feels like an eternity to the speaker but reads as thoughtful composure to the listener. Matt Abrahams, who discussed speaking clearly on the Huberman Lab podcast, emphasizes that pausing signals confidence rather than incompetence — listeners interpret silence as deliberation, whereas fillers signal panic. Start small: practice pausing for one full second before answering any question, even casual ones. It will feel absurd for about a week. Then it becomes automatic, and the automatic version is what shows up under pressure.

Pair the pause with slower overall pacing. Elderspeak research — which examines how speech patterns change in caregiving contexts — notes that slowing down reduces both repetition and filler density, because a slower rate gives working memory more time per syllable. Aim for roughly 120 to 150 words per minute in formal settings; most anxious speakers exceed 170, which crowds out thinking time and forces fillers into the gaps.

Practical Training Steps That Actually Work

Structured practice beats passive awareness. Toastmasters remains the gold standard here: clubs like Electric City Toastmasters in Scranton offer free public meetings where members give short speeches and receive explicit feedback on filler counts — some clubs literally tally your 'ums' on paper during your speech. That immediate, external count is uncomfortable and effective. If no club is nearby, replicate the mechanism by having a friend click a counter every time you fill, or by reviewing your own AI-generated transcripts weekly as described above.

Specific drills worth adopting: first, the 'answer in three beats' drill, where you structure every response as point one, point two, point three, announced explicitly ('I'll cover three things'). Pre-announced structure eliminates the mid-answer search that triggers fillers. Second, record and re-record the same sixty-second answer until you can deliver it with fewer than two fillers — usually takes three to six takes. Third, practice reading aloud at deliberately slow speed with full stops between paragraphs, training your mouth to tolerate silence. Fourth, replace verbal fillers with physical ones: a nod, a sip of water, a glance at your notes buys the same thinking time without the sound.

Timeline expectations matter. Expect visible improvement in two weeks of daily ten-minute practice, and near-automatic results around week six to eight. People who plateau usually skipped the measurement step or practiced only in low-stakes conditions; the skill transfers only if you deliberately rehearse in situations that mimic your real pressure environment.

Comparing Your Options: Coaching, Clubs, Apps, and Self-Recording

Different approaches suit different budgets and personalities, and none is universally superior. Here's how the main options compare:

FeatureToastmasters / Speaking ClubPrivate Speech CoachAI Transcription Self-ReviewDictation & Voice Apps
Typical costFree to ~$90/year dues$75–$300 per hourFree tiers; paid plans often $10–$30/monthOne-time purchases (~$50–$150) or subscriptions
Feedback typeLive human tallies and peer critiqueExpert, personalized, real-timeObjective word-level data, timestampedClean text output; limited filler analytics
Pressure simulationHigh — real audienceHighLow — solo recordingLow
Time commitmentWeekly meetingsScheduled sessions10–15 min dailyPassive/daily use
Best forHabitual practice and community accountabilityHigh-stakes deadlines (executives, interviewees)Data-driven self-trackersWriters converting speech to text
Self-recording with transcription is the cheapest entry point and the best measurement tool, but it lacks external pressure, so pair it with at least occasional live practice. Private coaching accelerates results dramatically — a good coach will spot breathing problems and structural issues you'd never see in a transcript — but at $75 to $300 per hour it's hard to justify without a specific deadline. Clubs split the difference: modest cost, real audiences, and built-in repetition over months. Note also that AI dictation tools aimed at productivity (Wispr Flow, ShoutFlow, Android's Gemini-based Rambler) optimize for clean output rather than self-improvement, so don't confuse 'the app removes my fillers from text' with 'I removed them from my speech.'

Common Mistakes That Make Filler Words Worse

The most common mistake is overcorrection into silence avoidance — becoming so focused on not saying 'um' that you rush, mumble, or truncate answers, which damages clarity more than the fillers did. Another frequent error is substituting new fillers: eliminate 'um' and watch 'like,' 'right,' 'sort of,' and 'basically' rush in to fill the vacuum. Track all of them in your transcripts, not just the obvious ones. Repeated phrases ('what I mean is,' 'at the end of the day') deserve equal scrutiny; elderspeak research flags repetition alongside fillers as credibility drains.

People also misdiagnose the cause. If your fillers cluster at the start of responses, your problem is answer-opening, and the fix is rehearsed openers, not general confidence work. If they cluster mid-explanation, your explanations lack structure, and the fix is outlining, not pausing drills. Treating every filler as the same problem wastes months. Finally, beware of perfectionism: chasing zero fillers creates performance anxiety that increases hesitation. Broadcast professionals accept roughly one filler per minute; aim for that band, not zero.

One modern caveat: as AI-generated content floods the internet — what critics call 'AI slop,' defined as digital clutter prioritizing speed over substance — audiences have grown more sensitive to anything that sounds unpolished or evasive. That raises the bar slightly for live speakers, but it also means genuinely human, well-paced delivery stands out more than it used to. Don't let the trend push you toward robotic delivery; measured pauses and natural rhythm remain the goal.

When to Act and What Results to Expect

Start measuring now, regardless of whether you have a big event coming, because the habit loop takes weeks to retrain and pressure amplifies whatever habits already exist. If you have a presentation, interview, or wedding toast in less than three weeks, prioritize the pause technique and pre-structured answers over broad drills — those two interventions produce the fastest visible change. If you're playing a longer game, join a club or commit to a twelve-week self-recording program with weekly transcript reviews.

Realistic outcomes: dedicated practitioners typically cut filler rates from five-plus per minute to under two within six to eight weeks, and report secondary benefits including slower, clearer articulation and reduced speaking anxiety — the same benefits Toastmasters has advertised for a century. The measurement habit itself tends to stick, since running a recording through a transcription service takes under two minutes and gives you a permanent, comparable record of progress. Treat filler reduction not as a one-time fix but as ongoing maintenance: check your rate monthly, and intervene whenever it creeps back above your threshold.