Why speech-to-text is the most overlooked AI unlock

Every time I mention that most of what I write these days starts as speech, not typing, I get the same reaction: a slightly confused look, like I’ve admitted to something a bit odd. It isn’t odd. It’s just faster, and I think it’s one of the most underrated unlocks AI has handed us, more overlooked than half the flashier stuff people build demos around.

flipsnack cQKceh3huwY unsplash 1 1

The maths nobody thinks about

The average person types somewhere around 40 words a minute. Talk instead, and that jumps to 120 to 150. That’s not a marginal gain, it’s three or four times the throughput, for something you already know how to do. You don’t need a faster keyboard layout or a year of touch-typing practice. You just need to stop typing.

I’m a case in point. Years of football left a couple of fingers that never healed quite right, and the keyboard has always been my natural enemy because of it. So when speech-to-text got genuinely good, it wasn’t a novelty for me, it was a fix for something that had been slowing me down for years. But you don’t need bad hands to benefit. You just need to accept that talking is the format your brain already runs in, and typing is the bottleneck you’ve gotten used to tolerating.

The two tools I actually run

I use two, and the choice mostly comes down to hardware.

VoiceDash is the cloud one, and it’s the one I reach for by default. Because the processing happens in the cloud, it doesn’t care whether you’re on a five-year-old laptop or a maxed-out MacBook, it works the same either way, and it follows you onto Android and iPhone too. That hardware independence and cross-device reach is the actual reason I default to it over Handy, more than the AI polish it adds on top. It listens, cleans up the filler words, fixes the grammar, and hands you finished text in real time, across whatever app you’re in. Personal dictionary for names and jargon, snippet library for recurring phrases. Free plan caps you at 1,000 words a month on Mac and Windows; Pro removes the cap and unlocks every platform for $12 to $15 a month depending on billing.

Handy is the opposite instinct: free, open source, runs entirely on your own machine, Whisper or Parakeet doing the transcription locally, nothing ever leaving your laptop. Press a shortcut, talk, the text lands wherever your cursor is. No subscription, no cloud, no account. The trade-off is hardware. On a powerful laptop, especially with a dedicated graphics card, it’s genuinely fast. On an older one, the local model asks for more than the hardware wants to give, and you feel it in the lag. That’s the gap VoiceDash’s cloud processing closes.

VoiceDashHandy
Where it runsCloud (OpenAI models)Locally, on your machine
CostFree tier, then $12 to $15/monthFree, open source
Hardware dependenceNone, same speed on any deviceDepends on your machine, best with a dedicated GPU
PlatformsMac, Windows, Android, iPhoneMac, Windows, Linux
Best forConsistent speed everywhere, older or weaker machinesPrivacy and offline use, if your hardware can keep up

It’s the workflow, not the microphone

The bit that actually moves the needle isn’t transcription accuracy, every decent tool gets that right by now. It’s what happens after the words land. A personal dictionary that stops your tool mangling client names every single time. A snippet library that turns a two-minute explanation into a two-second voice trigger. Text that pastes straight into Slack, or your notes app, or this very document, without you touching the clipboard once. That’s the difference between dictation as a party trick and speech-to-text as an actual part of how you work. Set that up properly and you’re not just talking instead of typing, you’re skipping half the admin around writing altogether.

The catch

Dictation is brilliant for a first pass and terrible for anything that needs heavy restructuring while you write it. Typing lets you see the sentence, backtrack, cross something out mid-thought. Speech is linear, once it’s out, it’s out, and you can’t glance at a paragraph and rearrange it the way you can on a screen. So I still type when an argument needs building in layers. For everything else, the rambling, the drafts, the emails, the notes I’ll tidy up later, it’s speech first, every time.

Why this matters more than people give it credit for

AI gets credit for the flashy stuff: agents, image generation, code that writes itself. Speech-to-text doesn’t get a fraction of that attention, and it’s probably touched my actual output more than most of it. Three to four times the words per minute, for a skill you already have. That’s not a gimmick. That’s just a faster way to work, sitting there mostly ignored.

Frequently asked questions

  • Is dictation actually faster once you count corrections?

    Yes, for most writing. You’ll fix the odd misheard word, but that’s a fraction of the time you save typing at 40 words a minute instead of speaking at 120 to 150. It gets slower relative to typing only when the piece needs heavy live restructuring, which is a different problem than raw speed.

  • Do I need a powerful laptop to use speech-to-text well?

    Only if you’re running something local, like Handy, where transcription happens on your own hardware. A cloud tool like VoiceDash processes on its own servers, so it runs at the same speed whether you’re on a five-year-old laptop or a new one.

  • Is a free speech-to-text tool good enough, or do I need to pay?

    Handy is genuinely free and open source with no cap. VoiceDash’s free plan caps you at 1,000 words a month, which is fine for testing but not for daily use, so most people end up on its $12 to $15 a month Pro plan once they rely on it.

Spread the word!