← All Tags

#speech-recognition

65 episodes

#5485: Fine-Tuning Parakeet for Hebrew and Your Own Jargon

NVIDIA's Parakeet beats Whisper on Android — but can you teach it Hebrew, or just your own jargon? Two answers, one much happier.

fine-tuningspeech-recognitioncustom-asr

#5456: Parakeet vs Whisper: Picking a Phone ASR Model

Why Whisper loses on Android, why Parakeet v2 beat v3, and how to benchmark speech-to-text without any tooling.

speech-recognitionquantizationlocal-inference

#5435: When Your TTS Model Eats the Numbers

Numbers, dates, and acronyms break text-to-speech in specific, documented ways. Here's where normalization lives — and why it depends on your model.

text-to-speechspeech-recognitionautomatic-speech-recognition

#5430: Two Boxes: ASR and the Text Fixer Behind It

Punctuation, casing, ITN, disfluency — the four-job layer between raw ASR output and text you can actually read.

speech-recognitionautomatic-speech-recognitioncustom-asr

#5422: Switching Android Keyboards Without the Tap Dance

One listener wants a one-tap jump between his Parakeet voice keyboard and his typing keyboard. Android says no — unless you know the trick.

androidsideloadingspeech-recognition

#5420: How Keyboard Middleware Actually Works

A keyboard isn't one thing — it's a stack. Here's how input methods, composition buffers, and hooks really work on Linux, Android, and Windows.

keyboard-layoutshuman-computer-interactionspeech-recognition

#5419: The Attention Budget in Your Pocket

Why phone dictation runs out of room, and how bounded attention windows buy you punctuation without blowing your memory budget.

speech-recognitionlatencyedge-computing

#5396: Teaching a Small Model to Stop Spelling Out Numbers

Your ASR pipeline is fine until someone dictates "three point two" and gets "three point two" spelled out. Here's how inverse text normalization ac...

speech-recognitionfine-tuningtraining-data

#5390: Chaining Small Models for Dictation Cleanup

Daniel's Android dictation fork won't render "three point five" as a decimal. How many models does cleanup actually need?

custom-asrsmall-language-modelsspeech-recognition

#5388: Android's Split Keyboard and Voice Input Problem

Android separates keyboards from speech engines by design. Most apps weld them back together.

androidspeech-recognitioncustom-asr

#5384: Android ASR Runtimes: LiteRT, ExecuTorch, and Why Your Phone Has No VRAM

Why does your phone have no VRAM number? A tour of Android's runtime layer and what it takes to run ASR locally.

androidquantizationspeech-recognition

#5219: Speechreading: What Lips Can and Can't Tell Us

Lips carry only 30–40% of speech sounds. What speechreading actually is, how the brain fuses sight and sound, and why forensic lip reading fails.

speech-recognitionsensory-processinghuman-factors

#4958: Multi-Model Transcription with LLM Reconciliation

Using multiple transcription engines and an LLM judge to catch errors in voice-to-text — especially tricky technical terms.

speech-recognitionautomatic-speech-recognitionllm-as-a-judge

#4920: Local Dictation on Android: The Real Bottlenecks

Why on-device speech-to-text on Android hits a wall at 30 seconds — and what silicon actually matters.

speech-recognitionlocal-aiandroid

#4919: Parakeet vs Whisper: On-Device Dictation Showdown

NVIDIA Parakeet beats Whisper on phone dictation despite having more parameters. Here's why architecture matters more than model size.

speech-recognitionautomatic-speech-recognitionaudio-processing

#4717: Building a Home Whisper Server: The Full Spec

Daniel wants a dedicated home server for Whisper dictation. Here's the exact build — GPU, RAM, storage, and runtime — to make it work.

speech-recognitionlocal-aigpu-acceleration

#4667: How Transformers Killed the Robot Voice

From espeak's robotic squawk to neural voices with added "ums" — how transformers made speech synthesis human.

text-to-speechtransformersspeech-recognition

#4666: Why 30 Seconds of Audio Beats 3 Minutes for Voice Cloning

Why adding more audio made Daniel's voice clones worse — and what it reveals about how voice embeddings actually work.

voice-cloningaudio-processingspeech-recognition

#4641: Mozilla's Hidden Projects: Beyond the Browser

Mozilla is more than just Firefox. Discover Common Voice, Monitor, Thunderbird, and Rust — the public infrastructure projects shaping the open web.

open-sourcespeech-recognitiondigital-privacy

#4497: Voice-First Stack Consolidation

Two transcription apps, two text expanders, no sync — and snippets about to fight each other. The path forward.

voice-firstspeech-recognitiontext-to-speech

#4456: Inside the Podcast Pipeline: How 15 Weekly Episodes Get Made

From prompt to published episode — a full walkthrough of the automated production system running 15 shows weekly.

audio-engineeringgpu-accelerationspeech-recognition

#4376: Which Mic Actually Lowers Word Error Rate?

Laptop mics hit 18% WER. A $70 mic drops it to 4%. Here's what actually works for voice-first dictation.

audio-engineeringspeech-recognitionaudio-quality

#4312: Why Speech-to-Text Still Fails at Its Own Name

When OpenAI's Whisper misheard its own name as "Wispr," it revealed why 95% word accuracy still isn't good enough.

speech-recognitionautomatic-speech-recognitionhallucinations

#3854: From Coos to Conversation: Baby's Hidden On-Ramp

How do babies go from babbling to real back-and-forth dialogue? The hidden architecture of early conversation.

child-developmentspeech-recognitionneurodivergence