Audio & Speech

Speech recognition, TTS, voice cloning, audio engineering

102 episodes RSS Feed

The technology of voice and sound. From text-to-speech systems and voice cloning to speech recognition and audio engineering, this channel covers the cutting edge of how machines learn to speak, listen, and sound convincingly human.

#5573: Audio Post-Processing for AI Podcast Pipelines

Silence trimming, loudness normalization, EQ, and tempo fixes — the CLI tools that turn raw TTS output into a finished episode.

audio-processingaudio-engineeringaudio-quality

#5570: Why Phone Trees Break and AI Voice Agents Can Fix Them

AI is good at routing calls and bad at resolving them. That single distinction explains why menu-replacing voice agents work and human-replacing on...

conversational-ailatencyspeech-to-speech

#5551: Why Learning Makes Us Happier at Any Age

The credential isn't what makes learning feel good — choosing it is. What the research says about self-directed learning, aging, and hard times.

neuroplasticityaudio-processinghealth

#5545: Jet Engines on Bicycles: The Physics of Bad Ideas

A pulsejet strapped to a bike sounds like a fighter jet — and burns 55 gallons an hour to move one guy who could have pedaled.

aerospace-engineeringmechanical-engineeringaviation-technology

#5487: Why Your Spreadsheet Mangles José's Name

A deep dive into character sets, from ASCII to UTF-8, and why José becomes "José" in your spreadsheet.

unicodebidirectional-textdata-integrity

#5485: Fine-Tuning Parakeet for Hebrew and Your Own Jargon

NVIDIA's Parakeet beats Whisper on Android — but can you teach it Hebrew, or just your own jargon? Two answers, one much happier.

fine-tuningspeech-recognitioncustom-asr

#5466: Hebrew Words Hidden in English Text

Daniel wants a classifier that spots Hebrew written in Latin letters — and it turns out nobody's built one.

linguisticsbidirectional-textautomatic-speech-recognition

#5465: Packaging AI Pipelines So They Actually Get Reused

Daniel's podcast pipeline works — so why can't he reuse it? Recipes, containers, and Hebrew code-switching TTS.

text-to-speechdockerdependency-management

#5464: Your Keyboard's Hidden Data Problem

Your keyboard knows your email address, your phrases, your habits — and you can't take any of it with you.

keyboard-layoutsandroidsideloading

#5461: TTS Can't Pronounce Hebrew Inside English

Your TTS reads Hebrew words with English phonetics. Here's why — and why the obvious fix doesn't work yet.

text-to-speechbidirectional-textlarge-language-models

#5460: Four Small Models, One Android Phone: Does It Actually Work?

A chained on-device dictation pipeline — VAD, ASR, cleanup — and why "it feels smooth" isn't the same as knowing it works.

on-device-asrquantizationvoice-to-text

#5459: Rebuilding the Podcast: Turn-Taking, Buttons, and Chatterbox

Daniel wants a push-to-interrupt button and a live voice loop. The research says the button is easy and the live part is a research project.

speech-to-speechhuman-computer-interactiontext-to-speech

#5456: Parakeet vs Whisper: Picking a Phone ASR Model

Why Whisper loses on Android, why Parakeet v2 beat v3, and how to benchmark speech-to-text without any tooling.

speech-recognitionquantizationlocal-inference

#5435: When Your TTS Model Eats the Numbers

Numbers, dates, and acronyms break text-to-speech in specific, documented ways. Here's where normalization lives — and why it depends on your model.

text-to-speechspeech-recognitionautomatic-speech-recognition

#5430: Two Boxes: ASR and the Text Fixer Behind It

Punctuation, casing, ITN, disfluency — the four-job layer between raw ASR output and text you can actually read.

speech-recognitionautomatic-speech-recognitioncustom-asr

#5422: Switching Android Keyboards Without the Tap Dance

One listener wants a one-tap jump between his Parakeet voice keyboard and his typing keyboard. Android says no — unless you know the trick.

androidsideloadingspeech-recognition

#5413: Who Actually Picks Your In-Flight Movie?

Eight in ten passengers use seatback screens — but at Southwest, one person picks everything. Inside the tiny teams behind airline entertainment.

aviationintellectual-propertyinternational-trade

#5397: Chaining Small Models for Voice Cleanup

Six cleanup stages at 97% accuracy each compound to 83% end-to-end. So how many small models can you actually chain?

small-language-modelsautomatic-speech-recognitionai-orchestration

#5396: Teaching a Small Model to Stop Spelling Out Numbers

Your ASR pipeline is fine until someone dictates "three point two" and gets "three point two" spelled out. Here's how inverse text normalization ac...

speech-recognitionfine-tuningtraining-data

#5390: Chaining Small Models for Dictation Cleanup

Daniel's Android dictation fork won't render "three point five" as a decimal. How many models does cleanup actually need?

custom-asrsmall-language-modelsspeech-recognition