Audio & Speech

Speech recognition, TTS, voice cloning, audio engineering

102 episodes Page 2 of 6

#5386: Morse Code Between Two Phones: Does It Actually Work?

Two phones, no data, one room. Morse over sound and flashlight — what the research says about echoes, timing, and whether anyone will hear you.

audio-processingsignal-processingdata-over-sound

#5384: Android ASR Runtimes: LiteRT, ExecuTorch, and Why Your Phone Has No VRAM

Why does your phone have no VRAM number? A tour of Android's runtime layer and what it takes to run ASR locally.

androidquantizationspeech-recognition

#5378: SIP Speakers, Multicast, and the One-Way Page

How SIP one-way paging actually works — and why multicast turns 500 speakers into a single call.

telecommunicationsnetworkingaudio-engineering

#5328: A Shabbat Radio That Can't Be Paused

A soundproof box that plays podcasts continuously — no pause, no skip, no volume. Just open the lid and listen.

diyadhdhome-lab

#5319: Focus Stacking Without the Phone App Guesswork

Your phone app wants three shots. The optics want three hundred. Here's how to actually stack macro focus.

computational-audioopen-sourcedocker

#5259: When Actors Watch Themselves: Tarantino's Laugh Track

Tarantino laughs at his own films in Tel Aviv theaters. Most actors can't even watch theirs once. What's behind the gap?

parasocial-interactionhuman-factorssensory-processing

#5128: The 120 Hz Problem: Why Your Window Can't Stop That Jackhammer

A bedroom-shattering 120 Hz tone reveals why low-frequency noise ignores windows, mocks active noise cancellation, and demands real physics.

audio-engineeringsignal-processingsound-transmission

#4799: How to Build Your Own TTS Audiobooks

From M4B containers to chapter timing drift — the technical pipeline for creating audiobooks with synthetic voice.

text-to-speechaudio-engineeringandroid

#4668: Why Podcast AI Voices Sound Too Perfect

We dig into why AI podcast voices sound too clean—and how TTS is learning to stumble, overlap, and interrupt convincingly.

text-to-speechmultimodal-aiconversational-ai

#4666: Why 30 Seconds of Audio Beats 3 Minutes for Voice Cloning

Why adding more audio made Daniel's voice clones worse — and what it reveals about how voice embeddings actually work.

voice-cloningaudio-processingspeech-recognition

#4456: Inside the Podcast Pipeline: How 15 Weekly Episodes Get Made

From prompt to published episode — a full walkthrough of the automated production system running 15 shows weekly.

audio-engineeringgpu-accelerationspeech-recognition

#4376: Which Mic Actually Lowers Word Error Rate?

Laptop mics hit 18% WER. A $70 mic drops it to 4%. Here's what actually works for voice-first dictation.

audio-engineeringspeech-recognitionaudio-quality

#4312: Why Speech-to-Text Still Fails at Its Own Name

When OpenAI's Whisper misheard its own name as "Wispr," it revealed why 95% word accuracy still isn't good enough.

speech-recognitionautomatic-speech-recognitionhallucinations

#3446: Where to Clip a Speaker for the Best Sound

Tiny placement changes dramatically alter sound. Learn the physics of where to clip your speaker for the best audio.

audio-engineeringacousticsspeaker-placement

#3189: Drawing the Melody: SSML's Hidden Power

How SSML gives developers narrative control over AI voices — and why ElevenLabs became its center of gravity.

text-to-speechaudio-engineeringconversational-ai

#3020: How Chatterbox Locks Your Voice Clone Across Thousands of Generations

Why most single-shot TTS models drift over time—and how Chatterbox's cached embedding approach solves it.

voice-cloningopen-source-aispeech-to-speech

#2982: Why Your TTS Model Nails "Shabbat" but Not "Keren Hishtalmut

Why multilingual TTS models handle loanwords but fail at niche vocabulary — and what you can do about it.

text-to-speechtokenizationfine-tuning

#2914: Can AI Read the Room? TTS Prosody Explained

Can TTS models truly infer emotion from text, or just mimic patterns? We break down the science of prosody.

text-to-speechspeech-to-speechaudio-processing

#2886: How Acoustic Cameras Catch Honking Drivers

Can an acoustic camera pinpoint one honk in a traffic jam? The tech is real, and fines are being issued.

audio-processingsignal-processingurban-planning

#2781: When Voice AI Features Enable Fraud

Voice AI platforms now let you simulate background noise, hesitation, and natural conversation — and that's a problem.

voice-cloningai-ethicsfinancial-fraud