LoRA
Episodes
-
#5533: Why AI Characters Drift Between ShotsWhy does your AI character look like a different person by shot four? The answer is structural — and the fix isn't what you'd expect.Main topic -
#5411: Fine-Tuning at 4-Bit vs 16-Bit: What It Really CostsQLoRA cuts fine-tuning VRAM 15x and cost up to 85% — but you pay in training time, quality, and safety alignment.Main topic -
#5410: Adapters: 102KB That Reshapes a 403GB ModelA 102KB adapter file changes how a 403GB base model behaves — without ever merging into it. Here's how model adapters actually work.Main topic -
#5485: Fine-Tuning Parakeet for Hebrew and Your Own JargonNVIDIA's Parakeet beats Whisper on Android — but can you teach it Hebrew, or just your own jargon? Two answers, one much happier. -
#5408: When Small NLP Models Beat the LLMFeature extraction, fill-mask, token classification — the classic NLP tasks still have a job. Here's when a small model beats a frontier API. -
#5393: Fine-Tuning a Model on 100 Hand-Edited AnswersYou don't need 10,000 examples to make a model sound like you. The real number is closer to 100 — if the edits are opinionated. -
#5467: Small Models, Big Schemas: When JSON Constraints BackfireSmall models plus strict JSON schemas should be a safe bet. A 15,000-generation study found the opposite. -
#5450: Heat, Fans, and the Limits of Air-Gapped SecurityDozens of ways to leak data from air-gapped machines exist — but every real breach still used a USB stick. -
#5416: Xet, Buckets, and Auto-Pulling Model WeightsHugging Face swapped its storage backend to Xet with barely a ripple. So why can't a bucket auto-pull new upstream weights?