LLaMA
large language model by Meta AI
Episodes
-
#5443: Heads vs Layers: How Model Merging Actually WorksHeads aren't the Lego bricks of model merging — layers are. Here's what heads really do and how frankenmerges get built. -
#5437: Hunting the Tells That Vanish When NamedThere's no tool for finding an LLM's verbal tics — you have to build one. Here's how keyness analysis works. -
#5431: The Other Half of Hugging Face: Why BERT Still Out-Downloads LlamaEncoder models pull over a billion downloads a month. Decoder models pull 397 million. The AI conversation and the download counter are describing ... -
#5393: Fine-Tuning a Model on 100 Hand-Edited AnswersYou don't need 10,000 examples to make a model sound like you. The real number is closer to 100 — if the edits are opinionated. -
#5182: DeepSeek V4.1 Flash: 1M Context, 437x Smaller KV CacheDeepSeek V4.1 Flash landed with a 1M-token window and a KV cache 437x smaller than V1. Here's what actually changed — and why the middle of your co... -
#5830: Serving Your Own Fine-Tuned Model in the CloudYou fine-tuned an open-weight model. Now how does anyone actually talk to it? Dedicated GPUs vs serverless inference, and the math that decides it. -
#5410: Adapters: 102KB That Reshapes a 403GB ModelA 102KB adapter file changes how a 403GB base model behaves — without ever merging into it. Here's how model adapters actually work. -
#5389: Hugging Face vs Kaggle: Where Models Actually LiveHugging Face and Kaggle aren't rivals — one is infrastructure, one is a practice field. Here's how the two platforms actually differ.