H100 GPU
computer chip by NVIDIA
Episodes
-
#5411: Fine-Tuning at 4-Bit vs 16-Bit: What It Really CostsQLoRA cuts fine-tuning VRAM 15x and cost up to 85% — but you pay in training time, quality, and safety alignment. -
#5830: Serving Your Own Fine-Tuned Model in the CloudYou fine-tuned an open-weight model. Now how does anyone actually talk to it? Dedicated GPUs vs serverless inference, and the math that decides it. -
#5624: One Forward Pass: Building a Bounded ClassifierNo generation, no parsing — just label scores in a single pass. How to build a small encoder classifier that actually works. -
#5431: The Other Half of Hugging Face: Why BERT Still Out-Downloads LlamaEncoder models pull over a billion downloads a month. Decoder models pull 397 million. The AI conversation and the download counter are describing ... -
#5412: Editing vs. Note-Taking for AI Fine-TunesHand-editing a model's output gives three training signals at once. Writing notes gives one — and a weaker one at that.