SGLang
open-source framework for large language model inference
Episodes
-
#5830: Serving Your Own Fine-Tuned Model in the CloudYou fine-tuned an open-weight model. Now how does anyone actually talk to it? Dedicated GPUs vs serverless inference, and the math that decides it. -
#5436: Small Models as Rewriters, Not WritersWhy "don't say X" prompts backfire, and how a tiny grammar-constrained model can scrub a script without breaking its grammar. -
#5410: Adapters: 102KB That Reshapes a 403GB ModelA 102KB adapter file changes how a 403GB base model behaves — without ever merging into it. Here's how model adapters actually work.