← All Tags

#llm-as-a-judge

7 episodes

#4444: Testing the Unpredictable: QA for Agentic AI

How QA adapts when your AI system gives different answers to the same question every time.

ai-agentsai-safetyllm-as-a-judge

#2640: Why Instructional Models Beat Conversational for Batch AI

Beyond cheaper tokens—how batch inference changes AI workflows and why instructional models beat conversational ones for automated jobs.

llm-as-a-judgebatch-inferenceinstruction-following

#2405: LLM Benchmarks Are Full of Noise: Statistical Rigor in AI Evals

Why most benchmark claims in AI are statistically indefensible — and what to do about it.

benchmarksinterpretabilityllm-as-a-judge

#2007: AI Grading AI: The Snake Eating Its Tail

We asked an AI to write this script. Then we asked another AI to grade it. Here’s what happens when the judges have biases.

llm-as-a-judgehallucinationsai-ethics

#2006: How Do You Measure an LLM's "Soul"?

Traditional benchmarks can't measure tone or empathy. Here's how to evaluate if an AI model truly "gets it right."

llm-as-a-judgeai-ethicsai-safety

#2005: Beyond Vibes: The Hard Science of LLM Evaluation

Running the same LLM on different GPUs can produce different results. Here’s why that happens and how to test for it.

llm-as-a-judgeragcontext-window

#81: When AI Judges Can't Tell Humans from Bots

Can a robot tell if you’re human? Herman and Corn explore the "Reverse Turing Test" and why being "messy" might be our best defense.

large-language-modelsllm-as-a-judgeai-detection