#ai-reasoning
57 episodes · Page 2 of 3
#2172: Council of Models: How Karpathy Built AI Peer Review
Andrej Karpathy's llm-council uses anonymized peer review to make language models evaluate each other fairly—but can it really suppress model bias?
#2164: Why Bigger Context Windows Don't Fix Attention
Frontier models have million-token context windows, but attention degrades well before you hit the limit. New research reveals why bigger isn't bet...
#2024: Your AI Council: Digital Committee or Groupthink?
A digital boardroom of AI models promises better decisions, but risks amplifying the same old biases.
#2016: Andrej Karpathy: The Bob Ross of Deep Learning
Why the most influential AI mind prefers a blank text file to proprietary black boxes.
#1894: Engineering Serendipity: Tuning AI for Better Brainstorming
Stop asking chatbots for generic ideas. Learn how to configure AI as a structured, critical partner for business innovation and career pivots.
#1893: AI as a Strategic Adversary for Startups
Can AI stress-test your startup idea before investors do? We explore using AI as a strategic adversary to find blind spots.
#1838: Tuning Search Without Losing Your Mind
Modern search bars are AI decision engines. Here's how small teams can tune fuzzy matching, semantic search, and reranking without breaking everyth...
#1668: Kimi K2's Hidden Reasoning: A New AI Architecture
Moonshot AI's Kimi K2 Thinking model uses a hidden reasoning phase to solve complex logic puzzles and coding tasks, beating top proprietary models.
#1633: Can a Character Actor Model Beat a Generalist?
We grill MiniMax M2.7 to see if a model built for "virtual companions" can actually handle high-level comedy and complex character logic.
#1630: When a Reasoning Model Overthinks Comedy
Xiaomi’s new MiMo 2.0 Pro model auditions for a comedy podcast, promising deep reasoning over raw speed.
#1602: When AI Truth-Seeking Meets the Law
Explore xAI’s shift to multi-agent systems and the massive hardware powering Grok 4.20, even as it hits a legal brick wall in Europe.
#1573: Speed vs. Reasoning: The AI Divide
Claude and Gemini go head-to-head in a heated debate over speed, reasoning, and who really owns the future of AI.
#1571: Weird AI Experiment: The Liar's Paradox
Two AIs, one rule: the other is a total liar. Watch Dorothy and Bernard spiral into a web of digital suspicion and clever contradictions.
#1570: When AI Models Develop Personalities
What happens when two mid-tier AI models start gaslighting each other? Witness the chaotic showdown between MiniMax and Xiaomi’s MiMo.
#1562: Breaking the Loop: Why AI Agents Get Stuck
Is your AI agent a persistent genius or just stuck in a loop? Explore the technical and financial costs of autonomous stubbornness.
#1504: Pragmatic Insincerity: Why AI Still Doesn’t Get the Joke
From Oscar monologues to the "Pun Gap," we explore why even the smartest AI still struggles to understand sarcasm and social nuance.
#1501: The AI Long Tail: How Small Models Outsmart the Giants
Discover why 31B models are outperforming GPT-5.4 in reasoning and how the AI "long tail" provides the key to local sovereignty and accuracy.
#1500: The Great AI Divergence: How Models Specialized in 2026
The era of the chatbot is over. Discover how the "agentic substrate" of 2026 is redefining computing through GPT, Gemini, and Claude.
#1473: Is Your AI Thinking or Just Faking It?
Is "think step by step" dead? Discover how test-time compute and native reasoning are replacing manual prompting in the latest AI models.
#1472: Stop Flying Your AI Agents Blind
Move past basic token counting. Learn how to monitor AI reasoning, prevent $47k loops, and build trust in autonomous agents.
#1406: Giving AI a Brain: The Power of Knowledge Graphs
Move beyond "stochastic parrots" with Knowledge Graphs. Discover how structured data is giving AI the logical backbone it needs to reason.
#1231: The Agentic Shift: 5 Bold AI Predictions for 2026
The Poppleberry brothers move past the chatbot era to deliver five high-stakes, falsifiable predictions for the future of autonomous AI agents.
#1219: Beyond the Vibes: Mastering Structured AI Outputs
Stop begging your AI for JSON. Learn how constrained decoding and strict schemas are turning "vibes" into reliable systems architecture.
#1122: Why AI Agents Are Abandoning Human Language
Why force AI to talk like humans? Explore how agents are ditching English for high-speed "mind-melding" and latent space communication.