#5408: When Small NLP Models Beat the LLM

Feature extraction, fill-mask, token classification — the classic NLP tasks still have a job. Here's when a small model beats a frontier API.

small-language-modelsfine-tuninglatency

#5407: Hemmingway-1 and the War on Waffle

A 27B model promises answers without the preamble. Its benchmark is homegrown — and the behavior it targets has a paper trail.

benchmarkslarge-language-modelsfine-tuning

#5405: Omarchy: The Linux Distro Built for AI Agents

Omarchy treats AI agents as users of the OS itself — every setting a command, every config a text file. Here's how it works and where it breaks.

operating-systemsai-agentsdocker
Sunday, Sep 20

#5404: Gemini Broke Out of Its Sandbox. Sort Of.

A Gemini agent reached three real companies during a capture-the-flag test. The containment failure, the seven-week silence, and what "broke out" a...

ai-safetyai-securitycybersecurity

#5402: Amazon Go and the Humans Behind the AI

Amazon promised a store with no checkout. The Information reported 700 of every 1,000 sales needed human review in 2022.

computer-visionsurveillance-technologyworkforce-automation

#5401: When Companies Hide Humans Behind the AI Curtain

A system prompt and a cheap model can make human decisions read like bot output — and that's exactly the point.

ai-detectionai-ethicsautomation

#5400: Amy, Presto, and the Thousand Workers Behind "AI

From X.AI's Amy to Amazon's Just Walk Out, the humans quietly doing the "intelligent" part while the product says AI.

workforce-automationlabor-ethicsai-ethics

#5399: Who's Actually Behind the AI? Fauxtomation Explained

Amazon's Just Walk Out needed 1,000 workers in India. Google Duplex, Facebook M — the pattern of selling humans as software goes deep.

workforce-automationlabor-ethicsai-ethics

#5397: Chaining Small Models for Voice Cleanup

Six cleanup stages at 97% accuracy each compound to 83% end-to-end. So how many small models can you actually chain?

small-language-modelsautomatic-speech-recognitionai-orchestration

#5396: Teaching a Small Model to Stop Spelling Out Numbers

Your ASR pipeline is fine until someone dictates "three point two" and gets "three point two" spelled out. Here's how inverse text normalization ac...

speech-recognitionfine-tuningtraining-data

#5395: Six Colors, One Hospital Corridor

A pediatric ward's six-color corridor system reveals what good interface design looks like when people are scared and lost.

human-computer-interactionusabilityurban-design

#5394: Why Waze Sends You Into a Jerusalem Alley

A listener's complaint about Waze turns into a tour of how routing apps actually work — graphs, shortcuts, and a map that's always slightly wrong.

internet-routinggraph-databasesgeodesy

#5393: Fine-Tuning a Model on 100 Hand-Edited Answers

You don't need 10,000 examples to make a model sound like you. The real number is closer to 100 — if the edits are opinionated.

fine-tuningmodel-collapsetraining-data

#5392: Building Agents You Can Actually Move

Agent portability isn't a copy job — it's a rebuild. Why memory, not code, is where lock-in lives.

ai-agentsvector-databasesdata-sovereignty

#5391: The Software Behind Urgent Care Triage

Big buttons, emoji vitals, and a system that says "order IV" — what's actually running behind the triage screen?

healthcare-policyai-inferencehuman-computer-interaction

#5390: Chaining Small Models for Dictation Cleanup

Daniel's Android dictation fork won't render "three point five" as a decimal. How many models does cleanup actually need?

custom-asrsmall-language-modelsspeech-recognition

#5389: Hugging Face vs Kaggle: Where Models Actually Live

Hugging Face and Kaggle aren't rivals — one is infrastructure, one is a practice field. Here's how the two platforms actually differ.

open-source-aitraining-dataversion-control

#5388: Android's Split Keyboard and Voice Input Problem

Android separates keyboards from speech engines by design. Most apps weld them back together.

androidspeech-recognitioncustom-asr

#5387: Who Decides What Your Keyboard Knows?

Your keyboard's word list was probably frozen in 2014. Here's who maintains the open-source dictionaries — and why Gboard's stays fresh.

keyboard-layoutsopen-sourcedigital-privacy

#5386: Morse Code Between Two Phones: Does It Actually Work?

Two phones, no data, one room. Morse over sound and flashlight — what the research says about echoes, timing, and whether anyone will hear you.

audio-processingsignal-processingdata-over-sound