Part 4 · The world in 2026 · 3 of 4
Lean AI: Matching Model Size to Task Difficulty
Small language models can be ~100x cheaper than large ones per conversation; mature teams route by difficulty rather than defaulting to the biggest model.
1 min read
Per-token pricing looks fine at demo scale and gets expensive fast at real scale. As of Q2 2026, industry figures put processing one million conversations through a large frontier model at roughly $15,000–$75,000, versus $150–$800 through a small model — a swing of roughly 100x.
The framing that's replacing the "small vs. large" debate: narrow AI (small, task-specific, embedded in a disciplined data platform) for the high-volume, well-defined slice of work, reserving big AI (general frontier models) for the harder, open-ended fraction that actually needs that range. The Toyota Production System parallel being used in the industry: the goal isn't a smaller model, it's producing more value with fewer resources without sacrificing quality — the same question Taiichi Ohno asked of manufacturing waste seventy years earlier.
Takeaway for a DPM: cost-per-outcome, not raw capability, is increasingly the metric a board wants — treat model choice as a product-tiering decision, not a one-time technical pick.
Source: State of Data Products, Q2 2026 (Modern Data 101).
Where this shows up
Lessons
1 flashcard in Review