Practice · 2023–now · AI Present

LLM eval pipelines

Regression tests for nondeterministic models. The unglamorous infrastructure that separates demos from products.

LLM evals stuck because shipping prompts without measurement is shipping bugs with confidence intervals. The practice failed early adopters who treated eyeball checks as QA. What remains is boring CI for model behavior — and that is the point.

Context

The stack gets a co-pilot

AI pair programming is already changing how code is written. Agent frameworks and vector stores are still sorting winners from demos. The durable layer will look familiar: evals, retrieval quality, product UX, and ownership. Autopilot rewrites without tests are just big-bang migrations with better slides.

Compare with

Related

© 2026 Fadstack · Shane Code

Opinionated history · not a ranking