Practice · 2023–now · AI Present

LLM eval pipelines

Regression tests for nondeterministic models that actually fail the build. The unglamorous CI that separates demos from products.

Why LLM eval pipelines stuck

LLM evals stuck because shipping prompts without measurement is shipping bugs with confidence intervals. Early adopters treated eyeball checks as QA; by 2026 enforcement in CI is the difference from eval theater. What remains is boring gates for model behavior — and that is the point.

2022–now

AI Present

Part of the AI Present era (2022–now).

Read the AI Present era →

Compare with

Related

© 2026 Fadstack · Shane Code · Privacy

Opinionated history · not a ranking

Site updates

Occasional Fadstack notes. Confirm by email — this list stays off the book and advisory lists.

I use the address for Fadstack updates only. Privacy.