
Your eval score went from 89% to 92%. That's not an improvement.
instagram.comReel
A 3-point improvement in eval scores within a margin of error of ±8 points (at 100 test cases) is statistically insignificant noise. The post outlines 12 key AI evaluation concepts: golden sets and versioned expected behavior, performance slices that reveal where averages hide failures, sample size requirements (roughly 1,000 cases needed to reliably detect 2-point differences), written rubrics for consistent scoring, LLM judges that can score 500 cases cheaply in seconds, judge bias including position bias (10-15 points depending on ordering) and length preference, pairwise comparisons that generate stronger signal than scales, human label calibration on samples not full sets, inter-rater agreement metrics like Cohen's kappa (0.61-0.80 is substantial), groundedness (claims linked to source spans), trajectory scoring for agents (right answer at 3x cost is still failure), and online evaluation against production data. Critical insight: LLM judges are not neutral and exhibit position bias; running evaluations in both orders can produce 10-15 point swings on identical content. The advice: always report confidence intervals rather than raw numbers, and screenshot evaluation grids before claiming improvements.
#ai engineering#llm#prompt engineering#machine learning#evaluation#statistics












































