← Glossary
Eval (Evaluation Harness)
A structured test set used to measure how well an AI system performs on real, representative examples.
An eval is a set of test cases with known correct (or acceptable) answers, run against a model or pipeline to measure accuracy before and after a change — the AI-application equivalent of a test suite. Without one, 'did this get better or worse' after a prompt or model change is just a guess.
Related terms
Where this shows up in practice
Have a project in mind?
Tell us what you're trying to automate or build — we'll reply with next steps, not a sales pitch.