Project 04 / Quality · Regression · Observability
AI Evaluation Framework
A practical test harness for retrieval, citations, structured outputs and tool selection across changing prompts and models.
The idea
A small evaluation framework for catching regressions in the parts of an AI application that ordinary unit tests cannot see.
What it explores
- Golden datasets
- Retrieval metrics
- Prompt regression
- Latency and token tracking
This project is in development. The page will grow alongside the working demo and its evaluation results.