Project 04 / Quality · Regression · Observability

AI Evaluation Framework

A practical test harness for retrieval, citations, structured outputs and tool selection across changing prompts and models.

The idea

A small evaluation framework for catching regressions in the parts of an AI application that ordinary unit tests cannot see.

What it explores

  • Golden datasets
  • Retrieval metrics
  • Prompt regression
  • Latency and token tracking

This project is in development. The page will grow alongside the working demo and its evaluation results.