Everything here that touches evaluation.
Can a model learn to reply the way I reply? The pipeline, the evaluation harness, and the first real numbers — including the parts where the numbers say not yet.