Build the evaluation systems, datasets, ground-truth loops, and benchmarks that measure what Shelf actually understands about a person and how reliably that understanding improves.