Skip to content
Golden-Set Evals
core ai

Golden-Set Evals

Golden-set evals are automated tests that run an AI agent against a fixed, human-approved set of real customer messages and the answers they should produce. The score shows whether a new prompt, model or knowledge update made the agent better or worse.

The golden set is the reference answer key. A team collects a few hundred representative inputs, real WhatsApp messages and call transcripts with personal data removed, and writes or approves the ideal response for each: the right price, the right policy, the right refusal, the right handoff. Each item may also carry checks such as 'must mention the 14-day return window' or 'must not promise a delivery date'. The set is versioned and kept stable so that scores are comparable over time.

For Arabic this is not optional. A model that answers well in Modern Standard Arabic can still misread a Najdi 'وش' or an Egyptian 'إزاي', mishandle Arabizi, or drift into a register that sounds foreign to a customer in Kuwait. Generic benchmarks do not measure any of that. A golden set built from your own customers' messages, split by dialect, is the only way to know how the agent performs for the people who actually write to you. Nano AI tests every WhatsApp agent against golden sets for Najdi, Hijazi, Emirati, Egyptian and Arabizi before go-live.

In practice the evals run automatically whenever something changes: a new system prompt, a model upgrade, a refreshed product catalogue. Each candidate answer is graded by rules, by a stronger model acting as judge, or by a person for the hard cases, and the pass rate is compared with the previous version. A drop blocks the release. This is how monthly tuning stays safe: it is the difference between 'we improved the prompt' and 'we can show that the prompt improved'.

Chat on WhatsApp