
Golden-Set Evals
Golden-set evals are automated tests that run an AI agent against a fixed, human-approved set of real customer messages and the answers they should produce. The score shows whether a new prompt, model or knowledge update made the agent better or worse.
The golden set is the reference answer key. A team collects a few hundred representative inputs, real WhatsApp messages and call transcripts with personal data removed, and writes or approves the ideal response for each: the right price, the right policy, the right refusal, the right handoff. Each item may also carry checks such as 'must mention the 14-day return window' or 'must not promise a delivery date'. The set is versioned and kept stable so that scores are comparable over time.
For Arabic this is not optional. A model that answers well in Modern Standard Arabic can still misread a Najdi 'وش' or an Egyptian 'إزاي', mishandle Arabizi, or drift into a register that sounds foreign to a customer in Kuwait. Generic benchmarks do not measure any of that. A golden set built from your own customers' messages, split by dialect, is the only way to know how the agent performs for the people who actually write to you. Nano AI tests every WhatsApp agent against golden sets for Najdi, Hijazi, Emirati, Egyptian and Arabizi before go-live.
In practice the evals run automatically whenever something changes: a new system prompt, a model upgrade, a refreshed product catalogue. Each candidate answer is graded by rules, by a stronger model acting as judge, or by a person for the hard cases, and the pass rate is compared with the previous version. A drop blocks the release. This is how monthly tuning stays safe: it is the difference between 'we improved the prompt' and 'we can show that the prompt improved'.
Related terms
Related services
WhatsApp AI Agents for Businesses in Saudi Arabia & the Gulf
WhatsApp AI agents that answer in your customer's dialect, capture orders, and recover carts — from $199/mo with a monthly revenue report.
Arabic Voice AI Agents: Every Call Answered, Every Booking Captured
Arabic voice AI receptionist that answers clinic calls in your patient's dialect, books appointments, and sends confirmations. From $149/mo per line.
AI Implementation Company — Real Systems That Ship, Not Pilots
Fixed-scope AI implementation for the Gulf: production systems at $10,000-50,000 with acceptance criteria, Arabic-first evals, and a live outcome dashboard.
Looking for Custom Advice?
Let us help you understand and implement these technologies tailored to your business goals.
Book a Discovery Call