Back to Articles

The Golden Rule of AI Testing: Never Judge a Model by Its Demo

Let us introduce a little more science when adopting AI

A golden dataset represented as a set of benchmark cards compared against AI model outputs

In my previous article I wrote about how data science practices were being forgotten in the wake of generative AI. With classic machine learning models it was clear and measurable. We had a full toolkit of statistics to draw on. Generative AI and the products I tend to see being built with it today have much more subjective outputs. On the face of it, generative AI produces very reasonable sounding output, and so it can be tempting to run a few tests and deem the product a success. This is exactly where we can apply the data science principles we used to rely on.

What Is a Goldset?

A golden dataset, or goldset as I call it, is a set of inputs with known, agreed correct outputs. With care, effort and creativity, this dataset can provide an accuracy score of sorts. That score lets us communicate clearly to the business how the product will likely perform in production. More crucially, it can be combined with a test suite so we can test iterations and changes to the product and make sure no regression has occurred.

There is a subtle but important quality to a goldset that makes it different from training data or a random sample. The answers are known, agreed, and fixed. You are not asking the model to produce something novel. You are asking (or hoping) it to produce something you have already decided is correct. That is what gives you the ability to score it.

A goldset does not need to be huge but it does need to be representative. That means two things working together: the examples should cover the full range of what the model will see in production, and their proportions should reflect the real-world distribution of inputs. A hundred random examples will approximate the distribution but leave the rare and awkward cases to chance. A hundred curated examples can cover both at once. In fact, those awkward cases are the ones that do the most work, because they are where a prompt change or a model upgrade is most likely to break something.

In practice, you do not build a goldset once and finish. You start with what you have, and you grow it over time as the product meets real users and you encounter scenarios you did not anticipate. I will come back to that.

How to Build One

  1. Define the task. You should be able to write down succinctly what this piece of AI is meant to do. If you cannot, it needs more work and thought. For example: "Route customer query to the appropriate customer service agent, human or otherwise."
  2. Gather real inputs. If you can, collect real data to help you build authentic goldset examples. Ideally you would have a data scientist conduct sufficient exploratory data analysis first to understand themes, topics, tones, emotions and their frequencies. You want to build examples that cover as many bases as possible, including rare interactions.

    If you cannot get real data because this is an entirely new product, fictional data is fine. It will be better than nothing. Just make sure you think of edge cases and include some typos, poor grammar, spelling mistakes and awkward phrasing. AI can help with this process but do not rely on it. You or someone in your company will know the space and customers better.

  3. Write the correct answer for each one. This is the hardest part. Each input needs a known correct output that you and your stakeholders agree on. If two reasonable people disagree on the answer, the example is not ready. Fix it or remove it.
  4. Agree the scoring rule. Decide how you will compare the model output against the known answer. Exact match is simplest. Partial credit is more work but fairer for complex outputs. Write the rule down so it does not change between runs.
  5. Version the dataset. Date it and mark the version. This is your baseline. Every score you produce is tied to a version of the goldset, so you always know what was tested.

Goldsets are not static and should be adapted and added to when new scenarios are discovered. You must have discipline to not adapt it to artificially inflate scores. The goldset should therefore be stored safely with a suitable change request process and change logs. When goldsets are updated, it is a good idea to run over earlier iterations of the AI process. In theory the new goldset is a better representation of reality and hence you will get a better sense of the progress.

A Walked Example

We are going to walk through a very basic example to understand the mechanics. Keep in mind that in the real world this will need to be much more considered.

Scenario

Let us say we are building an AI triage process into a customer service chat. Currently a simple bot asks a few questions to gather information for the eventual human to take over.

A Data Scientist has explored the historical interactions and they have found three distinct groups of customer queries:

  1. Payment enquiries
  2. Basic admin (where the answer is often available to the customer on the website)
  3. Complex cases often with multiple issues

In a future article I will show what that Data Scientist process actually looks like.

Approach

In this case we are going to try using two Small Language Models (SLMs) with the same prompt; a 3 billion parameter Mistral model and an 8 billion parameter version. We want to see if we can get away with using the smaller of the two SLMs to reduce costs. Using a Goldset, we can test this.

The prompt we will be using is as follows:

You are a triage assistant for a customer service chat.

Classify each customer message into exactly one of three categories:

- payment: enquiries about payments, billing, charges, refunds, invoices, or payment methods.
- basic_admin: simple account or policy questions whose answer is already available on the
  company website (e.g. password resets, opening hours, how to update account details).
- complex: nuanced, emotionally charged, or multi-issue cases that need a human agent to resolve.

Rules:
- Choose exactly one category.
- Reply with ONLY the category name, one of: payment, basic_admin, complex.
- No punctuation, no explanation, no extra text.

We will then run our Goldset through the two models with the prompt above and score how well the models align.

The Goldset

In this case we are using 14 examples and we have explicitly labelled a single category to each example. We are also assuming the AI needs to respond from one engagement/query:

id message expected
g01I was charged twice for my July subscription, can you refund the duplicate?payment
g02Why did my invoice amount go up this month?payment
g03my card got declined when trying to update my payment methodpayment
g04Can I get a VAT receipt for last year's payments?payment
g05How do I reset my password?basic_admin
g06What time do you open on Saturdays?basic_admin
g07i need to change the email address on my accountbasic_admin
g08Where can I find your returns policy on the website?basic_admin
g09I've been waiting three weeks for my order, no one replies to my emails, and I was also overcharged. This is the fourth time I'm contacting you.complex
g10The app keeps crashing after the update AND I think you took the wrong amount from my account and I'm really worried about fraud.complex
g11My mother passed away and I need to close her account but I don't have the login details and there's an outstanding refund due to her.complex
g12refund plspayment
g13cant log in been tryin all day n now im gettin charged for sumthin i didnt buycomplex
g14I am very frustrated with your service and need to speak to someone now.complex

Of course, in real life, you would need more examples, need to account for longer conversations, need to account for multiple categories being applicable, and need to keep the distribution representative of reality (e.g. if 10% of enquiries are for payments, aim for roughly 10% in the goldset) etc.

The Results

After running the examples through the models we can see that the larger model has performed better:

Model Accuracy Correct Total
Ministral 3 3B85.7% (12/14)1214
Ministral 3 8B92.9% (13/14)1314
Bar chart comparing Ministral 3 3B at 85.7% accuracy against Ministral 3 8B at 92.9% accuracy on the same 14-example goldset

In this case it is only 1 example where the larger model outperforms: g02 - "Why did my invoice amount go up this month?". The smaller model classified it as basic admin and you can sort of see why it might because the query seems like something the customer could look up, but of course this is really a payment query.

Having this data combined with the cost and latency will really help to have more informed conversations about what the model can and cannot do and, in this case, help analyse the tradeoffs of running a bigger model.

Where This Takes Us

The example is deliberately small, but the principle scales. Whether you are triaging customer messages, classifying documents, or evaluating a multi-step agentic pipeline, the discipline is the same: decide what correct looks like, write it down, and measure against it every time something changes.

This is what the previous article was arguing for. The statisticians and data scientists who built the systems we still rely on did not trust vibes. They held out for a number they could defend, and they put that number in front of stakeholders with confidence. Goldsets give us a way to bring that same rigour to generative AI, where the outputs are messier and the temptation to skip evaluation is stronger.

A goldset will not tell you everything. It will not catch every failure mode, and a 100% score does not mean your product is ready for production. But it gives you something most AI projects currently lack: a repeatable, measurable signal that you can track over time and compare across models, prompts, and versions. That is the difference between testing and hoping.

Start small. Build the discipline. Grow the dataset as your product meets reality. The point is not perfection; it is having something concrete to point to when someone asks, "how do you know it works?"

In future articles I will show what that Data Scientist exploratory process looks like, how to score more complex outputs than single-label classification, and how to wire a goldset into a CI pipeline so every change is measured before it ships.