Skip to content

How to Test an AI Feature Before You Launch It: Evaluations Explained for Business Owners

  Posted on 05 Oct, 2026
  Artificial Intelligence
How to Test an AI Feature Before You Launch It: Evaluations Explained for Business Owners

A demo of an AI feature almost always looks good. Someone types five sensible questions, the answers read well, and the room agrees it is ready. Then real customers arrive with misspelled, half-explained, off-topic or hostile requests, and the feature behaves in ways nobody saw in the meeting.

The discipline that closes that gap is called evaluation, usually shortened to "evals". If you are commissioning custom software development that includes an AI assistant, a document summarizer or an automated reply tool, evals are how you find out whether it works on your cases before your customers do.

This guide explains how to test an AI feature before launch in plain terms: why the usual testing methods fall short, how an evaluation set is built and graded, what to measure, and what to ask your developers for.

Why AI features cannot be tested like normal software

Ordinary software is predictable. The same input produces the same output every time, so a tester can write a rule such as "when a customer adds two items, the total equals the sum of both prices" and the software either passes or fails.

Language models break both halves of that arrangement. OpenAI's evaluation best practices guide points out that models sometimes produce different output from the same input, which makes traditional software testing methods insufficient. There is also rarely one correct answer. A refund reply can be written a hundred acceptable ways and a thousand unacceptable ones.

So the question changes from "does it pass?" to "how often is it good enough across a realistic sample, and how bad are the failures?" That is a measurement, and the acceptable level is a business decision, not a technical one.

What an evaluation set is and how to build one

An evaluation set is a collection of realistic inputs, each paired with a description of what a good response looks like. Think of it as the exam paper for the feature. Anthropic's documentation on building evaluations advises designing tests that mirror the real mix of tasks the feature will face, including edge cases.

The best raw material is already in your business: past support tickets, emails, chat transcripts, submitted forms and the documents the feature will read. OpenAI's guide recommends combining this kind of production data with cases written by people who know the domain. Remove personal data before the cases are used for testing. A useful set contains:

  • Typical requests, in roughly the proportion they really occur
  • Hard cases: ambiguous, incomplete, very long or badly written
  • Out-of-scope requests the feature should decline or hand to a person
  • Deliberate misuse, such as attempts to extract data or override its instructions

Start with a few dozen cases you understand well and grow from there. The expected outcomes should be written by whoever knows what the right answer is, usually a support lead or operations manager, not by the developer alone.

How the answers are graded

Once the feature has answered every case in the set, something has to decide whether each answer was good. There are four methods, and most projects combine them.

Exact checks

Where there is a definite answer, code can check it: the right category was chosen, the invoice number was extracted correctly, the output is in the required format, a forbidden phrase is absent. These checks are fast, cheap and objective, so use them wherever the task allows.

Rubric scoring

Open-ended output needs written criteria. Does the reply agree with the source document? Did it answer the question that was asked? Is the tone right for your brand? Each criterion is scored yes or no, or on a short scale. A rubric turns "it feels fine" into something two reviewers can agree on.

Human review

A person who knows the subject reads the outputs against the rubric. This is the most trustworthy method and the slowest and most expensive, so reserve it for high-stakes cases and for checking the other methods.

Model-graded with spot checks

A second AI model can apply the rubric to far more outputs than a person could read. Both Anthropic and OpenAI describe this approach, and Anthropic's guidance favors a larger number of automatically graded cases over a small hand-graded set. The condition is calibration: OpenAI's guide advises scaling a model grader up only once it consistently agrees with human judgments. In practice, people grade a sample, the results are compared with the model grader, the rubric is corrected where they disagree, and spot checks continue after launch.

What to measure

Accuracy on your cases comes first: the share of cases that meet the rubric, reported by case type. A single average can hide one category that fails badly.

Refusals matter in both directions. A feature that declines legitimate requests frustrates customers, and one that confidently answers questions it should decline, such as legal or medical questions outside your business, creates risk. Test both.

Harmful or off-brand output covers invented policies or prices, leaked data, promises your company cannot keep, and the wrong tone. These failures may be rare but they are the costly ones, so count them separately. It is reasonable to require none at all in the test set for the most serious kinds.

Finally, measure cost and latency. Run the evaluation set and record the cost per request and the response time, including the slowest responses and not only the average. If the feature answers from your own documents, also check that it found the right document; our guide to retrieval-augmented generation for business explains how that part works.

Regression testing when the model or prompt changes

An AI feature can change behavior without any change to your application code. A small edit to the prompt, meaning the instructions given to the model, can fix one case and quietly break three others.

The model underneath will also change. OpenAI's deprecations page states that it regularly retires older models and that software relying on them may need updates. Anthropic's model deprecations page says requests to retired models will fail and commits to at least 60 days' notice before retiring publicly released models. Sooner or later, a model switch is forced on you.

Regression testing means rerunning the whole evaluation set before and after any change to the prompt, the model or the source documents, then comparing scores by category. OpenAI recommends running evaluations on every change and growing the set over time. A good habit: every failure found in production becomes a new test case.

Common mistakes, and when lighter testing is enough

The most common mistake is judging the feature on a handful of friendly demo questions, which OpenAI's guide lists as an anti-pattern under the name "vibe-based evals". Others follow close behind: leaving developers to define what a good answer is, reporting one overall score, trusting a model grader nobody has checked against human review, and treating evaluation as a one-time task before launch.

Not every feature needs the full process. An internal drafting tool, where a person reads and edits every output before it is used, can launch with a smaller set and lighter grading. A feature that replies to customers or takes actions without review needs the rigor described here.

A go / no-go checklist

Before launch, you should be able to answer yes to each of these:

  • The evaluation set is built from real cases, including hard, out-of-scope and misuse cases, with expected outcomes approved by the business
  • The agreed target is met for every case type, with no failures in the most serious category
  • The model grader has been checked against human review
  • Cost per request and response time fit the budget at expected volume
  • There is a fallback: handoff to a person, a way to switch the feature off, and logging for later review

If any item is missing, the honest options are to delay or to run a limited launch, for example to a small group of users or with a person approving each response.

What to do next

Start with the part only your business can do. Collect real examples of the requests the feature will handle, write one sentence for each describing a good response, and decide which failures are unacceptable.

Then ask your developers or vendor three questions: what is in the evaluation set, how are the answers graded, and what happens when the model changes? If the work is outsourced, make the evaluation set and its results a named deliverable, so that you own it and can rerun it later.

Conclusion

AI features cannot be signed off with a demo, because their output varies and there is seldom a single right answer. What replaces the demo is an evaluation set drawn from your real cases, graded by a mix of exact checks, rubrics, human review and checked model grading, and rerun whenever the prompt or model changes.

Entrant Technologies builds websites, web applications, mobile apps and custom software. If you are planning an AI feature and want testing scoped in from the start, you can request a quote and describe the cases it needs to handle.

Entrant Technologies
Post written by
Entrant Technologies is one of the leading web, software, iPhone & Android app development company which deliver robust results for great brands worldwide. We deliver software solutions that meet the customers and business expectations.
View all posts by Entrant Technologies →
Latest Blogs
 
A demo of an AI feature almost always looks good. Someone types five sensible questions, the answers read well, and the room agrees it is ready. Then real customers arrive with misspelled, half-explai ...
on 05 Oct, 2026 Read More
 
Google's Android developer verification requirement reached its first enforcement date on September 30, 2026. From that date, according to Google's developer verification guide, apps must be registere ...
on 04 Oct, 2026 Read More
 
On September 30, 2026, the UK's data protection regulator changed its legal form. The single office of Information Commissioner was replaced by a board-led body called the Information Commission, a ch ...
on 03 Oct, 2026 Read More