A convincing demonstration is not an acceptance test. Enterprise AI needs a repeatable set of tasks with known expectations, including situations where the correct response is to ask a question, decline an action or admit that the evidence is insufficient.
Build a test set from actual work
Ask users for questions they answered last week and the sources they used. Remove personal information that is not needed. Record the expected evidence, acceptable answer and any action that is allowed. Add cases with conflicting documents, missing information and restricted access.
Keep a portion of these tasks hidden from everyday prompt tuning. If you repeatedly optimise against the same small set, you can improve the demonstration without improving real performance. Add new failure cases from production, but preserve a stable set for comparing versions.
Score the parts separately
Retrieval quality asks whether the necessary evidence was found. Answer quality asks whether the response is supported and useful. Permission testing asks whether any restricted material was exposed. Tool testing asks whether the system attempted only the authorised action with the right parameters.
A single average can conceal a serious problem. Report critical permission or action failures separately from minor style issues. For a policy assistant, a short, correct answer with a valid citation is preferable to a fluent paragraph that invents an exception.
Test the awkward operational cases
What happens when the source system times out, the model is unavailable or the user repeats a request? A booking workflow must not create two bookings after a retry. A document assistant should not silently switch to an unsupported answer when retrieval fails.
Measure end-to-end latency under representative concurrency. Include long documents and mixed Arabic-English questions where relevant. Test the system with real user roles and revoked permissions, not only an administrator account that can see everything.
Agree an acceptance record
- Task completion and the proportion accepted without correction.
- Unsupported claims and incorrect or missing citations.
- Access-control and unauthorised-action failures, reported individually.
- Human review time, latency and cost per accepted task.
- A documented fallback, escalation owner and rollback procedure.
The acceptable thresholds depend on the consequences of error. A draft marketing idea and an internal financial workflow need different review levels. Record that difference explicitly instead of applying one score to every use case.
Keep evaluating after launch
Changes to documents, tools, prompts and models can alter behaviour. Run the relevant regression set before each change reaches users, review a sample of live failures and make it easy to report incorrect answers. Evaluation is part of operating the service, not a one-off procurement exercise.
Sources & further reading
Primary references for the technical background and regional statements in this guide. Planning examples and checklists are Novacom’s practical guidance; examples are illustrative unless explicitly identified as project experience.