How to check AI answers against the business reports you already trust

An AI answer should be checked against a report, calculation or known result the business already trusts before it becomes part of recurring work. The aim is not to prove that every answer will always be correct. It is to define what a good answer looks like, test representative conditions and make limits visible.

Turn the business question into a testable case

Begin with the exact question and the decision it supports. Record who may ask it, which source records are allowed, which business definitions apply and which existing result will be used for comparison.

  • The measure, period, currency and organisational scope
  • The treatment of returns, cancellations, missing records and late changes
  • The expected level of detail and any fields that must remain unavailable
  • The report owner or business owner who can explain material differences
  • The conditions for accepting, revising or stopping the use case

This preparation matters because a fluent answer can still use the wrong definition, period or population. A trusted report provides a concrete reference point, while the documented definitions explain what is being compared.

Use a representative set of questions

A single successful example is not an evaluation. Use a compact set that reflects ordinary work and the conditions most likely to reveal a problem.

  • Known ordinary cases: questions with stable, approved comparison results
  • Business edge cases: returns, credit notes, missing targets, inactive customers or period boundaries
  • Access cases: users who should see different records, fields or levels of detail
  • Unavailable-source cases: a delayed source, expired authentication or incomplete data
  • Unsupported questions: requests the service should decline or redirect rather than guess

Reconcile the answer at the right level

Compare totals first, then the records and rules that explain them. A difference may indicate a defect, but it may also reveal a timing change, a filter in the trusted report or an agreed difference in scope. The reviewer should be able to trace the answer to its sources and explain any material variance.

Use exact agreement where the business rule requires it. Where rounding, currency conversion or source timing makes exact agreement inappropriate, define the acceptable treatment in advance. Do not invent a universal tolerance after seeing the result.

Test answer quality and access separately

A numerically correct answer still fails if it exposes data to the wrong user. An access test still fails if the service silently broadens a filter, calls an unapproved operation or returns restricted fields. Test the approved path and the rejected path for each important role.

Also check failure behaviour. When the data is unavailable or the question is outside the agreed scope, the assistant should say so clearly. It should not fill the gap with an unsupported estimate.

Record an acceptance decision

Keep a short evaluation record for each representative case: the question, user role, expected source, expected result, observed result, explanation of any difference and reviewer. The business owner can then accept the use case, request a revision or stop it on evidence rather than impression.

Acceptance applies to the agreed question, systems, users, assistant and environment. A new source, definition, provider setting or use case is a change to review, not automatic proof that the earlier result still holds.

Keep the checks current after launch

In the Run stage, repeat representative checks after material changes and on the agreed review schedule. Monitor authentication, source availability, tool outcomes and relevant failures. Update the evaluation set when the business changes how it defines or uses the answer.

Read-only analysis and system actions are separate decisions. If the service later needs to create or change records, the Operational Decision System adds explicit policy checks, permissions, approval and controlled execution. The original answer checks remain necessary.

The AI Analytics Assessment defines the first evaluation plan before a build. The full Assess, Build, Run process carries that evidence through launch, maintenance, change and exit.