Classical validation was built for a fixed dataset, a metric, a threshold and a sign-off. A generative AI application produces non-deterministic output, is scored by another model, and changes when someone edits a prompt or a vendor updates the model underneath it. The validation question stays the same. The evidence does not.
The question a validator asks does not change: does it do what it is supposed to, and how do we know? What changes is what counts as an answer. For a generative AI application it is five artifacts, produced by the engineering work rather than written for the review: a golden dataset the business owns, evaluation results per version, a release gate, per-call traces, and sampled production review with a measured agreement rate.
A model validator knows the shape of a validation file: the data the model was built on, the metric, the threshold, the signature. Put a generative AI application in front of the same person and every element wobbles. The same input gives a different output tomorrow. “Correct” is a judgement, often made by another model. And the thing under review changes without a release — a prompt edited on Tuesday, a vendor rotating the model version on Thursday.
The easy response is to file generative AI under “tooling”, where nobody asks. The market largely has: Risk.net’s Model Risk Benchmarking 2026, run with Moody’s across 44 banks, found that about a third of banks keep no systematic log of their generative AI use. The better reading is that validation is entirely possible; it consumes different evidence. Here is what that evidence looks like, at the level a second-line reviewer can ask for and recognise.
Each item is an artifact, not a practice. A practice is something a team says it does; an artifact is something a reviewer can open.
What the artifact is. Tens to hundreds of cases drawn from real, anonymised data in the system’s domain, each with the expected outcome: the right answer to a document question and the passage it must be grounded in, or a refusal where refusal is right. Owned by the business, because the business knows the right answer; engineers facilitate. Versioned like code. Split into a development set and a held-out set, so the prompt is not tuned to the test. Grown from production: every wrong answer that reached a person becomes a case.
What it provesThat “correct” was defined before the test, by someone accountable, with a version and a date. The inventory line reads “validated against dataset <version>, <n> cases, last revised <date>” rather than “tested”.
What the artifact is. The scored output of running the golden dataset through one specific version of the system. Metrics come in three families: deterministic (exact match, schema conformance — cheap and hard), semantic (accuracy against the label; for retrieval-based systems the RAG triad — faithfulness to the retrieved sources, context precision, answer relevance — usually scored by a judge model), and safety (the rate of claims no source supports, personal-data leakage, prompt-injection and jailbreak robustness, refusal correctness). Each run records four versions together — dataset, prompt, model, metric — without which the number cannot be reproduced. Thresholds are versioned configuration. Because outputs vary run to run, a threshold is read against repeated runs at a fixed temperature — a pass rate with its interval, not a single pass. For retrieval-based systems the retriever is scored separately from the generator (context recall against the passage the golden case names). The suite runs automatically on every change to a prompt, model, retrieval index or tool definition.
What it provesThat this version was measured against this definition of correct, and the measurement can be repeated. It is a regression suite; a validator reads it like a backtest that runs every time something changes instead of once a year.
What the artifact is. The pipeline record showing that a score below threshold blocks deployment. Prompts and configuration live in version control, not in a tool’s web interface; the pipeline runs the suite; the result is stored with the build; a failing change does not ship. An override needs an approval and leaves a record.
What it provesThe negative claim validators usually cannot get for generative AI: a change could not have reached production below threshold. The override log is part of the evidence — a gate that has never been overridden is either very good or not really in the path.
What the artifact is. A per-call record: the input (or its hash where the data is sensitive), the retrieved context, the tools called and their arguments, the output, the model, prompt, retrieval-index and judge versions, latency and cost. The OpenTelemetry semantic conventions for generative AI give this record an open shape, so it need not be tied to one vendor’s tool.
What it provesThat a reviewer can reconstruct a decision after the fact — which documents the system saw, which tools it used, which model version answered, for a given case on a given day. Without traces an incident review is an exercise in memory. This is the gap the Risk.net logging figure describes.
What the artifact is. A sample of production outputs — random, plus targeted: low confidence, borderline scores, complaints — routed to a queue for human review. The reviews are stored, feed the golden dataset, and yield three numbers: agreement between reviewer and system, override rate, time to review.
What it provesThat the oversight described in the policy operates and has an effect. Risk.net’s benchmark found only about one in five banks measure the effectiveness of human-in-the-loop review. An override rate sitting at zero month after month says reviewers are clicking through — a finding that is visible only if the rate is measured.
None of the five is a document somebody writes. The dataset, the eval results, the gate record, the traces and the review sample are produced by running the system properly. The validation report is a view over them. That is the shift: from documentation as a deliverable to evidence as a by-product.
Most semantic metrics — faithful, context-precise, relevant — are scored by a second model reading the first model’s output against a rubric. Standard practice, and the point where a regulated second line rightly narrows its eyes: AI checking AI?
The answer is that the judge is a model and is treated as one. Its rubric is a versioned prompt. Before it is trusted, a sample of its scores is compared with a domain expert’s ratings and the agreement is recorded as a statistic, not an impression — an inter-rater reliability measure such as Cohen’s kappa for categorical rubrics, or precision, recall and F1 against the expert’s labels — together with the sample size and the date. Its model version is pinned; when that changes, calibration is repeated. The judge is kept independent of the system under test — a different model family where possible, since judges favour their own family’s output — and agreement between two human raters on the same sample is the ceiling any judge score is read against. A record of how closely the judge tracks a human, on which sample, on which date, is what turns “an LLM scored it” into evidence.
For an agent — a system that calls tools and takes steps toward a goal — the final answer is the least informative thing to score. What matters is the trajectory: which tools were called, in what order, how many steps, whether the goal was reached within a step and token budget, whether the agent stayed inside its permitted scope. Trajectory evaluation runs scenarios against mocked or sandboxed tools and asserts on the path: must call this, must never call that, must finish within so many steps.
The same budgets and limits are controls at runtime. The agent’s inventory entry states purpose, permitted tools and limits; the trajectory tests show the limits hold before release; the runtime limits enforce them after. Two layers, one set of rules. The rest of that entry — owner, data, oversight, validation status — is set out in our AI model inventory guide.
Much of classical validation transfers intact: an independent reviewer, the business owning the ground truth, thresholds set before the test, the question of whether the approach fits the use. What breaks is the cadence and the shape of the evidence.
| Classical habit | Generative AI equivalent |
|---|---|
| A fixed test set, built once | A growing golden set, versioned, fed by production incidents, with a held-out portion the prompt has never seen |
| One metric, one threshold | Several families — deterministic, semantic, safety — because a system can be accurate and ungrounded at the same time |
| Annual review, or on material change | Every change runs the suite; a cadence remains for judge calibration and the review sample |
| Documentation written by the validator | Generated evidence assembled from run records; the validator’s writing is the interpretation on top |
| Sign-off by a named person | Unchanged — but the signature is on the artifacts above, not on a description of them |
The failure mode is inheriting the classical checklist unchanged. Several validation checklists in circulation lack groundedness and hallucination tests: the list asks about data quality and stability, as it did for a scorecard, and never whether the answer was supported by the source. If your framework descends from SR 11-7 / SR 26-2, one more caution: the 2026 guidance is non-binding and leaves generative AI outside its scope, so the framework you apply here is the one you choose.
The EU AI Act compliance checklist walks one system through classification and lists the obligations that attach, each with the evidence an assessor usually asks for. Whether the testing described here is a statutory duty or a matter for your own framework depends on where that classification lands.
Open the checklistFive questions for the team that owns a generative AI application. Each has an artifact for an answer; the absence of the artifact is an answer too.
Then attach the answers to the system’s inventory entry — dataset version, last eval result, gate status, trace retention, review cadence and agreement rate. That is where an auditor will look first.
In Model Inventory for Jira, each AI system is a Jira work item: owner, classification, purpose and status on the item, validations and reviews as linked work, and an immutable change history behind it. The dataset version, the last evaluation result and the review cadence sit on the entry where the question will be asked. The judgement stays with your people; what changes is that the decision, the date and the person are recorded when they happen.
See how it worksNo, and you need both. Evaluation is the engineering practice: the dataset, the metrics, the judge, the suite that runs on every change. Validation is the governance act: someone independent of the build examines that evidence, decides whether it supports the claim that the system does what it is supposed to, and signs. Evaluation without validation is testing nobody has accepted. Validation without evaluation is a signature with nothing underneath it.
A curated set of input cases, each with the outcome the system is expected to produce, drawn from real, anonymised data in the domain where the system operates. Owned by the business people who know the right answer, versioned, split into development and held-out parts so the prompt is not tuned to the test, and grown from production incidents.
It depends on the system’s classification, not on the technology. For a high-risk AI system the Act attaches accuracy, robustness and record-keeping duties to the provider, and separate duties to the deployer, so testing and logging are part of what is expected. For a generative AI application that is not high-risk, the Act is not the driver of testing; your own model risk framework, your auditor and your customers are.
This page is a practical explanation, not legal advice. Confirm obligations against the official text of Regulation (EU) 2024/1689, the standards that apply to you, and where the stakes warrant it, qualified counsel.