Generative AI is increasingly being used in products, services and professional work, but evaluating whether it is helping requires more than checking its outputs. This post draws on established approaches to planning, indicators, evaluative rubrics and monitoring, evaluation and learning. A composite consultation example shows how an AI-supported output can be followed into the work and decisions that come next.

Many AI evals focus on whether a system produced an accurate answer, followed its instructions or completed a defined task successfully. Test sets, reference material and rubric-based grading can provide useful evidence about factual support, relevance, completeness and other qualities of an output. Current eval tools also include model-based graders that assign scores or labels against specified criteria. (OpenAI Platform) These checks matter. An inaccurate answer, unsupported analysis or misleading summary is unlikely to help anyone do better work.
But output performance is only part of the evaluative picture. Whether AI is used by an individual practitioner or built into a product, service or organisational process, a satisfactory result at the point of generation may not show what changed around it. An AI-assisted report may be factually sound while giving readers too much confidence in an uncertain conclusion. A product feature may produce accurate responses while leaving users unclear about its limitations or how much reliance to place on it. An AI-supported service may reduce response times while making unusual or higher-risk cases harder to identify and escalate.
A satisfactory output does not show, by itself, that the work improved.
These are not entirely new evaluative problems. Approaches to planning, indicators, rubrics and monitoring, evaluation and learning have developed over decades in fields dealing with contested purposes, partial evidence and different legitimate views of what matters. Generative AI introduces new technical possibilities and risks, but it does not require us to start again. These traditions can help clarify what an AI use is meant to contribute, identify relevant evidence, support judgement and improve the practice.
Start with what the use is meant to contribute
Evaluation begins before measures are selected or outputs tested. It begins with knowing what we are trying to achieve and how the proposed use of AI is expected to help.
Clarifying the intended contribution means asking what should improve around the output, not simply what the system should produce. An AI-enabled product feature may be intended to help users complete a task while understanding the limits of the response and knowing when other support is needed. An AI-supported service may aim to reduce response times without making unusual or higher-risk cases harder to recognise and escalate. In professional work, the intended contribution may be to retain conflicting evidence or reduce routine effort while continuing to support thinking and judgement.
These descriptions provide a reference point for evaluation. Without that clarity, it is easy to measure what is readily available: user numbers, task-completion rates, response times, output volumes or reported time savings. Such evidence may tell us about uptake, activity and efficiency without establishing that the product, service or work has become better. It may also leave assumptions untested, such as whether users recognise uncertainty or human review recovers material softened during automated synthesis.
This is where planning and monitoring, evaluation and learning need to remain connected. Planning clarifies where we are trying to go and why we expect AI to help. Monitoring and evaluation provide evidence about what is happening, while learning enables people to revisit their assumptions and adjust the practice. My material on Monitoring, evaluation and learning approaches this as an ongoing process of reflection, shared interpretation and adaptation.
Move from evidence to judgement
Once the intended contribution is clear, we can consider what evidence would help us judge whether it is occurring.
Indicators can help organise that evidence, but they remain partial and need to be interpreted in relation to purpose and context. Faster completion does not necessarily mean a better service. Increased use may show that a feature is convenient without demonstrating that it is valuable. A high accuracy score may be reassuring without showing whether uncertainty was handled appropriately or people placed more confidence in an answer than the evidence justified.
The aim is not to create a large dashboard for every AI application. It is to identify a small amount of evidence that helps people understand what is happening and make decisions about the use. This reflects the approach developed in Effective indicators for place-based initiatives – indicators need to be connected to purpose, interpreted alongside other evidence and used to support learning and adaptation. They do not provide conclusions on their own.
Some aspects of performance can be measured directly. Others require judgement. A technically accurate response may still handle uncertainty poorly, while a faster service may become less responsive to people whose circumstances do not fit common patterns.
Rubrics are already used in AI evaluation to judge outputs against stated criteria. The further question is what the rubric is designed to evaluate. Where the intended contribution concerns a wider product, service or practice, criteria may need to address whether users understand the limits of a response or unusual cases reach appropriate human support.
Developing such criteria requires more than extending an output-scoring rubric. Teams need to examine real cases and clarify what good performance would mean in the wider setting. My post on Using rubrics to plan and assess complex tasks and behaviours shows how criteria and descriptions of quality can be developed with the people involved and refined as understanding grows.
Together, indicators and rubrics help extend evaluation beyond the immediate output. Indicators point towards evidence of what is happening in the product, service or practice. Rubrics help people judge what that evidence means in relation to the contribution they were seeking.
Follow an issue beyond the output
One way to evaluate an AI-supported practice beyond the immediate output is to follow an important issue through the workflow. This can show whether it remains visible as the output is reviewed, used to frame later discussion and carried into decisions. The following composite example, drawn from recurring patterns in facilitation, synthesis and review and adapted to an AI-supported process, illustrates this approach. It does not describe a single project.
Suppose an AI-generated consultation summary is clear, concise and broadly accurate. It captures the main themes and areas of agreement, with few factual errors. One disagreement, however, is reduced to a short qualification within a larger theme. The summary is then used to prepare the next meeting, where the disagreement does not appear on the agenda, is not discussed further and is absent from the eventual decision record.
An output-level assessment might reasonably conclude that the summary performed well. To assess whether its use supported the wider work, we would also need to follow what happened to the disagreement. We could examine how it appeared in the original material, how it was represented in the synthesis and whether it remained available for later consideration. Much of the evidence may already exist in the notes, generated summary, human revisions, meeting agenda and decision record.
Tracing the issue would not prove that AI caused it to disappear. Human-written summaries have always selected, condensed and sometimes flattened what people said. The original notes may not have captured the disagreement clearly. The AI summary may have softened it, while a human reviewer chose not to restore it. It may have remained visible in the final summary but been omitted from the next agenda. Each finding points towards a different part of the work requiring attention.
The trace might also show that the disagreement remained clearly represented, reached the next discussion and was visibly considered in the eventual decision. That is equally useful evidence. The purpose is not to assume that something has gone wrong, but to understand where important issues are being retained, altered or lost.
The same approach can be adapted to other intended contributions. A service organisation, for example, might follow an unusual case through automated triage, staff review and final resolution to see whether it was recognised and escalated appropriately.
Looking across several consultation cases might eventually support a provisional description such as:
Consequential differences remain visible enough to be understood and considered at the next relevant stage.
This could become a reference point within an evaluative rubric, but only after being tested and discussed in the setting where it would be used.
Who decides what good looks like?
Once evaluation moves beyond the immediate output, the question of who helps define good performance becomes central.
A developer might judge the consultation summary by its factual accuracy, adherence to instructions and consistency of format. A facilitator may notice whether an unresolved issue has been absorbed into a majority theme. Participants may care not only that their views appear in the document, but whether consequential concerns remain available for later consideration. Decision-makers may need to know where conclusions remain contested and how those differences have been handled. Each sees a different part of what good performance involves.
What counts as consequential cannot be determined by frequency alone. A view may matter because it identifies a serious risk, represents people likely to bear the consequences of a decision, challenges a central assumption or draws on knowledge that others do not hold.
Asking people to correct a completed summary may reveal particular omissions. Involving them in deciding what the evaluation should look for may reveal a different set of criteria altogether. This need not become an elaborate exercise. A small trial might involve reviewing a few routine and difficult cases with people who bring different perspectives, then agreeing on one or two provisional criteria for the next round.
Use evaluation to improve the practice
The point of evaluating generative AI use is not simply to determine whether a tool, feature or workflow passes or fails. Findings may lead to changes in an interface, explanations of uncertainty, routes to human support, escalation procedures, review responsibilities or the circumstances in which AI is used. They may also lead people to narrow the intended use, reconsider the assumptions behind it or decide that AI is not appropriate for a particular task.
This is what it looks like when evaluation feeds back into practice rather than ending with a verdict. Evidence is interpreted and used to change the tool, the surrounding work and sometimes the decision to use generative AI at all.
An output eval might tell us that the consultation summary was accurate, readable and broadly representative. To understand whether its use improved the work, we would also need to know whether the disagreement remained visible, reached the next discussion and could be followed through the resulting decision.
The same principle applies across an AI-enabled product, service or professional practice. The task is not to invent evaluation again, but to connect purpose, evidence, judgement and learning closely enough to see whether using AI actually improved the work.
For related material, see Planning, monitoring and evaluation for connecting intended outcomes with evidence and learning; Effective indicators for place-based initiatives for developing useful indicators and interpreting them in context; Rubrics as tools for reflection, learning and evaluation for making evaluative criteria and judgements more explicit; and Monitoring, evaluation and learning for wider resources on reflection, adaptation and learning in complex settings.
[* Image by Jintana / Adobe Stock]