Before You Scale AI: Define What a Good Result Looks Like
An AI demonstration can be impressive without showing whether the tool is ready for everyday business use. A polished summary, quick response or plausible recommendation may conceal omissions that become visible only when someone tries to act on it. The management challenge is to define acceptable performance before adoption expands.
Start with the work the system will support. A tool drafting internal meeting notes needs different checks from one producing information for a customer or recommending a consequential decision. The evaluation should reflect both the purpose of the output and the effects of an error.
Treat quality as a business requirement
NIST released its Generative Artificial Intelligence Profile on 26 July 2024 as a companion to its AI Risk Management Framework. It provides a voluntary resource for identifying and managing risks associated with generative AI. [1] It does not certify that a particular product is suitable for a business task.
For managers, the practical implication is to evaluate the complete working arrangement: the approved data, the task, the output, the review and the action that follows. The testing approach below is a suggested business practice, rather than a NIST certification procedure.
Define the output before choosing the score
Write a short description of a usable result. For a management summary, requirements might include accurate figures, traceable references, clear treatment of uncertainty and no missing material qualifications. For a customer response, the requirements could include consistency with current policy and an appropriate escalation route.
Distinguish unacceptable errors from minor defects. An invented contractual commitment has different consequences from an awkward sentence. Counting both as one error would obscure the information managers need when deciding whether a process is reliable enough.
Agree these requirements with the people who use and check the output. A technically accurate result can still be unusable if it arrives in the wrong format, omits a required approval or assumes information that the organisation cannot lawfully provide.
Test representative work, including difficult cases
Build a permitted test set that reflects the task's real variation. Include ordinary requests, incomplete inputs, conflicting documents and situations in which the correct response is to ask for clarification or decline to make a recommendation.
Use approved, appropriately protected information. Where real records are unsuitable, create clearly labelled fictional examples that preserve the relevant complexity without exposing personal or confidential data. Ensure the test does not become a route for introducing information into an unapproved service.
Compare results against a reference prepared or checked by a qualified person. Repeat a selection of cases to assess consistency. A small pilot can reveal obvious weaknesses, but it cannot establish reliability for every future situation. Record where the evidence remains limited.
Include the cost of human review
Measure the whole task from initial request to accepted output. Drafting time alone can exaggerate the benefit if employees then spend considerable effort checking sources, correcting numbers and rebuilding the structure.
For illustration, suppose a task previously took 40 minutes. An AI-assisted version takes five minutes to draft and 25 minutes to review and correct. The measured time saving is ten minutes, assuming comparable quality and excluding other implementation costs. It is not a 35-minute saving.
Ask reviewers to record recurring corrections. If they repeatedly repair the same omission, the process may need better input instructions, clearer reference material or a narrower task. Extra review should not silently become a permanent substitute for addressing a known weakness.
Decide what happens when the output is uncertain
Assign an accountable owner for the result and specify when specialist review is required. A reviewer needs relevant expertise, access to the underlying evidence and enough time to challenge the output. Adding a human approval step without those conditions offers limited reassurance.
Define a fallback process. If required information is missing, the system changes unexpectedly or quality drops, employees should know how to complete the task using an approved alternative. Record who can suspend use and who decides that the problem has been resolved.
For systems capable of taking actions, separate permission to produce a recommendation from permission to execute it. Spending money, changing records or communicating externally requires controls appropriate to the action and the organisation's policies.
Expand only within the tested scope
Document what the pilot covered, what it excluded and which conditions made its results acceptable. Approval for one task should not be interpreted as approval for unrelated uses or more sensitive information.
Retest when material changes affect the tool, reference documents or workflow. Continue sampling completed work after adoption so that emerging weaknesses are visible. A useful AI system needs ongoing ownership as well as a successful demonstration.
Explore GLOMACS Artificial Intelligence (AI) for Leaders and Managers training course to build the understanding needed to evaluate business applications and lead responsible adoption.
Related Training Courses
View all Training CoursesRelated Training Subjects
Browse Popular Training Subjects
- Certified Training Courses
- Management & Leadership
- Oil and Gas
- Finance & Budgeting
- Project Management (PM)
- Contract Management
- Corporate Governance, Compliance, & Risk Management (GRC)
- Human Resource (HR) Management
- Personal Effectiveness
- Master Classes
- Health, Safety & Environment (HSE)