Skip to content

ISO 42001 Clause 9 explained: performance evaluation

5 min read by John Bagnall

You can operate an AI management system diligently and still have no idea whether it works. Clause 9, “Performance evaluation,” is the part of ISO 42001 that fixes that. Having run the system in Clause 8, Clause 9 makes you step back and check it: measure it, audit it, and put it in front of leadership for review. This is where governance stops taking its own word for it, and it feeds directly into the improvements you make in Clause 10.

The three parts of Clause 9

Clause 9 is the system’s feedback loop, in three escalating layers:

  • 9.1 Monitoring, measurement, analysis and evaluation (measure it)
  • 9.2 Internal audit (independently check it)
  • 9.3 Management review (leadership judges it)

9.1 Monitoring, measurement, analysis and evaluation

Clause 9.1 requires you to determine what to monitor and measure, how (the methods), and when to measure and evaluate, so the results are valid and comparable. You choose the indicators, but they have to tell you two things: whether you are meeting the AI objectives you set in Clause 6.2, and whether your controls are effective.

For AI, this is where operational evaluation belongs, and where uptime is not enough. Monitoring an AI system means watching its actual behaviour against the metrics that matter, faithfulness, quality, drift, error rates on the cases you care about, not just whether the service is up. The evaluation harness you built before shipping is what makes 9.1 meaningful: you can only alert on a drop if you defined what good looks like in the first place.

9.2 Internal audit

Clause 9.2 requires internal audits at planned intervals to check that the system conforms both to your own requirements and to ISO 42001, and that it is effectively implemented and maintained. The mechanics matter:

  • run an audit programme, not ad-hoc spot checks, with defined frequency and method;
  • set the scope and criteria for each audit;
  • use auditors who are objective and impartial, crucially, they should not audit their own work;
  • report results to the relevant managers; and
  • retain the results as documented information.

The value of internal audit is that it finds problems on your terms. A nonconformity caught by your own audit is a corrective action; the same nonconformity caught by a certification auditor is a finding, and caught by an incident, it is a headline. Impartiality is what makes it work: an audit where people mark their own homework finds nothing.

9.3 Management review

Clause 9.3 brings it back to leadership. At planned intervals, top management must review the AI management system to confirm it remains suitable, adequate and effective. The standard is unusually specific about the inputs and outputs.

Inputs to the review include: the status of actions from previous reviews; changes in external and internal issues relevant to the system; information on performance, including monitoring results, audit results, and whether objectives are being met; feedback from interested parties; and opportunities for improvement.

Outputs are decisions: on improvement opportunities and on any changes the system needs, along with the resources to make them happen.

This is the clause where leadership’s Clause 5 commitment becomes evidenced. A management review that actually examines audit findings, missed objectives and changing context, and produces real decisions, is oversight. A meeting that rubber-stamps a status slide is the governance equivalent of the rubber-stamped AI output: present in name, absent in effect.

Common mistakes

  • Measuring the wrong thing. Tracking uptime and cost but not whether the AI is actually behaving well or the objectives are being met (9.1).
  • No real evaluation baseline. Monitoring with no defined “good,” so a drop in quality is invisible.
  • Auditors marking their own work. Internal audits run by the people who built the thing, defeating the impartiality 9.2 requires.
  • Audit as theatre. Running audits but never acting on the findings, so the same issues recur.
  • A hollow management review. A review that skips the required inputs or produces no decisions, failing 9.3 and undermining Clause 5.

The short version

Clause 9 is how you find out whether the AI management system actually works, rather than assuming it does. You monitor and measure the right things, tied to your objectives and to genuine AI performance, not just uptime (9.1); you audit the system independently and impartially so problems surface on your terms (9.2); and you put it in front of leadership for a real review that examines the evidence and makes decisions (9.3). Together they form the feedback loop that keeps governance honest. Done properly, Clause 9 is where a system proves itself; skipped, it is where a system quietly decays while everyone assumes it is fine. Whatever it surfaces then flows into Clause 10, where you actually improve.

Want your evaluation to hold up to an auditor? The ISO 42001 guide walks the full standard, the checklist turns it into a to-do list, and our ISO 42001 consulting helps you build monitoring, internal audit and management review that mean something. For a quick baseline, the free AI governance check takes about ten minutes, no email.

Previous: Clause 8, operation. Next: Clause 10, improvement.

Frequently asked questions

What is Clause 9 of ISO 42001?

Clause 9, "Performance evaluation", is where you check whether the AI management system is actually working. It has three parts: 9.1 monitoring, measurement, analysis and evaluation (decide what to measure and measure it); 9.2 internal audit (independently check the system against ISO 42001 and against your own requirements); and 9.3 management review (top management formally reviews the system's performance and decides what to change). Together they are the standard's feedback loop, the mechanism that stops governance running on assumption.

What does Clause 9.1 require you to monitor?

Clause 9.1 requires you to determine what needs to be monitored and measured, the methods to use, and when to measure and evaluate the results, so that the results are valid. You choose the indicators, but they should tell you whether the system is meeting its AI objectives (from Clause 6.2) and whether controls are effective. For AI systems specifically, this is where operational evaluation lives: monitoring model and system performance against the metrics that actually matter, not just uptime.

What is an internal audit under ISO 42001?

An internal audit (Clause 9.2) is a planned, independent check that the AI management system conforms both to your own requirements and to ISO 42001, and that it is effectively implemented and maintained. You run an audit programme at planned intervals, define the scope and criteria for each audit, use auditors who are objective and impartial (they should not audit their own work), report results to relevant management, and keep the results as evidence. It is how you find problems before a certification auditor, or an incident, does.

What is a management review in ISO 42001?

A management review (Clause 9.3) is a formal review of the AI management system by top management, at planned intervals, to confirm it is still suitable, adequate and effective. The standard specifies the inputs, including the status of past actions, changes in context, performance and monitoring results, audit results, feedback from interested parties, and improvement opportunities, and the outputs, which are decisions on improvements and any changes the system needs. It is the point where leadership's Clause 5 commitment turns into evidenced oversight.

How is Clause 9 different from Clause 8?

Clause 8 operates the system; Clause 9 evaluates it. In Clause 8 you perform the risk assessments, treatments and impact assessments and run the controls. In Clause 9 you step back and ask whether all of that is actually working: you monitor and measure performance (9.1), audit the system independently (9.2), and have leadership review it and decide on changes (9.3). Clause 8 produces the activity; Clause 9 judges it and feeds the verdict into improvement under Clause 10.


Building something you need to govern?

Start with a fixed-scope AI Opportunity & Risk Audit.

Meet an Expert