Insights
What your ISO 42001 assessor will accept as competence evidence
Picture this for a moment. You have a date in the calendar. You have an AI policy, a training rollout, and a folder of completion certificates. Clause 7.2 of ISO 42001 does not ask what your people attended. It asks for documented evidence that they are competent, and those are not the same thing. It frustrates me how often the two get confused. Here is what actually closes that gap, and what you can realistically do in 90 days.
If ISO 42001 is new to you. It is the international management system standard for artificial intelligence, published in 2023. Unlike most AI guidance it is certifiable, which means an external assessor visits, asks for evidence, and writes findings. That single property is why it reaches organizations that no regulation currently touches. Our requirements page sets out how it compares to everything else asking about AI competence.
The moment it goes wrong
The scene is always the same. The assessor asks for evidence that the people using AI are competent to use it. Someone pulls the training export. There is a short pause while everyone in the room works out that the export answers a different question than the one that was asked.
Nobody did anything wrong. Things are moving very fast, the training was real, the policy was written, the records were kept. The problem is that the clause asks for something different from what the market knows how to produce.
So what does clause 7.2 actually ask for?
Four things, in plain terms:
- Determine the competence needed by the people whose work affects AI performance
- Ensure those people are competent
- Retain documented information as evidence of that competence
- Clause 7.3 adds that people know the policy and what happens when it is ignored
One honest note before going further. ISO does not license free reproduction of its clause text, so that is a summary rather than a quotation, and you should work from a licensed copy before you rely on it. Plenty of template vendors quote the standard freely and should not.
The third bullet is the one that grabs us. It is also the one nobody can satisfy, mainly because all of this is so new. Everything else in this article is about why, and about what to do before your audit.
Here are the four things the market offers instead
Ask any AI assistant what ISO 42001 clause 7.2 requires and you will get roughly the same four categories. It takes thirty seconds and you should do it rather than take my word for it. I mean that literally: go and check.
| What gets produced | What it actually evidences |
|---|---|
| Role and skill profiles | That you defined the requirement |
| Gap assessments | That you noticed a gap |
| Training and completion records | That people attended |
| Peer reviews, risk assessments, audit evaluations | Competence, inferred from something else |
The first three are genuinely required and genuinely useful. None of them is evidence of competence. The fourth is the one that claims to be, and every item in it is an inference from a different activity altogether.
That is the market's working answer to "show me competence." Find a proxy and hope the assessor accepts it. Most of the time, unfortunately, they do.
Why the proxies do not hold
Completion records are attendance with a timestamp, and we all know that does not cut it.
Self-assessment is the one that surprises people. When researchers gave 288 people both a self-report AI literacy scale and an objective scenario-based test, the correlation between what people said about their own ability and what they actually demonstrated ran between r = 0.07 and r = 0.24. Self-report explained under six percent of the variance in demonstrated competence.
That study was run on teachers rather than on an enterprise workforce, and I will not pretend otherwise. No equivalent study exists on a corporate population, which is itself worth sitting with: the method most organizations rely on has never been validated against the workforce they are using it on.
Here is the uncomfortable part. The standard tools organizations reach for when they want to measure AI confidence across a workforce are employee survey platforms. Qualtrics, Culture Amp, Perceptyx. Good products, used for the wrong job. It is 2026 and we are still asking people to rate themselves on a skill that changes every quarter, while some of them are quietly worried about their job security.
Usage telemetry is the other popular substitute. The Kirkpatrick model that most L&D teams work from defines behavior change as adoption rates, prompt frequency and error reduction, and those are easy to pull from a dashboard. They measure whether people use the tools. They do not measure whether people use them well.
Which produces the failure mode worth stating plainly: an organization whose most overconfident staff use AI constantly will score beautifully on adoption. That population is not your success story. It is your next incident report.
There is evidence for that. In a study of 319 knowledge workers across 936 real work episodes, 114 cross-referenced an AI output against another source, and 23 checked the sources it cited. Higher confidence in the AI predicted less critical thinking, not more.
You have met this person. They stand up, they talk fast, they are completely confident, and the room takes what they say as true because nobody wants to admit they are less sure. Nothing gets checked. We are all so keen to feel like we understand AI that we stop asking whether any particular claim about it is actually correct. That person is who this whole article is about.
What "demonstrated competence" would have to mean
Criticism is easy. Here is the specification, derived from the failures above rather than asserted.
Performance, not self-report. Because asking people how good they are explains about six percent of how good they are. We are not going to fix that by asking more politely.
Judgment, not usage. Because adoption metrics reward the population you should be most worried about. Using AI a lot is not an indicator of judgment, and on the evidence above it may be a mild counter-indicator.
By role. Because clause 7.2 scopes to people whose work affects AI performance, and because an assessor will eventually ask why one group got more attention than another. Proportionality is the thing that gets argued about, and it is genuinely hard, because every team uses these tools differently.
Dated and repeatable. Because competence decays quickly as the tools change, and a single measurement has no trend behind it.
And one more that nobody puts in a specification, which is the reason most attempts fail: people have to be willing to answer honestly. If individual results reach their employer, they will not. That is a design constraint, not a courtesy.
What to actually do, with 90 days on the clock
If your audit is a quarter away, you are not going to build a measurement program. Here is what is achievable, in order.
Week one: ask your assessor. In writing, ask what they will accept as competence evidence for your scope. Their answer is the only one that matters, and most organizations never ask. You will also have created a record of having asked, which is worth something on its own.
Week one: write down who is in scope and why. Which roles affect AI performance, and what competence each needs. Most organizations have never done this, it is the foundation of everything else in the clause, and it is a document you can write yourself in an afternoon.
Weeks two to four: keep the training records and stop treating them as the answer. They are required. They are just not sufficient. Filing them under competence evidence is what draws the finding.
Weeks two to eight: get one real measurement, even a small one. A single dated assessment of one population, showing what people actually understood, is worth more in an audit than three years of completion logs. It does not have to cover everybody. It has to exist, be dated, and be honest.
Throughout: be able to explain proportionality. Why this group and not that one. Why this depth. An assessor who disagrees with your judgment will usually accept a documented reason. An assessor who finds no reasoning at all will not.
What not to do: do not buy a template that quotes the standard back at you. You already know what the clause says. That is not the part you are missing.
That fourth step is what we built Testamur to produce, and if it is useful to you, the instrument is here. If you get there another way, the argument in this article still holds.
What this does not solve
Measurement tells you where you stand. It does not make anyone competent, and a vendor implying otherwise is selling you the same promise the LMS did.
There is a harder limit worth admitting, because we ran into it last week. Writing questions that measure judgment rather than recall is genuinely difficult. Our first eight assessment items were written against a documented psychometric blueprint, reviewed, and mechanically checked. The first think-aloud interview found one clean item in eight. Not one wrong in eight. One clean in eight, where the respondent chose the right answer for the reason the item intended.
Three of the items he got right, he got right by working out what the test wanted rather than by reasoning about what he would do at work. That failure is completely invisible in answer data. It scores as a success.
We are rewriting them. That is what the process is for, and it is the reason to be suspicious of anyone who tells you that measuring this is simple. The folder is easy. The measurement is the hard part, which is exactly why the folder is what the market sells you.