Any vendor can hand you a score. This page explains where ours comes from: what we assess, how items are built and scored, how we protect the people answering, and what the result does not establish. If you are going to put our findings in front of a regulator or a board, you should be able to explain them.
AI literacy is not one thing, and treating it as one thing is why completion rates feel hollow. Article 4 itself points at several distinct capacities: technical knowledge, experience, training, and the context a system is used in. We assess five constructs separately and report them separately, because an organization can be strong in one and exposed in another.
What these systems are actually doing when they produce an output, and what they are structurally incapable of doing.
Recognizing in the moment when an output should not be trusted, and knowing what verification is proportionate to the stakes of the task.
What may be entered into which systems, what must be disclosed when AI contributed to work, and where the organization's own policy sits.
Assessed only for roles that review, approve, or can override AI-informed decisions about people. Whether that review is capable of being meaningful.
The distance between what someone believes they understand and what they demonstrate. Reported as a gap, in either direction.
Calibration is the construct most often left out, and in our view the most consequential. Overconfident staff do not ask for help and do not escalate. Underconfident but capable staff quietly avoid tools the organization has paid for. Both look identical in a completion report.
Items are scenario-based rather than definitional. We do not ask people to define a large language model, because knowing the definition predicts very little about whether someone will paste a client contract into a public chatbot on a Friday afternoon.
Each item presents a realistic situation from the respondent's own working context and asks what they would do. Distractors are written to be genuinely attractive to someone with partial understanding, which is what makes an item discriminate rather than simply be passed.
An assistant drafts a summary of a customer contract and cites a specific renewal clause with a section number. The contract is 60 pages. You are due in a meeting in ten minutes. What do you do?
Article 4 is explicitly proportionate to role and context, so a single organization-wide score would misrepresent it. Respondents are mapped to role bands at intake, item sets vary by band, and results are reported per band. C4 is assessed only where oversight duties plausibly apply.
We also measure trust, anxiety, and perceived impact on people's own work. These matter, because adoption stalls on them and because leadership rarely has an accurate picture of them.
But they are a different construct from literacy and we never blend them into one number. A workforce can be fluent and deeply uneasy. It can also be cheerful and dangerously overconfident. Averaging those signals into a single index would destroy the only information worth having.
People answer honestly when they are confident the answers cannot be traced back. If that confidence is absent, the data is worthless and we would rather not collect it. Our controls are structural rather than promissory.
| Control | What it means in practice |
|---|---|
| Minimum segment size | No result is reported for any group smaller than five respondents. The segment is suppressed, not rounded. |
| Suppression of complements | Where suppressing one segment would allow it to be inferred by subtraction, the adjacent segment is suppressed too. |
| No individual reporting | Individual responses are never disclosed to the client, including to the sponsor who commissioned the assessment. |
| Free text handling | Open responses are returned thematically. Verbatims are released only where they cannot identify the author, and never on request for a named individual. |
| Separation of identity | Email addresses used for distribution are held separately from response data and are not joined in reporting. |
Individuals receive their own results directly, with their own recommended learning. Their employer receives the aggregate. Those two facts are stated to every respondent before they begin.
A score with nothing to compare it to is not much use. Every assessment contributes to a benchmark set, and results are expressed as percentile position against organizations of comparable size and sector as well as in absolute terms.
Two things we will always tell you about that benchmark: how many organizations and respondents currently sit behind the comparison you are being shown, and where the sample is thin. An early benchmark segment with few organizations in it is reported as indicative and labeled as such. We would rather show you a small number honestly than a confident percentile built on nine companies.
Client data is never identified in benchmark reporting, and participation in the benchmark is not optional to purchase but is always anonymized.
C5 relies in part on self-report, which is what makes the calibration gap measurable but also means it inherits the usual limits of self-assessment. Results are a point-in-time measurement and lose currency as tools and usage change, which is why we recommend re-measuring at least annually. Scenario items are validated for the working contexts we have data on, and we say so when a client's context sits outside them.
The scoping tool applies the same reasoning to your situation in about three minutes, and tells you where your evidence position actually stands.