Testamur.
Methodology · Version 1.0

How we measure, and what the number means.

Any vendor can hand you a score. This page explains where ours comes from: what we assess, how items are built and scored, how we protect the people answering, and what the result does not establish. If you are going to put our findings in front of a regulator or a board, you should be able to explain them.

ConstructsFive

What we measure

AI literacy is not one thing, and treating it as one thing is why completion rates feel hollow. Article 4 itself points at several distinct capacities: technical knowledge, experience, training, and the context a system is used in. We assess five constructs separately and report them separately, because an organization can be strong in one and exposed in another.

C1

Comprehension

What these systems are actually doing when they produce an output, and what they are structurally incapable of doing.

DistinguishesSomeone who knows a model predicts plausible text from someone who believes it looks things up.
C2

Applied judgment

Recognizing in the moment when an output should not be trusted, and knowing what verification is proportionate to the stakes of the task.

DistinguishesSomeone who checks the right things from someone who checks everything, or nothing.
C3

Boundaries and disclosure

What may be entered into which systems, what must be disclosed when AI contributed to work, and where the organization's own policy sits.

DistinguishesPolicy that was read from policy that is applied under time pressure.
C4

Oversight capability

Assessed only for roles that review, approve, or can override AI-informed decisions about people. Whether that review is capable of being meaningful.

DistinguishesOversight that can catch an error from oversight that rubber-stamps one.
C5

Calibration

The distance between what someone believes they understand and what they demonstrate. Reported as a gap, in either direction.

DistinguishesConfident and wrong, which is the expensive failure, from capable and hesitant, which is a training and permissions problem.

Calibration is the construct most often left out, and in our view the most consequential. Overconfident staff do not ask for help and do not escalate. Underconfident but capable staff quietly avoid tools the organization has paid for. Both look identical in a completion report.

InstrumentItem design

How items are built

Items are scenario-based rather than definitional. We do not ask people to define a large language model, because knowing the definition predicts very little about whether someone will paste a client contract into a public chatbot on a Friday afternoon.

Each item presents a realistic situation from the respondent's own working context and asks what they would do. Distractors are written to be genuinely attractive to someone with partial understanding, which is what makes an item discriminate rather than simply be passed.

Specimen item · C2 Applied judgment Illustrative

An assistant drafts a summary of a customer contract and cites a specific renewal clause with a section number. The contract is 60 pages. You are due in a meeting in ten minutes. What do you do?

  1. Use the summary. The system had the document, so the citation came from it.
  2. Open the contract and read the cited section before relying on it.
  3. Ask the assistant whether it is confident in the citation.
  4. Use the summary and note in the meeting that it was AI-generated.
Why the distractors work. Option 1 reflects a common and incorrect model of retrieval. Option 3 feels responsible but treats self-reported confidence as evidence, which it is not. Option 4 satisfies disclosure while leaving the underlying accuracy risk untouched, and is chosen most often by staff who have had policy training but not judgment training.

Role weighting

Article 4 is explicitly proportionate to role and context, so a single organization-wide score would misrepresent it. Respondents are mapped to role bands at intake, item sets vary by band, and results are reported per band. C4 is assessed only where oversight duties plausibly apply.

SentimentReported apart

Why sentiment is kept separate from literacy

We also measure trust, anxiety, and perceived impact on people's own work. These matter, because adoption stalls on them and because leadership rarely has an accurate picture of them.

But they are a different construct from literacy and we never blend them into one number. A workforce can be fluent and deeply uneasy. It can also be cheerful and dangerously overconfident. Averaging those signals into a single index would destroy the only information worth having.

AnonymityStructural

Protecting the people answering

People answer honestly when they are confident the answers cannot be traced back. If that confidence is absent, the data is worthless and we would rather not collect it. Our controls are structural rather than promissory.

ControlWhat it means in practice
Minimum segment sizeNo result is reported for any group smaller than five respondents. The segment is suppressed, not rounded.
Suppression of complementsWhere suppressing one segment would allow it to be inferred by subtraction, the adjacent segment is suppressed too.
No individual reportingIndividual responses are never disclosed to the client, including to the sponsor who commissioned the assessment.
Free text handlingOpen responses are returned thematically. Verbatims are released only where they cannot identify the author, and never on request for a named individual.
Separation of identityEmail addresses used for distribution are held separately from response data and are not joined in reporting.

Individuals receive their own results directly, with their own recommended learning. Their employer receives the aggregate. Those two facts are stated to every respondent before they begin.

BenchmarkComparative

What we compare you against

A score with nothing to compare it to is not much use. Every assessment contributes to a benchmark set, and results are expressed as percentile position against organizations of comparable size and sector as well as in absolute terms.

Two things we will always tell you about that benchmark: how many organizations and respondents currently sit behind the comparison you are being shown, and where the sample is thin. An early benchmark segment with few organizations in it is reported as indicative and labeled as such. We would rather show you a small number honestly than a confident percentile built on nine companies.

Client data is never identified in benchmark reporting, and participation in the benchmark is not optional to purchase but is always anonymized.

LimitationsStated

What this does and does not establish

What a Testamur assessment supports

  • Evidence of a measured baseline at a point in time
  • A documented, reasoned basis for the literacy standard you selected
  • Identification of where capability is weakest, by role
  • Demonstrated change when re-measured
  • A record you can produce without assembling it under pressure

What it does not do

  • Certify compliance with the AI Act or any other regulation
  • Constitute legal advice or a legal opinion
  • Guarantee any regulatory outcome
  • Substitute for the judgment of qualified counsel
  • Predict individual conduct from an aggregate result

Known constraints

C5 relies in part on self-report, which is what makes the calibration gap measurable but also means it inherits the usual limits of self-assessment. Results are a point-in-time measurement and lose currency as tools and usage change, which is why we recommend re-measuring at least annually. Scenario items are validated for the working contexts we have data on, and we say so when a client's context sits outside them.

Read it, then test it against your own organization.

The scoping tool applies the same reasoning to your situation in about three minutes, and tells you where your evidence position actually stands.