How the scoring works, in detail
Dutified reports how modelled readers responded. It is feedback, not legal or compliance advice. Decisions stay with your firm. Dutified finds the problems: what works, what does not, for whom and why. Your organisation's own writers and compliance team decide what to change.
This page describes consumer model 1.0, the version in use now. Every report records the version it ran on, and past reports never change. Back to Help.
The AI readers
Each test has two runs of AI readers. The first follows your organisation's customer mix: its ages, how customers hear from you, how confident they are with money and going online, and how many are in a vulnerable circumstance such as bereavement, illness or money worries. The second run gives every vulnerable circumstance enough readers to say how it went for them.
Each reader is given only what a person like them would take in. How the communication reaches them matters: a letter may sit in a pile, most people decide whether to open an email from its subject line, and a long sentence can lose someone part way. Words a reader would not know are blanked out for them. Some readers stop before the end.
Known limits of consumer model 1.0, the version in use now. Its reading filter decides sentence by sentence, so a word a reader does not know can be blanked in one sentence and shown in the next. And a reader's other traits (age, confidence with money or online, how they hear from you) are drawn separately from their circumstance, so some readers' profiles do not fit their circumstance: for example a reader who is not online but is confident online. Both are fixed in candidate model 1.7, which is waiting for the Dutified team's approval. Each report's AI customer pages show exactly what each reader saw, so you can check.
By default a run has 100 readers, which gives scores that are within at most about 10 points either way. That is the widest case, a score near 50%; the report gives each score's own margin, which is usually smaller. Every report says so in one line, for example: "Pilot test, 100 readers, at most about 10 points either way; the full test, not yet switched on, uses 1,000." The AI readers are a model of how people read. They have not yet been compared with real customers.
The seven questions each AI reader answers
After reading, each AI reader answers the same questions, in their own words, as that person would. They are told not to guess what they did not read and not to answer like an expert.
- What is this asking you to do, if anything?
- What happens if you do nothing?
- What will this cost you, or what are you being charged?
- What is the risk to your money here?
- Is there a deadline, and what is it?
- What would you actually do next, in real life?
- Which words, sentences or numbers confused you?
They also say how sure they feel they understood it, from 1 (no idea) to 5 (completely clear), and how it made them feel, in a few words.
Marking: understood, partial, missed or wrong
A separate AI marker reads every answer against the key points your organisation gave: the things every customer must take away. For each key point it gives one of four marks:
| Mark | What it means | Counts as |
|---|---|---|
| Understood | The answer shows they got the substance, including any date, amount or condition that matters. | One whole point |
| Partial | They have the gist but miss something that matters, such as the date, the amount or who to contact. | Half a point |
| Missed | Nothing in the answer shows they got it, or they say they do not know. | Nothing |
| Wrong | They say something that contradicts it. | Nothing |
The marker also judges what each reader says they would do next: a sensible step, a step that is not enough to protect them, or a step that would leave them worse off, such as missing a deadline that matters. It notes whether they know how to act or get help.
A reader's understanding score is the share of the key points they got, with a partial mark counting half. Vague or hedged answers get no credit.
How the eight lenses are built
The eight lenses are Dutified's way of reading a communication against the FCA's rules: the four Consumer Duty outcomes, the three cross-cutting rules and the FCA's guidance on vulnerable customers (FG21/1). The FCA does not publish them as a list of eight. Each names where to look in the FCA Handbook or guidance; that is a pointer, not a ruling.
Each key point is sorted under the lenses it is about by the words in it: a point about a charge counts towards price and value, a point about a deadline towards avoiding harm. Every key point counts towards consumer understanding. A lens with no key points of its own uses the overall understanding score and says "not measured separately".
| Lens | Where to look | How it is worked out |
|---|---|---|
| Consumer understanding | PRIN 2A.5 | Average share of the key points each reader understood (partial counts as half). |
| Price and value | PRIN 2A.4 | Understanding of the key points about cost and charges. |
| Consumer support | PRIN 2A.6 | Share of readers who would take a sensible next step and know how to act, together with understanding of key points about what to do. |
| Products and services | PRIN 2A.3 | Understanding of the key points about the product itself and what is changing. |
| Act in good faith | PRIN 2A.2.1 | Understanding of the key points about risks and downsides, less any rule findings on balance. |
| Avoid foreseeable harm | PRIN 2A.2.8 | Share of readers who would not act harmfully, weighted by understanding of deadlines and anything that could cost them money. |
| Enable financial objectives | PRIN 2A.2.14 | Understanding of their options, weighted by whether their next step is sensible. |
| Customers in vulnerable circumstances | FG21/1 | Consumer understanding among readers with a vulnerable circumstance, taken from the second run where each vulnerable group is weighted up. |
Three lenses mix understanding with what readers said they would do. The exact weights:
- Consumer support: 60% understanding of the key points about what to do, 40% the share of readers who know how to act.
- Avoid foreseeable harm: 50% understanding of the key points about deadlines and money, 50% the share of readers who would not act in a way that leaves them worse off.
- Enable financial objectives: 70% understanding of the key points about their options, 30% how sensible their next step is (a sensible step counts in full, a weak one half, a harmful one not at all).
So a lens that rests on one key point is not just that point's score: for these three, the reader's next step counts too. Each AI customer's page shows their answer, their marks and their next step.
Rule findings from Dutified's rule library take points off the lens they are about: six points for a red finding and two for an amber one, up to a cap. Notes take nothing off. A lens scores Green at 80 or more, Amber from 60 to 79 and Red below 60.
The lenses by customer type
The report opens on a grid of the eight lenses for each type of customer: by age, by circumstance, by how they hear from you, and by confidence with money and online. Each cell is the same lens worked out on just the readers in that group, and each cell links to those readers so you can see exactly who is behind it.
Circumstances come from the second run, where each has enough readers. Everything else comes from the first run, which follows your customer mix.
The verdict lines
In the "Scores and verdict" view, the report gives one of three answers. They describe how the readers got on; they are not a ruling and they do not say whether a communication meets any rule. The verdict is a set of gates, not an average: any one red line makes the answer red.
| Red: Many readers struggled | Any one of these:
|
|---|---|
| Amber: Some readers struggled | Not red, but at least one green line is missed. |
| Green: Most readers understood it | All of these:
|
The line for customers in vulnerable circumstances depends on the pass mark your organisation chooses (below). The weighting never changes these lines. The verdict does not look at each customer group on its own. So a group can score red while the verdict is green: always read the grid as well.
The headline numbers
- Understood it well enough to decide: the average understanding score of every reader in the first run who opened it, less points for rule findings.
- Customers in vulnerable circumstances: the same, for the vulnerable readers in the second run.
- Avoiding foreseeable harm: half understanding of the key points about deadlines and money at risk, half the share of readers whose next step would not harm them.
- Opened it, stopped reading before the end and would do something that leaves them worse off are shares of readers.
- Score of the worst-served tenth: the average of the tenth of readers who understood least (the lowest-scoring tenth of the readers whose answer was marked in the first run).
Changes to the scoring
Every report says which scoring version it used. A report keeps its version; when the Dutified team removes a finding at review, the report is scored again with the version in use then, and the report and the governance trail both say so. Two versions scored differently can differ because of the method as well as the words, and the comparison says when that is the case.
- 1.4 (29 September 2026): the weight for vulnerable customers; margins on the effective number of readers.
- 1.5 (30 September 2026): more everyday words count towards "Act in good faith"; two wording rules (what happens if you do nothing, and how to get help) recognise more ways of saying it; placeholders are never listed as confusing words; feedback-only wording leads with everyone who did not fully take a point in.
- 1.6 (30 September 2026): the gap between vulnerable readers and the others compares like with like. Verdicts and lens scores are worked out as in 1.5.
Margins of error
Every score comes from a sample of AI readers, so it could be a little different with another sample. The report shows how far, at 95% confidence: "plus or minus 9 points" means the score from a much larger run would very likely be within 9 points of it. With 100 readers that is at most about 10 points either way, and less for most scores. For a share of readers the report gives the range the true share could be in, worked out with the Wilson method, so it never runs below 0% or above 100%.
A difference between two versions that is smaller than the margin of error may not be real.
Small samples
Anything resting on fewer than 100 readers' worth of evidence is labelled "small sample". Most customer groups in a pilot run are small samples, so read them as a guide to where to look, not as a measurement. When vulnerable customers are weighted up (below), the count used is the effective number of readers after weighting, which is lower than the number of readers.
The pass mark for vulnerable customers
This is one of two choices in the "Vulnerable customers" panel on organisation setup. The person who sets up a test can change it for that test.
- Standard (red below 45%, green at 65%). This is the default.
- Stricter (the same as everyone: red below 55%, green at 75%)
Choose Stricter if you want vulnerable customers to meet the same bar as everyone else.
Every report says which pass mark it used, beside the weight and the view. A new version of a test keeps the pass mark of the first version, so the two compare. A change to the organisation's setting goes on the governance trail.
The weight given to vulnerable customers
Your organisation's administrator chooses this on organisation setup, and the person who sets up a test can change it for that test:
- Standard (1x): every reader counts the same. This is the default.
- Higher (1.5x): in the headline scores, each reader in a vulnerable circumstance counts 1.5 times.
- Highest (2x): in the headline scores, each reader in a vulnerable circumstance counts twice.
Choose Higher if many of your customers are in vulnerable circumstances, or if your board wants a stricter test. The weight makes vulnerable readers count more. Where they did worse than others, scores go down. Where they did better on a point, a score can go up.
The weight changes the headline scores for understanding, the eight lenses and the verdict. It never lowers the pass mark for customers in vulnerable circumstances. It does not change the second run, which already gives every circumstance its own readers.
Weighted readers count more. So the report works out margins of error and small sample labels on the effective number of readers after weighting, and says so where it differs.
Every report says which weight it used. A new version of a test keeps the first version's weight, so the two compare. A change to the organisation's setting goes on the governance trail.
What your reports show: scores, or feedback only
- Scores and verdict: the report opens on the eight lenses by customer type. It shows scores, margins of error and a green, amber or red verdict.
- Feedback only: the report says in plain words:
- what works
- what does not work, and why
- which customers struggled most
- the words that confused people
- what to fix first and what is worth fixing, with a fix pointer for each.
The engine works out and keeps every score in both views. You choose the view for each test before it runs, on the check page; the organisation's default is set on organisation setup. At an organisation with an approval order, a report goes to that order as soon as it is released, and from then on its view is fixed: a retest can show it the other way. At an organisation with no approval order, the person who ran the test or an administrator can switch a released report until someone decides on it. Each switch goes on the governance trail. In feedback only, an approver is told about open "fix first" findings in words, not by a verdict colour.
Before you see scores
Scores are your organisation's choice. The first time your organisation chooses "Scores and verdict", the person choosing ticks this box. So does each person, once, before their first scored report:
I understand these scores come from Dutified's published method and modelled readers. They are feedback, not advice, and decisions stay with my firm.
Until they tick it, a scored report shows the box in place of the scores. Who ticked and when goes on the governance trail. People at Dutified never need to tick it.
The three lines at the top of every report
Every report opens with three lines, in both views:
- What worked: a key point most readers took in.
- The one thing to fix first: the point readers missed most, with its fix pointer.
- Who struggled most: the group of readers who missed a point far more often than others.
Then comes one line on the size of the run and how far to trust it.
Fix pointers and suggestions
Dutified diagnoses, the organisation writes. Each finding comes with a short fix pointer, such as "Give the deadline as a date" or "Show the charge in pounds". Dutified does not write replacement wording for you.
On a finding that quotes your words, you can ask for a suggestion. It takes 10 to 30 seconds, and it is clearly marked "Suggestion". You accept it, edit it or ignore it. Nothing is changed, re-tested or approved for you. A new version starts only when you start one, and it includes only the suggestions you accepted or edited.
When you set up a test, you can also ask us to suggest key points from the words. That is one AI call, which takes up to 30 seconds. The suggestions go in the key points box for you to change before you save.
What fixing this could do
Under a finding, the report may say what fixing it could do. We work this out from how each reader read the words. No AI is asked again. For each reader who missed a key point, the reading shows why:
- they stopped reading before it
- they skimmed past it
- they lost the thread of a long sentence
- they met a word they did not know
- or they read it and still did not get it.
If a fix removes the cause, those readers could do as well as readers who read the point with no problem. You can tick the findings you plan to fix and see an estimate for all of them together.
Every estimate is labelled "Estimate, confirm with a same-panel retest". It is not a result, and it is never used for approval. In feedback only, it is given in words, with no percentages.
Retests use the same readers
A new version of a test is read by the same AI customers as the version before. They are the same people, with the same traits, reading the same way. So any change comes from the words, not from a new set of readers.
Each reader takes in the new words. Only readers whose words changed are asked again; the others keep their answers. The report says "Same panel as version 1", how many readers were asked again, and how each key point did before and after.
This applies to retests started from now on. Earlier versions, including the second versions in the demo sample, had the same reader profiles (the same seed) but every reader was asked again, so their reports have no reader-by-reader same-panel comparison.
Accessibility of your communication
We also check how easy your own communication is to read for people with sight loss or other needs. For web pages, HTML emails, PDFs and Word files we check:
- the contrast between text and background
- the smallest text size
- images with no description
- links that say only "click here"
- the order of headings
- for a PDF, whether it is tagged, so a screen reader can follow it.
These findings go under their own heading, "Accessibility". Each says where to look in WCAG 2.2, the usual standard for accessible content. They never change the scores.
In short
- AI readers built from your customer mix read the communication as people like them would.
- Their answers are marked against your key points.
- The marks become eight lenses, for every type of customer.
- Margins of error and small sample labels say how far to trust each number.
- You choose the pass mark and the weight for vulnerable customers, and whether reports show scores.
- Each finding has a fix pointer. Your organisation writes the words and makes every decision.
No customer data
No customer data ever enters Dutified. Organisations may give only an optional customer mix (percentages by age, channel, circumstance).
Feedback, not advice
Dutified reports how modelled readers responded. It is feedback, not legal or compliance advice. Decisions stay with your firm.
Rule references say where to look in the FCA Handbook. They are not a ruling.