What "Not evidenced" means, and why it is scored like a gap
Every check in an evidence review ends in one of four states: pass, partial, gap, or not evidenced. The last one surprises people. Here is what it means and why it costs you the same as a gap.
The four outcomes
| Outcome | Meaning | Example |
|---|---|---|
| Pass | The records show the control working at or above threshold | A model version is recorded on 99.9% of decisions |
| Partial | The records show it, below the pass threshold | A reviewer is recorded on 87% of the decisions the review rule covers |
| Gap | The records show the control is missing or not followed | Overrides exist and none carries a reason |
| Not evidenced | Nothing submitted shows the control either way | No answer about an incident process; no reviewer field in the export |
Why it is scored like a gap
Because that is how it is treated in the room. An examiner working from Section 4 of Notice 2024-04 asks for "documentation of compliance." A control that exists but cannot be shown is, for the purpose of that request, a control that does not exist. The 2026 NAIC examiner draft makes the point structurally: its checklist version "added request to provide document name and page #" for each answer. An answer without a document is not an answer.
So in a readiness score, Not evidenced counts against you with the same weight as a Gap. The one difference is in the remediation plan: a Gap needs the control built; Not evidenced often needs only the record produced, which is faster and cheaper, if the control is real.
The two ways it happens
- The practice exists; the record does not. The underwriters do review referrals; the system does not store who reviewed what. The fix is a field, not a policy.
- The answer was "unknown." Nobody on the call could say whether the incident process covers model errors. The fix is to find out, then write it down.
Turning it into a pass
- Make the decision record carry the evidence: decision ID, timestamp, model and version, output, score, reviewer, override and reason, reason codes on adverse outcomes. Every one of those is a field an examiner can count.
- Answer practice questions with the document name, not an adjective. "Yes: AI policy v1.2, approved by the Risk Committee 2026-03-14" is evidence. "Yes, we have a policy" is not.
- Keep the records tamper-evident, so the evidence survives the question "how do we know nobody changed this?" (See how to verify a decision log.)
The Evidence Pack uses exactly this scale: 24 checks, 9 measured from your logs and 15 from your answers, each ending in pass, partial, gap or not evidenced, with the record or the missing record named.
Related
- Pennsylvania Insurance Notice 2024-04: the AI documentation checklist
- The NAIC AI Model Bulletin: what it expects insurers to keep
- How to verify an AI decision log, and prove nobody edited it
- The AI Workflow Evidence Pack: find out which of these you can evidence today, from one decision-log export.