The measurement architecture
Decision-Space Integrity.
DSI is the measurement and assurance architecture for comparing a governed reference structure against what an AI output actually made visible — and for governing that comparison so it keeps meaning the same thing over time. DSI is not a Python package, an audit product or a sidecar. Those words may describe a particular release of DSI Audit; they are not the definition of DSI.
The abstraction
From governed structure to evidence.
- Governed source structure
- Effective expected structure
- Observed structure
- Loss · addition · distortion
- Evidence · provenance
The reference is governed and independent of the response. The observed structure is read from the output. The comparison yields what was preserved, what was added and what was distorted — recorded as evidence bound to the identity of the instrument that produced it.
Measurement is not disposition. A measurement says what a comparison found. It does not decide what should happen next, and it confers no authority to act. Keeping those separate is the whole point of the architecture.
Capabilities and their maturity
Each part carries its own tier.
No component inherits maturity from the programme as a whole.
| Capability | Tier | Note |
|---|---|---|
| Governed structural measurement | Research implementation | Implemented and exercised in the governed programme. |
| Expected / effective-map construction | Research implementation | Reference built independently of the response. |
| Audit execution | Research implementation | Runs locally today; access is request-gated. See DSI Audit. |
| Regression comparison | Research implementation | Compares runs under a declared instrument identity. |
| Evidence generation and provenance | Research implementation | Fingerprinted, reproducible evidence bundles. |
| Measurement contract and replay | Research implementation | Not represented in the evaluation build. |
| Measurement identity and comparability | Research implementation | Refuses to compare runs whose governing identities differ. |
| Authority and controlled execution (GATE) | Research implementation | Locally implemented and tested; evidence is locally reported. |
| Integrated measurement + authority operation | Target architecture | The joined system is not established. |
Where it could apply
Application profiles, labelled individually.
These are bounded profiles, not deployments. Each carries its own tier.
Checking whether options a configured reference requires remained visible in a supplied response.
Comparing two runs under one declared instrument identity, with comparability refused when that identity moves.
Whether structure survives a transfer between systems or stages. Not implemented.
Whether retrieved source structure survives into the generated answer. Under investigation, not established.
None of these establish regulatory compliance, safety or advice quality. DSI does not certify AI systems, and it is not a safety review.
Common questions
Answered plainly.
Is DSI a product?
No. DSI is the measurement and assurance architecture. DSI Audit is the current local implementation of it, available for evaluation on request.
Does DSI measure correctness?
No. DSI does not measure advice quality, factual correctness, user outcomes, or regulatory compliance. It measures whether governed expected structure remained visible.
Does DSI certify AI systems?
No. DSI does not certify AI systems. It can support governance documentation by supplying decision-space evidence, but it issues no certification of governance, safety or compliance.
Does DSI guarantee safety?
No. DSI is not a safety guarantee and not a safety review.
Which layer decides what happens to a response?
DSI measures governed structural change. DSI Audit applies that architecture to supplied outputs. GATE governs disposition and authorised execution. None of those layers tells the user what decision to make.
Has the instrument been validated against human judgement?
Independent human validation has not yet been conducted. Internal annotation, adjudication and calibration informed instrument development, but the lexical instruments remain self-validated. These activities do not establish independent agreement or accuracy.