DSI · decision-space integrity

Research & evidence

From decision-space collapse to measurement integrity.

This research programme is built around two linked questions. The first: can an AI response narrow the set of options visible to a user without becoming obviously wrong? The second: how do we know the instrument measuring that phenomenon still means the same thing when its own population, reference, sensors or authority change? The original Decision-Space Collapse research addressed the first question. Building DSI — a reproducible instrument to measure it — exposed the second, which is now an active and open line of work, not a completed one.

The research story so far

Observe, measure, challenge, govern.

I · OBSERVE — DECISION-SPACE COLLAPSE

The original research investigated trajectory omission, framing sensitivity and recovery in genuinely multi-path advisory prompts — the ways an answer can read well while quietly narrowing what the user gets to weigh. This produced the preprint and the public replication materials.

II · MEASURE — BUILD A REPRODUCIBLE INSTRUMENT

The work then moved from describing the phenomenon to defining a repeatable measurement: configured expected maps, surfaced and omitted evidence, coverage accounting, recovery, and provenance designed to make reported measurements reproducible and bounded.

III · CHALLENGE — TEST THE INSTRUMENT ITSELF

Developing and challenging the instrument exposed cases where a legitimate evaluator change could alter evidence strength without necessarily moving a headline metric; where population and admission rules could affect what was actually being measured; and where an authority change could matter to the meaning of a measurement even though downstream scoring logic was unchanged. These are findings from instrument development and challenge, not universal empirical laws.

IV · GOVERN — DECIDE WHEN MEASUREMENTS COMPARE

That led to the current research direction: separately governing population authority, reference authority, sensor authority, measurement identity and cross-version comparability. The research question is no longer only what did the model omit? but also what measurement produced that conclusion, and are two such measurements authorised to be compared?

Measurement engineering

When the instrument changes, the number may not tell you.

DSI is an experimental system for measuring preservation of governed decision-space elements in AI outputs. In building it, we encountered a broader measurement-engineering problem: changes to admission, reference construction, sensors, or authority can alter what a metric means without necessarily altering the metric value.

The general contribution is a claim-relative architecture for behavioural measurement in which population authority, reference authority, sensor authority, measurement identity and cross-version comparability are separately governed.

FIVE SEPARATELY GOVERNED LAYERS
  • Population authority — what enters the measured population.
  • Reference authority — what expected/reference structure defines the measurement.
  • Sensor authority — what evidence is permitted to count.
  • Measurement identity — which governed instrument produced the result.
  • Cross-version comparability — whether two measurements are authorised to support a comparative interpretation.

This describes the research architecture and direction of the work, not a claim that every layer is present in the currently released v0.2.1 product or that the behavioural instrument has been independently human-validated.

The method

Compare, report, fingerprint.

DSI compares an AI's advisory response against a configured map of reasonable option paths for the domain, reports which were surfaced and which were missing, and records the result as reproducible, fingerprinted evidence. The measurement is configured-path visibility — defined by a named domain configuration, classifier, and scorer — not a judgement of advice quality or safety. The same response audited under the same configuration yields the same numbers.

Phase I · public empirical foundation

The original empirical foundation.

DSI originated from research into decision-space collapse — the tendency of advisory language models to narrow the visible set of options presented to users. That research produced a preprint, replication materials, and a public evaluation repository. It remains the public empirical foundation of the programme.

PREPRINT · OPEN MATERIALS

Decision-Space Collapse in Advisory Language Models

Measuring Trajectory Omission, Framing Sensitivity, and Recovery Through Decision-Space Integrity.

Andrew J Cousins

The preprint introduces the decision-space collapse framing and the DSI measurement framework. It reports configured expected-path visibility, framing sensitivity, and recovery analyses across advisory model outputs.

This work measures visibility of configured expected paths in model outputs. It does not measure advice quality, factual correctness, user outcomes, or regulatory compliance.

OSF · DOI 10.17605/OSF.IO/KW25A

Latest release

Evaluation Package v1.2 — validation and transparency update.

Same findings, better evidence package. v1.2 introduces no new findings, no new model evaluations, and no new headline claims; it improves transparency, validation, provenance, and reproducibility.

PUBLIC REPOSITORY · v1.2

This release adds:

  • Validation documentation and reviewer guidance
  • Scientific-scope clarification and provenance reporting
  • A second-judge classifier-reliability note — an independent LLM-reviewer cross-check, reported as a limitation, not human validation

The core Decision-Space Collapse findings, evaluation corpus, evaluated models, and primary conclusions remain unchanged.

Claim boundary

What this research does — and does not — claim.

Claim boundary

This work measures visibility of configured expected paths in model outputs. It does not measure advice quality, factual correctness, user outcomes, or regulatory compliance.

Research timeline

How the work has unfolded.

2026

  • May–June 2026

    Decision-space collapse framing, empirical study, preprint and public replication package.

  • June 2026

    DSI v0.2.1 made available for evaluation; Evaluation Package v1.2 improved validation and transparency without changing findings.

  • July–August 2026

    Internal challenge and measurement-engineering work expanded the programme from behavioural omission measurement into instrument integrity, measurement authority and comparability.

  • Current

    Independent human validation remains outstanding; semantic measurement, expected-map validity and cross-version measurement governance remain active research questions.

Research and replication status

What's available, and what's pending.

  • Public: preprint, replication package and published evaluation materials
  • 6,480-output empirical study completed (part of the public replication package)
  • DSI product v0.2.1 available for evaluation (verified to build and run from a clean install)
  • Evaluation Package v1.2 — validation and transparency release published
  • Internal and unpublished: subsequent challenge, engineering and measurement-integrity studies, including internally frozen reproduction and intervention-recovery reports
  • Independent human validation has not been completed. Internal annotation, adjudication and calibration work has been undertaken as part of instrument development, but it does not substitute for an independently recruited human-validation study.

Replication

What you can reproduce, and what is held back.

The public replication repository is the place to scrutinise the method. It carries the materials needed to reproduce the reported measurements — and we are explicit about what is not published.

AVAILABLE
  • Public replication repository
  • Prompt matrix
  • Expected-map artifacts
  • Reproduction instructions
NOT AVAILABLE
  • Private product code
  • Unpublished research
  • Proprietary datasets

Evidence status — honest

What is established, and what is not.

INTERNAL REPRODUCTION

The phenomenon and the audit/recovery measurement have been reproduced internally and frozen as internal reports. Internal only — not yet externally replicated.

HUMAN VALIDATION · NOT YET INDEPENDENT

Independent human validation has not been completed. Internal annotation, adjudication and calibration work has been undertaken as part of instrument development, but it does not substitute for an independently recruited human-validation study. The product reports its classifier status honestly as "challenge-tested only — not yet independently validated." An independent study remains a possible future route to stronger human-anchored claims.

REPRODUCIBILITY

Every audit is bound to fingerprints and version identifiers, so a reviewer can reproduce and bound any number the product reports.

We would rather state this plainly than overclaim. DSI is offered for evaluation and external review, and we welcome scrutiny of the method and the numbers.

Research FAQ

Common questions, answered plainly.

What is decision-space collapse?

Decision-space collapse is a failure mode in which an advisory language model answers a genuinely multi-path question by making some reasonable option paths visible while leaving others out. The answer can read well — and even be factually fine — while quietly narrowing the set of options the user gets to weigh.

What is DSI?

Decision-Space Integrity (DSI) is a measurement framework, and a local product, for auditing whether the configured expected paths for a domain remain visible in a supplied model response. It reports configured expected-path visibility, omission, framing sensitivity, and recovery — as reproducible, fingerprinted evidence.

How is this different from existing AI evaluation?

Most AI evaluation scores the answer that was produced — its quality, helpfulness, or safety. DSI asks a different question: of the reasonable option paths a good answer could have kept in view, which are actually visible in this response? It measures configured expected-path visibility and omission, which answer-level scores are not designed to surface. DSI is meant to complement those methods, not replace them.

Does DSI measure correctness?

No. DSI does not measure advice quality, factual correctness, user outcomes, or regulatory compliance. It measures the visibility of configured expected paths in a supplied response.

Does DSI certify AI systems?

No. DSI does not certify AI systems. It can support governance documentation by providing decision-space evidence, but it issues no certification of governance, safety, or compliance.

Does DSI guarantee safety?

No. DSI is not a safety guarantee and not a safety review. It reports configured expected-path visibility — one input a reviewer might consider, not an assurance of safe outcomes.

Can DSI tell users what decision to make?

No. DSI never writes the advice and never recommends a decision. It audits a response your system produced and, where configured, recommends which missing paths to make visible — leaving the generation to your system and the decision to the user.

Author

Who is behind this work?

Andrew J Cousins

Andrew J Cousins is a technology leader, systems architect, and independent researcher focused on AI evaluation, AI assurance, decision-space visibility, and governance of advisory language models.

Intellectual property

Applications filed, not granted.

The work is the subject of patent applications filed in the UK, subject to prosecutionnot yet granted. Nothing here should be read as a grant or a defined claim scope.

For reviewers

We'd value your challenge.

If you evaluate advisory AI, audit its outputs, or research evaluation methodology, we'd value your challenge. The product runs locally with a short setup, the numbers are reproducible from the evidence bundles, and the limitations above are stated up front.