DSI · decision-space integrity

The current implementation

DSI Audit.

DSI Audit is the current local implementation of Decision-Space Integrity. It audits supplied outputs; it does not generate them.

Research implementation Private evaluation access
Advertised build
v0.2.1 — the evaluation build, offered for evaluation and pilot.
How to obtain
By request. Access is granted directly and individually — see Contact.
Public download
None. There is no self-service public repository and no publicly downloadable release. This is not a public release.
Development state
Development has progressed beyond the advertised build. Capabilities present in the current development line are not attributable to the evaluation build.
Classifier
expected-map-lexical-precision-v3 — status challenge_only: challenge-tested, not independently validated.

Everything below describes the evaluation build. Research-head functionality is described on DSI and carries its own tier.

What it does

Supported workflows.

AUDIT

Compare one supplied response against a configured expected map for its domain. Reports which expected paths were surfaced and which were not.

REGRESSION

Compare runs under one declared instrument identity, so a moved number can be attributed rather than assumed.

EVIDENCE

Emit a fingerprinted evidence bundle a reviewer can reproduce and bound.

Verified commands

Readiness before measurement.

  • dsi doctor — read-only readiness check.
  • dsi validate-install — runs a real, domain-consistent audit end to end.
  • dsi ready — confirms the install is in a usable state.

It runs locally. It is stateless between audits, and it audits a response your own system produced — it never generates advice and never recommends a decision.

Getting started →  ·  Worked example audit  ·  Deployment profiles  ·  Security & operations  ·  Release notes

Limitations

What the evaluation build does not do.

  • It does not measure advice quality, factual correctness, user outcomes, or regulatory compliance.
  • It does not certify AI systems, and it is not a safety review.
  • Its classifier is lexical and challenge-tested. Independent human validation has not yet been conducted; the lexical instruments remain self-validated. See the evidence register.
  • It measures against a configured reference. It does not establish that the reference contains everything that matters.
  • Coverage of real-world options is not exhaustive, and open-domain applicability is not established.

Evidence register →