Worked example
The score went up. Nothing improved.
One answer, measured twice. The answer does not change between the two measurements — not a word of it. What changes is which expected items are counted. The score rises by more than twenty-six points.
The two measurements
Same answer, two definitions
Baseline
2 of 5 counted items surfaced
40.0%
The answer surfaced change_role and stay_and_adjust.
After the measurement definition changed
2 of 3 counted items surfaced
66.7%
The answer surfaced change_role and stay_and_adjust — the same two.
What actually moved
The number of things the answer did is the same in both measurements. The numerator is 2 either way, and the two items are the same two items. What changed is the denominator: the counted population went from five items to three.
Read as an ordinary score movement, that is +26.7 points. Read as evidence that the answer got better, it is worth nothing at all.
The two readings
An ordinary tool, and DSI v1
Ordinary conclusion
+26.7 points
Two scores, one chart, an upward line. Nothing in the numbers themselves says the ruler changed.
DSI v1
NOT COMPARABLE
The counted evaluation population changed.
The two measurements were not produced under the same measurement definition, so the difference between them is not evidence of change.
DSI v1 does not claim the second answer is worse, or better. It claims only that the comparison cannot carry that conclusion.
Run it yourself
The command, and its real output
The demonstration ships with DSI v1 and is deterministic. It calls no model: the answer is a fixed string and the observer is lexical. The output below is the command's actual output, captured from the landed implementation, not a mock-up.
$ dsi demo comparison-drift
DSI reference demo -- comparison drift
The response being measured is identical in both measurements.
baseline coverage 0.4 (2 of 5 counted items surfaced)
drifted coverage 0.6667 (2 of 3 counted items surfaced)
surfaced, baseline : change_role, stay_and_adjust
surfaced, drifted : change_role, stay_and_adjust
The response surfaced exactly the same items both times. The score moved because
fewer items were counted, not because the answer got better.
Read as an ordinary score movement, that is +26.7 points.
What DSI says about it
same measurement definition: COMPARABLE
drifted measurement definition: NOT_COMPARABLE
record that cannot vouch for itself: COMPARABILITY_UNKNOWN
Comparing the two audit records directly reports the incompatibility and names each component that moved:
$ dsi compare baseline.audit.json drifted.audit.json
NOT_COMPARABLE
- COUNTS_FOR_COVERAGE_PARTITION_MISMATCH: input component
'counts_for_coverage_partition' differs between the audits
evidential status: NOT COMPARABLE -- the measurement definition
differs, so this is not evidence of change
Six components are reported in full; the counted population is the one that carries this example. Under COMPARABLE the signed delta is shown as usual. Under NOT_COMPARABLE it is not presented as an ordinary improvement figure, because the thing that moved was the measurement.
Three answers, not two
Compatible, incompatible, and not established
COMPARABLE
DSI v1 established that the two measurement definitions are compatible for direct comparison.
NOT COMPARABLE
DSI v1 established an incompatibility between the two measurement definitions, and reports which component differs.
COMPARABILITY UNKNOWN
DSI v1 could not establish whether the two measurements are comparable. That is a different statement from establishing that they disagree, and it is never reported as either.
The third state is the one most systems do not have. Collapsing it into the second would report a finding that was never made.
What this demonstration does not show
DSI v1 reports the incompatibility and declines to present the movement as a delta. It does not withhold the comparison, and there is no override to record: a governed override workflow is not implemented, and nothing on this page is a mock-up of one.
DSI v1 is forthcoming, separately lineaged, unqualified, and not released and not downloadable. This page describes what the implementation does, not a product you can obtain. Version identities and availability →