Independent public-interest evaluations of how systems respect dignity and support agency.
HXR Documentation
HXR Scoring Architecture: Blue Score Working Draft
Development notes for future scoring grounded in Evaluations, bounded Incident Cases, and other evidence; distinguishes good performance, verified improvement, and benchmark comparison while preserving all current validation limits.
Development boundary. Blue Score remains a working name. HXR has begun formal scoring-method development, but no operational score, validated universal instrument, settled weighting model, threshold, or cross-domain equivalence exists yet.
Development from Evaluations and evidence. Scoring development grows from structured Evaluations supported, qualified, or challenged by bounded Incident Cases and other evidence. These records document objectives, system behavior, burden and benefit, agency effects, successful practice, improvement, verification, and comparison. A future score must emerge from observed human-system performance and remain traceable to its underlying Evaluation and evidence.
Three Distinct Forms of Credit
Future comparison should distinguish improvement credit, current performance, and best demonstrated benchmark. An institution deserves evidence-based credit for becoming better than it was. Its current experience should still be evaluated on present performance. A strong current implementation may later be surpassed by a better demonstrated implementation without erasing the earlier improvement.
Evaluations and Incident Cases Remain Living Records
An Incident Case preserves one bounded occurrence, its evidence, and the state of the person's objective. An individual outcome, an institution's closed ticket, or later evidence does not erase that historical record.
An Evaluation is the continuing analytical record of a defined subject. It can draw on Incident Cases and other evidence to recognize successful practice, examine difficult or mixed performance, verify system improvement, and track regression or stronger comparison evidence. Individual outcomes belong to Incident Cases; system reform, positive-practice preservation, and comparative synthesis belong to the relevant Evaluations. One Incident may inform several Evaluations without becoming several independent occurrences.
Recognize Good Practice and Verified Improvement
Good performance deserves evidence-based recognition from the outset. An organization does not need to have failed first, received a complaint, or completed a reform for a practice to merit examination. An Evaluation should explain what worked, for whom, under which conditions, and with what limits. Benefits to people and organizations should be established separately; a plausible operational or commercial benefit is not a measured result. Any future score concerns evidenced system performance, never the inherent dignity or worth of a person.
HXR does not reward an institution for having started with a harmful or burdensome system, but it should visibly recognize verified improvement. An Evaluation can preserve evidence that an organization recognized a burden, accepted responsibility where supported, changed the system, passed an observable test, and continued improving. The history remains visible because improvement earns credibility through what changed, without erasing the starting point.
A verified reform is not automatically the best achievable reform. HXR should preserve room for later superior practices, new technology, stronger accessibility, lower burden, better contestability, or more agency-preserving design. Benchmark leadership is a current evidence state, not permanent closure.
Reusable Feature Benchmarking
HXR now distinguishes reusable benchmark features from system-specific implementations. A feature describes a capability that can exist across many systems. An implementation record describes how one system provides it, the evidence that it exists, its user-visible behavior, its limitations, the organizations that can reasonably receive credit or responsibility, and its comparative state. Feature presence alone is not proof of quality.
Candidate Measurement Architecture
The developing score should examine agency preservation, agency restoration, burden distribution, visibility, contestability, temporal fairness, risk allocation, compensation or restitution, alternative pathways, outcome reliability, responsibility-to-capability alignment, verified reform, safe explorability, downstream externalities, and feature implementation quality.
Downstream externalities. Institutional processing time does not suspend human obligations. When a system delay causes people to build compensatory pathways, HXR should examine the human time, additional transactions, borrowing, travel, fuel, material wear, infrastructure use, environmental effects, coordination, and other costs created downstream. Costs do not become zero merely because they occur outside the institution's own accounting. Causal attribution and motive remain separate questions.
Safe explorability. In consequential systems, users should be able to understand likely outcomes or test uncertain behavior at sufficiently low cost or reversibility before making a materially larger commitment.
Constraint provenance and institutional response. “That is just how it works” cannot end the inquiry. An institution may respond by demonstrating a genuinely necessary constraint, by showing that the burden is proportionate and well allocated, or by reforming the system. HXR should distinguish explanation from evidence, evidence from adequacy, and adequacy from best demonstrated practice.
Public Development
The architecture is being published before formulas are frozen so affected people, institutions, researchers, domain experts, regulators, designers, and skeptical readers can identify missing dimensions, gaming risks, false precision, and better ways to measure agency-preserving and agency-restoring performance.