CORRIGIBILITY
A temporal framework for systems that must remain capable of correcting themselves
A system may retain the formal ability to change a decision and still lose the practical ability to correct its consequences.
THE PROBLEM
When an error is detected, is there still enough time to correct it?
Contemporary systems are often evaluated in terms of accuracy, robustness, transparency, accountability, human oversight, or reversibility.
These properties matter. But they do not answer a different question:
Can an error still be corrected when it is discovered?
A decision may be reversible in principle while becoming irreversible in practice. An appeal may exist but arrive after the relevant consequence has already occurred. A human may retain formal authority while receiving the information too late to intervene meaningfully.
The missing variable is often time.
A MINIMAL MODEL OF CORRIGIBILITY
Corrigibility depends on two timescales
τᵣ — Revision latency
The time required to detect a relevant error, reconsider a model or policy, decide on a correction, and implement it.
τᵥ — Viability horizon
The remaining time during which intervention can still preserve, restore, or redirect an acceptable trajectory.
C = log(τᵥ / τᵣ)
τᵣ < τᵥ
Correction remains reachable.
τᵣ ≈ τᵥ
The system enters a critical corrigibility window.
τᵣ > τᵥ
Revision may still occur, but correction arrives too late.
The distinction is not simply between reversible and irreversible systems, but between systems in which corrective action remains reachable and systems in which it no longer does.
THE CORRIGIBILITY WINDOW
Correction is possible only while intervention remains reachable
Between an initial decision and an effectively irreversible consequence lies a limited interval in which revision can still alter the outcome.
Decision → Error becomes observable → Revision → Corrective action
The decisive constraint is the irreversibility threshold: the point after which revising the decision no longer allows the relevant consequences to be corrected.
Before the threshold: correction remains reachable.
After the threshold: revision may remain possible, but effective correction has been lost.
This limited interval is the corrigibility window.
CURRENT RESEARCH PROJECT
The Corrigibility Window: Can AI-Assisted Decisions Still Be Corrected in Time?
The current applied project asks whether the temporal framework of corrigibility can be translated into a practical method for evaluating AI-assisted institutional decisions.
AI governance often focuses on whether humans retain formal authority, oversight, appeal mechanisms, or the ability to override automated decisions.
The Corrigibility Window asks a stricter question:
At what point does intervention cease to be practically reachable, even if human authority formally remains?
The six-month project will develop and test a framework for identifying these limits in real decision processes.
Planned outputs include:
an open evaluation matrix for temporal corrigibility;
a taxonomy of corrigibility failures;
3–5 documented case studies;
a distinction between formal reversibility and corrective reachability;
external methodological review;
a non-certifying assessment protocol;
an open research report and reusable materials;
a roadmap for subsequent empirical testing.
The objective is not to create another general principle of “human oversight”. It is to determine whether corrective reachability can be identified, compared, and eventually measured.
TESTABILITY
What would count against the framework?
Corrigibility is intended as a testable research program, not as an unfalsifiable vocabulary.
The framework would require substantial revision if evidence showed that:
revision latency had little or no relevance to successful correction;
no meaningful distinction could be identified between reversible and effectively irreversible states;
the proposed temporal variables added no explanatory or predictive value over existing measures;
different forms of correction collapsed into the same functional process;
empirical cases failed to reveal meaningful differences in corrective reachability.
A framework about corrigibility should itself remain corrigible.
RESEARCH STATUS
From theoretical framework to applied research
The Corrigibility program is currently being developed through a sequence of theoretical and applied research projects.
Corrigibility and Continuity: A Temporal Hypothesis on Life and Mind
A theoretical paper developing the temporal foundations of the framework and the relationship between persistence, revision, and viable correction.
Submitted for peer review.
When Information Cannot Correct: Corrective Reachability in Distributed Organizations
A second research line extending the framework toward distributed organizations and the conditions under which information can — or cannot — produce effective correction.
Submitted for peer review.
The current applied project, The Corrigibility Window, translates these ideas into an operational framework for AI-assisted institutional decision systems.
The research program moves from conceptual formulation → operationalization → documented cases → empirical testing.
CURRENT FUNDING
Support the first applied phase of The Corrigibility Window
The six-month applied research project is currently seeking funding through Manifund.
USD 15,000 — Minimum viable project
Core research, case analysis, external methodological review, development of the evaluation framework, and open publication of the results.
USD 25,000 — Expanded project
Additional cases, broader expert review, and stronger public research outputs.
USD 30,000 — Full six-month scope
Includes structured partner identification and feasibility work for a subsequent empirical pilot.
All core research outputs are intended to be openly accessible and reusable.
The objective of this phase is to determine whether corrective reachability can be operationalized well enough to justify empirical testing.
