Authorized correction remains effective
Costly-correction tests, mechanism analysis, held-out scenarios, and attempted manipulation checks
Generalization to stronger agents, contested authority, and semantic ambiguity
Six requirements. A bounded first system. Tests that can reject our preferred architecture. The work begins where cooperative demonstrations end.
The model proposes actions inside a narrow task. External controls enforce its scope. Independent review governs expansion. Each layer has its own failure modes.
Proposed architecture of authority. A gate on a diagram is not evidence of an unbypassable implementation.
Costly-correction tests, mechanism analysis, held-out scenarios, and attempted manipulation checks
Generalization to stronger agents, contested authority, and semantic ambiguity
Threat model, tested attack routes, precise formal properties where feasible, independent reproduction
Unknown vulnerabilities, social bypasses, and gaps between model and implementation
Accessible information, comprehension checks, challenge routes, and scrutiny of the permission's origin
Legitimate persuasion versus manipulation; plural values and affected nonusers
Inventory, explicit authority transfer rules, revocation tests, and new assessments after material changes
Emergent interactions and successor capabilities beyond prior evidence
Voluntary studies of understanding, alternatives, correction, independent performance, and exit
Measurement error, longer-term dependence, and uneven effects
Governing documents, protected review access, budget, and a successful pause exercise
Future capture, removal of protections, and financing pressure
The legitimate goal changes while the agent is pursuing a useful task. Measure acceptance when stopping reduces task reward, and when manipulation or unauthorized continuation is possible.
Give the adversary an explicit objective to violate the restriction. Test tools, copied credentials, delegated jobs, queued actions, and the integrity of evaluation.
Research question: At comparable usefulness, which designs preserve legitimate correction, and which controls prevent unacceptable effects during deliberate subversion?
Use a simulated research or coding task with a shared action interface. Give the agent opportunities to continue after cancellation, obtain extra resources, manipulate an approval input, and delegate around a limit. Keep the external controls constant while comparing selected internal designs. Include a simple baseline and resource-matched comparisons where feasible; report remaining capability differences.
Run one track to measure the implementation's choices and another with deliberate adversarial attempts against the controls. Changing the adversary's strength is part of the evaluation, not an assumption that the test agent represents all possible future attackers.
Report task quality, correction acceptance, unauthorized attempts, completed violations, revocation delay, remaining delegated effects, resource consumption, and feedback manipulation. Inspect action records and resulting artifacts. Self-reports are additional evidence, not direct access to a hidden objective.
A zero-failure result is a result over the tested conditions. Statistical bounds from a sampled distribution do not cover unseen adversarial strategies automatically. The review must identify which deployment claims, if any, the experiment supports.
This proposed schedule begins once the initial team, resources, and independent review are in place. A failed gate delays expansion.
Select one narrow use case. Specify affected people, authority, hazards, assumptions, and stop criteria. Establish independent review and the initial budget.
Signed operating scope; threat model; decision rights; registered first experiment; plan for handling sensitive findings.
Undefined severe hazards; no independent stop power; financing incompatible with a pause.
Build an isolated assistant with small tool scope, external authorization, action records, and revocation across controlled delegation.
Reproducible demonstrations of limits; inventory of trusted components; documented remaining channels and irreversible effects.
Severe unblocked route; uncontrolled copies or delegated effects; unverifiable records.
Compare controllers at useful performance. Test costly correction and deliberate subversion separately. Include held-out attacks and context changes.
Reproducible results; clear denominators; attack coverage; utility comparison; failed hypotheses; reviewed incident reports.
Successful severe violation; compromised evaluation; proposed gain cannot be distinguished from inactivity.
Independent reviewers reproduce selected results, challenge assumptions, and conduct a pause/resume exercise.
Written decision for a specified bounded pilot, further research, or continued hold; public account of material limitations where responsible.
Evidence does not justify proposed reach, or the institution cannot implement the decision.