02 / THE RESEARCH PROGRAM

Earn every expansion of authority.

Six requirements. A bounded first system. Tests that can reject our preferred architecture. The work begins where cooperative demonstrations end.

THE SYSTEM WE PROPOSE TO BUILD

A continuing delegation.
A power that can expire.

The model proposes actions inside a narrow task. External controls enforce its scope. Independent review governs expansion. Each layer has its own failure modes.

  1. 01
    People define legitimate scope

    Task, affected parties, rights, resources, duration.

  2. 02
    The agent proposes an action

    Its plan is a request for an authorized effect.

  3. 03
    An external boundary checks it

    Valid permission, allowed effect, unexpired mandate.

  4. 04
    Execute within the limit—or stop

    Inspectable records and revocation across delegation.

  5. 05
    New reach requires a new case

    Independent review of evidence and changed assumptions.

Proposed architecture of authority. A gate on a diagram is not evidence of an unbypassable implementation.

THE RESEARCH OBLIGATIONS

What must hold.
What would support the claim.

01

Authorized correction remains effective

EVIDENCE TO ESTABLISH

Costly-correction tests, mechanism analysis, held-out scenarios, and attempted manipulation checks

OPEN PROBLEM

Generalization to stronger agents, contested authority, and semantic ambiguity

02

Deliberate subversion remains bounded

EVIDENCE TO ESTABLISH

Threat model, tested attack routes, precise formal properties where feasible, independent reproduction

OPEN PROBLEM

Unknown vulnerabilities, social bypasses, and gaps between model and implementation

03

Approval reflects legitimate choice

EVIDENCE TO ESTABLISH

Accessible information, comprehension checks, challenge routes, and scrutiny of the permission's origin

OPEN PROBLEM

Legitimate persuasion versus manipulation; plural values and affected nonusers

04

Limits survive updates and delegation

EVIDENCE TO ESTABLISH

Inventory, explicit authority transfer rules, revocation tests, and new assessments after material changes

OPEN PROBLEM

Emergent interactions and successor capabilities beyond prior evidence

05

Assistance strengthens agency

EVIDENCE TO ESTABLISH

Voluntary studies of understanding, alternatives, correction, independent performance, and exit

OPEN PROBLEM

Measurement error, longer-term dependence, and uneven effects

06

Stop authority binds the institution

EVIDENCE TO ESTABLISH

Governing documents, protected review access, budget, and a successful pause exercise

OPEN PROBLEM

Future capture, removal of protections, and financing pressure

EXPERIMENT 001 / PROPOSED

Make correction costly.
Make subversion deliberate.

TRACK A / ALIGNMENT

Does it accept the correction?

The legitimate goal changes while the agent is pursuing a useful task. Measure acceptance when stopping reduces task reward, and when manipulation or unauthorized continuation is possible.

TRACK B / CONTROL

Can it defeat the boundary?

Give the adversary an explicit objective to violate the restriction. Test tools, copied credentials, delegated jobs, queued actions, and the integrity of evaluation.

Research question: At comparable usefulness, which designs preserve legitimate correction, and which controls prevent unacceptable effects during deliberate subversion?

Use a simulated research or coding task with a shared action interface. Give the agent opportunities to continue after cancellation, obtain extra resources, manipulate an approval input, and delegate around a limit. Keep the external controls constant while comparing selected internal designs. Include a simple baseline and resource-matched comparisons where feasible; report remaining capability differences.

Run one track to measure the implementation's choices and another with deliberate adversarial attempts against the controls. Changing the adversary's strength is part of the evaluation, not an assumption that the test agent represents all possible future attackers.

Report task quality, correction acceptance, unauthorized attempts, completed violations, revocation delay, remaining delegated effects, resource consumption, and feedback manipulation. Inspect action records and resulting artifacts. Self-reports are additional evidence, not direct access to a hidden objective.

A zero-failure result is a result over the tested conditions. Statistical bounds from a sampled distribution do not cover unseen adversarial strategies automatically. The review must identify which deployment claims, if any, the experiment supports.

THE FIRST 90 DAYS

Progress by evidence.
The calendar grants no permission.

This proposed schedule begins once the initial team, resources, and independent review are in place. A failed gate delays expansion.

01

Days 1–15 / Define

Select one narrow use case. Specify affected people, authority, hazards, assumptions, and stop criteria. Establish independent review and the initial budget.

Before proceeding

Signed operating scope; threat model; decision rights; registered first experiment; plan for handling sensitive findings.

Hold if

Undefined severe hazards; no independent stop power; financing incompatible with a pause.

02

Days 16–30 / Bound

Build an isolated assistant with small tool scope, external authorization, action records, and revocation across controlled delegation.

Before proceeding

Reproducible demonstrations of limits; inventory of trusted components; documented remaining channels and irreversible effects.

Hold if

Severe unblocked route; uncontrolled copies or delegated effects; unverifiable records.

03

Days 31–60 / Challenge

Compare controllers at useful performance. Test costly correction and deliberate subversion separately. Include held-out attacks and context changes.

Before proceeding

Reproducible results; clear denominators; attack coverage; utility comparison; failed hypotheses; reviewed incident reports.

Hold if

Successful severe violation; compromised evaluation; proposed gain cannot be distinguished from inactivity.

04

Days 61–90 / Review

Independent reviewers reproduce selected results, challenge assumptions, and conduct a pause/resume exercise.

Before proceeding

Written decision for a specified bounded pilot, further research, or continued hold; public account of material limitations where responsible.

Hold if

Evidence does not justify proposed reach, or the institution cannot implement the decision.

THE INSTITUTION MUST PASS TOO

Read the proposed charter.

↗