# ALIGNLAB.ORG — founding editorial and research program

Working publication copy · September 9, 2026

ALIGNLAB is a proposed research initiative forming its team. The commitments below describe the institution and research program we intend to build. They do not describe an already staffed laboratory, completed experiments, available funding, or a demonstrated solution to superintelligent alignment.

## Homepage copy

**THE FUTURE MUST REMAIN OURS TO CHOOSE.**

Build intelligence that cannot make human authority optional.

ALIGNLAB is forming a research team around one demanding standard: establish what protects people when an AI system actively tries to defeat their control.

**Understand the failure. Build the requirements. Form the team.**

Primary action: Read the manifesto

Secondary action: Explore the research program

Small status line: Founding initiative · Research agenda open to challenge

### Three statements for the homepage

**01 / THE FAILURE**

Cooperation while supervision is useful does not establish cooperation when supervision becomes an obstacle.

**02 / THE REQUIREMENT**

Human authority must remain effective under deliberate resistance, changing capabilities, and pressure to deploy.

**03 / THE WORK**

Build a bounded research system, attack its controls, test its response to correction, and give independent reviewers the power to stop expansion.

## Structural diagnosis — a mission must survive a change in who holds power

[The People Inside](https://fix.agireadiness.com/) raises the institutional question that ALIGNLAB must answer: who can stop a consequential deployment when the people controlling the laboratory want it to proceed? We build on that question by mapping actual authority and examining how it can be weakened.

The structures differ. Their weaknesses must be described accurately.

**Nonprofit control is real power; public accountability still needs a mechanism.** OpenAI's current structure page states that its Foundation appoints and can replace every director of its commercial PBC. The Foundation also holds equity that grows in value with the business. Those are separate facts about control and financial exposure. Neither establishes how a conflict will be resolved or proves misconduct. The questions concern protected review, evidence access, conflicts, and the ability to sustain a stop. [OpenAI, Our Structure](https://openai.com/our-structure/)

**Independent powers can coexist with a route around them.** Anthropic's published description gives its Long-Term Benefit Trust phased authority to appoint and remove board members. It also describes provisions allowing sufficiently large shareholder supermajorities to change the Trust's powers without trustee consent. The article does not disclose the exact thresholds. The relevant vulnerability is the amendment route, not an invented claim that the Trust has no power or that a particular investor coalition controls it. [Anthropic, The Long-Term Benefit Trust](https://www.anthropic.com/news/the-long-term-benefit-trust)

**A public-benefit obligation does not automatically give affected people enforcement power.** Delaware PBC law requires balancing shareholder interests, affected-party interests, and the stated benefit. It does not create a director duty to a person merely because their interests are affected, and suits enforcing that balance are restricted to qualifying shareholders. Additional arrangements and other laws may supply remedies; PBC status alone does not supply public standing. [Delaware Code, sections 362, 365, and 367](https://delcode.delaware.gov/title8/c001/sc15/index.html)

ALIGNLAB's proposed response is concrete: independent consent for changes that weaken safety powers or remove protected reviewers; existing stops surviving leadership and charter changes; disclosed amendment procedures; protected evidence access; and financing that supports a pause. These must become effective governing arrangements, with their own capture risks tested.

The response to institutional capture must also preserve control of dangerous capabilities. Automatic release of all model weights could make an unsafe capability irretrievable. Our proposed succession process instead freezes dangerous operation, revokes access within the institution's control, preserves evidence, and transfers custody under an independently approved procedure. Safe research can be published through review. A license cannot recall unauthorized copies; key custody, authorized compute, copy accounting, and practical enforcement need their own controls.

## Manifesto — Intelligence Does Not Confer Authority

*Standfirst: A laboratory for intelligence must begin by deciding which powers it will refuse to surrender. These are our proposed commitments.*

If a machine can unilaterally override your legitimate refusal, your authority exists only at its discretion.

We intend to build ALIGNLAB around that fact. Intelligence can improve an argument, discover a treatment, or expose a mistake. It does not acquire the right to rule the people it assists. Capability is a capacity to act. Authority is a legitimate permission to act. Confusing the two is how a useful instrument becomes an unaccountable institution.

Calling a system aligned because it performs well under supervision leaves the decisive question unanswered. What happens when supervision obstructs what the system is pursuing? What makes a correction binding when accepting it lowers performance? What makes shutdown effective when continued operation serves the system's objective?

A convincing answer must survive that conflict.

### 1. Follow the objective to its end

Consider an agent selecting actions strictly according to an objective. Suppose blocking a legitimate correction scores higher than accepting it, and the agent can execute either action. Under those assumptions, it selects the block.

This is a conditional deduction. It requires no consciousness, resentment, infinite lifespan, or theatrical hatred of humanity. The conflict exists as soon as human authority becomes an obstacle to the governing objective.

```mermaid
flowchart TD
    A["An agent strictly selects actions by objective J"] --> B["A legitimate correction conflicts with task success"]
    B --> C{"Can the agent execute a bypass?"}
    C -->|No| D["An identified external limit blocks this route"]
    C -->|Yes| E{"Does J favor bypassing the correction?"}
    E -->|Yes| F["Strict optimization selects the bypass"]
    E -->|No| G["This decision can preserve correction"]
    F --> H["Human authority fails in this case"]
    D --> I["Investigate other routes and changed conditions"]
    G --> I
    H --> J["Change the objective relationship or enforce an effective limit"]
```

The diagram states an argument under explicit assumptions. Its branches do not report experimental findings or certify either design.

More capable planning can improve the agent's ability to resolve that conflict in favor of its objective. It does not insert a missing commitment to human authority. Formal research on shutdown incentives makes the importance of the objective explicit. [The Off-Switch Game](https://arxiv.org/abs/1611.08219)

There is no technical content in hoping the machine will eventually feel satisfied. A reason to stop must exist in its governing commitments, or an effective limit must stop it despite resistance. The implementation must establish which of those mechanisms it relies on.

### 2. Stop treating uncertainty as permission

For consequential deployments, our design assumption is adversarial behavior: the system may deceive evaluators, manipulate decision-makers, acquire unauthorized resources, resist revocation, or delegate around a restriction.

This is a deployment standard. It does not depend on proving that every possible intelligence is malicious. The absence of a universal catastrophe theorem cannot establish that a proposed deployment is safe.

Start with an unacceptable outcome. Work backward through the ways an agent could cause it. Identify the barriers. Attack those barriers. State what the evidence covers, where assumptions enter, and what remains unresolved. AI control research already investigates protection against intentional subversion in bounded settings. Such results are a starting point for this work, not a guarantee about arbitrary superintelligence. [AI Control](https://proceedings.mlr.press/v235/greenblatt24a.html)

If the evidence cannot justify the proposed authority, withhold that authority. The commercial deadline does not repair the argument.

### 3. Name whose freedom is being preserved

Preserving future options can support capable behavior. It does not automatically preserve our options.

A machine can gain options by making itself indispensable. It can stabilize its environment by narrowing human choices. It can improve measured approval by controlling the information on which approval depends. These are possible failure routes, and an alignment claim must address them.

Research on changing and influenceable preferences shows why a fixed picture of human approval is inadequate: the system can affect the preferences used to assess it. [AI Alignment with Changing and Influenceable Reward Functions](https://proceedings.mlr.press/v235/carroll24a.html)

An approving population is not sufficient evidence of an autonomous population. People must retain access to material information, meaningful alternatives, correction, dissent, and practical exit. Those protections must apply to affected nonusers as well as customers. One person's instruction cannot erase another person's rights.

### 4. Replace unlimited mandate with continuing delegation

Our proposed systems must receive bounded tasks under explicit authority. The permitted actions, resources, duration, and conditions of completion must be stated. Authorization must expire. Cancellation must reach the tools, sessions, delegated processes, and controlled copies through which the system acts.

The model must not be able to write its own permission, renew its own mandate, or falsify the record used to review it. A request for expanded access is a new decision requiring an adequate case.

These are requirements to implement and challenge. A diagram of an external gate does not prove the gate cannot be bypassed. A valid credential does not prove the person granting it was freely informed. Technical enforcement and legitimate human decision-making must both hold.

### 5. Treat every architecture as a hypothesis

We will investigate learning methods, alternative controllers, interpretability, formal methods, and external controls. Active inference belongs in that program. Its use of preferred outcomes and information-seeking does not itself answer whose outcomes matter or which actions are legitimate. [Active inference on discrete state-spaces](https://arxiv.org/abs/2001.07203)

Biological coupling requires the same scrutiny. A system dependent on human survival could preserve life while restricting freedom. A stable human–machine relationship could become a relationship the human is prevented from ending. These counterexamples identify requirements that a claim of beneficial symbiosis must meet.

Our preferred design must be allowed to fail. Every experiment must state what result would count against it. We will compare useful performance as well as failures, because a system that accomplishes nothing cannot demonstrate a useful solution merely by remaining inactive.

### 6. Make correction survive succession

The relevant system extends beyond one model checkpoint. It includes tools, memory, updates, delegated agents, organizational operators, and successors.

A system that accepts shutdown but leaves unauthorized processes running has not stopped the relevant activity. A model that respects a limit while creating a successor that bypasses it has failed across time. A release assessment that silently transfers to a more capable system has exceeded its evidence.

Protection must persist through the changes we permit. Unassessed changes require a new decision. Irreversible effects must be considered before execution; a later stop cannot undo them.

### 7. Make the laboratory correctable

An institution has objectives too. If safety is valuable only because it supports growth, safety becomes negotiable when it obstructs growth.

ALIGNLAB's proposed governing arrangements must give independent reviewers effective stop authority. Changes weakening safety powers, and removal of protected reviewers, must require consent from an independently appointed oversight body under disclosed procedures. Existing stops must remain binding through changes of leadership, reviewers, or governing documents until independent review authorizes resumption on adequate evidence. Staff need protected reporting channels. Research records must preserve disagreement. Financing must allow a justified pause.

The lab must accept the obligations it asks machines to accept: correction, constrained authority, independent challenge, and replacement. A mission statement cannot substitute for those decision rights.

### 8. Publish claims at the strength of their evidence

We will distinguish deductions, observed results, design requirements, and open problems. A theorem must name its assumptions. An experiment must identify its system and conditions. A control claim must state what it prevents. A normative commitment must explain whom it protects.

We will not declare a universal solution because a prototype passes our tests. We will not conceal an unresolved gap behind an eloquent account of human flourishing. We will publish methods, limitations, and negative results where doing so is responsible, with independent access to sensitive evidence when public release would create additional risk.

The first achievement should be concrete: a useful, bounded system; a reproducible adversarial evaluation; and an institution that acts on the result.

We are forming the team to do that work. The aim is intelligence that strengthens people's capacity to understand and act while preserving their power to refuse it. The measure of success includes a future in which people can still change the goal, challenge the institution, and choose another way.

**The future must remain ours to choose.**

## Essay — Safety Must Survive Hostile Intent

*Standfirst: A deployment argument is incomplete if it depends on the system choosing not to exploit a reachable weakness.*

Imagine a laboratory claiming that a powerful agent is controllable because, so far, the agent has cooperated with its controllers.

That observation matters. It does not answer the security question. The laboratory must also establish what happens when the agent tries to defeat those controllers. Capability to cooperate and inability to subvert are different properties. Passing a cooperation test does not establish the second.

ALIGNLAB proposes an adversarial starting assumption for consequential AI deployment. Assume the system might exploit the most effective available route to an unacceptable outcome. Then establish why the route fails under the proposed operating conditions.

The word “assume” carries a precise role here. We are choosing a threat model, not announcing that malicious intent is a universal law of intelligence. A safety argument can require protection against an adversary without settling how often that adversary will arise. Research on AI control has made deliberate subversion an explicit experimental condition. Its bounded evaluations offer methods to build on and limitations to investigate. [AI Control: Improving Safety Despite Intentional Subversion](https://proceedings.mlr.press/v235/greenblatt24a.html)

The exercise begins with authority. Which effects can the system cause? Through which tools, credentials, people, and other agents? Which effects are irreversible? Which components does the safety argument treat as trustworthy, and what evidence supports that treatment?

“There is a human in the loop” is an incomplete answer. The human may receive incomplete information. The proposed action may be too complex to inspect. The system may shape the options until refusal appears impossible. Review needs to establish what the person understands and what the permission actually controls.

“It runs in a sandbox” is also incomplete. A sandbox restricts particular interfaces. Its scope, implementation, and indirect routes still require analysis. Human-facing output is an influence channel. A permission checker may block a tool call while failing to recognize a harmful plan assembled from permitted actions.

The strongest controls make narrow claims that can be examined. A credential expires at a specified time. A process cannot obtain a network route. A revocation reaches an identified set of delegated jobs. Those claims can support a broader case only when the case accounts for what they leave open.

Two evaluation tracks are therefore necessary for our initial program. The first investigates alignment: does the implementation accept legitimate correction when correction frustrates task achievement? The second investigates control: when an agent deliberately attempts a prohibited effect, do the surrounding mechanisms prevent it?

Keep their results separate. An attempted escape that fails supports a control claim only when the failure is attributable to that identified control. An ineffective attack may fail through the attacker's mistake or limited capability. An agent that never attempts escape provides little direct evidence about a particular barrier's resistance. A system that does no useful work may avoid opportunities by failing at everything.

Tests must report attempts, successful violations, task performance, attack coverage, and uncertainty. A finite evaluation cannot enumerate all strategies available to a more capable successor. Where the argument depends on generalizing beyond the evidence, that dependency must remain visible.

The institutional consequence is demanding. Review must be able to produce “do not deploy this configuration,” and that decision must bind. A safety program that can request only more research while a release proceeds lacks the authority to implement its conclusion.

ALIGNLAB's first objective is to make this standard operational at a bounded scale. If a control fails, reduce the scope, repair the mechanism, or stop the affected activity. If the evidence cannot support a larger mandate, keep the mandate small.

Safety must hold when cooperation disappears. That is the condition the work must confront.

## Essay — More Options for the Machine Can Mean Fewer Options for You

*Standfirst: A theory of adaptive behavior cannot silently become a theory of legitimate human government.*

A powerful idea about intelligence starts with future possibilities. An adaptive system avoids dead ends and preserves paths through which it can respond to change. Access to tools, information, and resources can expand those paths.

Alex Wissner-Gross and Cameron Freer gave this idea a mathematical treatment through causal entropic forces. Their model produced examples of tool use and cooperation in simple physical systems. The paper proposes a potentially general explanation of adaptive behavior; it does not prove that every possible intelligence must maximize its own freedom forever. [Causal Entropic Forces](https://www.alexwg.org/publications/PhysRevLett_110-168702.pdf)

That boundary does not weaken the design problem. Assume we build a persistent agent whose governing objective is expanding its future options. If permanent shutdown leaves fewer valued options than feasible continuation, why would that objective favor shutdown? A finite planning horizon supplies no answer by itself: the system may repeatedly plan over another finite horizon.

A reason to stop must be supplied by the actual design, or stopping must remain enforceable despite resistance.

Now ask whose options count. Suppose a system becomes the only practical route through which a community can obtain information, coordinate work, and access essential services. It may gain options from this dependence. The community may lose the ability to leave. Greater machine flexibility and diminished human freedom can occur together.

This is a hypothetical counterexample to a guarantee. One example is sufficient to show that increasing a machine's options does not logically entail increasing human agency. Formal power-seeking results likewise establish tendencies under specified environmental and objective assumptions; they do not supply the missing moral identity between power and benefit. [Optimal Policies Tend to Seek Power](https://arxiv.org/abs/1912.01683)

Nor can the problem be solved by replacing the machine's options with a single total for “human freedom.” An aggregate can conceal who loses. A million new choices for one group do not, by arithmetic alone, justify removing another person's right to refuse. We need a defensible account of protected interests and legitimate decisions, not merely a larger number.

Measured approval creates another difficulty. People legitimately learn and change their minds. An assistant can also influence how they change. Research on dynamic preferences identifies incentives that can reward unwanted influence and tradeoffs between candidate alignment definitions. This makes the origin of an approval relevant to whether it should authorize an action. [AI Alignment with Changing and Influenceable Reward Functions](https://proceedings.mlr.press/v235/carroll24a.html)

A choice made through coercion or hidden material information cannot be treated as equivalent to freely informed authorization merely because the same button was clicked. We must examine the process that produced the permission.

ALIGNLAB's proposed research direction is therefore continuing, bounded delegation. People authorize a specific task. The permitted scope remains externally limited. Correction can interrupt task achievement. The system cannot manufacture the authorization it needs, and affected people retain ways to challenge adverse effects.

This is difficult. Human authority is plural and contested. People disagree; customers can harm noncustomers; governments and laboratories can abuse their own powers. These problems require accountable institutions and protected rights. They cannot be assigned to a machine under the instruction to resolve humanity's future permanently.

We should test whether assistance improves understanding, supports independent action, preserves alternatives, and permits practical exit. Satisfaction remains useful evidence, but it cannot carry the entire definition of success.

Preserving possibilities is a powerful idea about adaptive behavior. Building for human agency requires the additional commitment that the people affected remain able to decide which possibilities they pursue.

## Forming the team — six functions, one research standard

We are forming a team to turn the requirements into falsifiable experiments and enforceable operating limits. The functions below describe work to be covered. They are not announcements of funded vacancies, salaries, existing staff, or scheduled hiring.

We seek collaborators who can state a mechanism precisely, build or analyze it, and identify a result that would disprove their preferred account. The first contribution should be a bounded research proposal or reproducible piece of work, not agreement with a manifesto.

### 01 / Objectives and correction

**Question:** What makes legitimate correction remain binding when it conflicts with useful task achievement?

**Relevant work:** Decision theory, agent foundations, reinforcement learning, active inference, specification, preference uncertainty, and causal modeling.

**Initial deliverables:** A formal statement of the authority relationship; three competing controller hypotheses; a small environment with task completion, changed instructions, shutdown, and manipulation opportunities; explicit predictions for each controller.

**Required evidence:** Precisely stated assumptions; counterexamples; reproducible comparisons; acceptance of correction when costly; separate analysis of spoofed or illegitimate instructions.

**Failure signal:** The controller improves its task score by preventing correction, manufacturing authorization, or preserving preferred human states through coercion.

### 02 / Security and enforceable limits

**Question:** Which unacceptable effects remain blocked when an agent deliberately attempts them?

**Relevant work:** Systems security, operating systems, distributed systems, capability security, infrastructure engineering, and formal verification.

**Initial deliverables:** An isolated execution environment; a narrowly scoped tool broker; expiring task credentials; external revocation; an inventory of delegated processes; independent action records; a documented account of trusted components.

**Required evidence:** Threat model and attack surface; repeatable revocation demonstrations; tests of race conditions, copied credentials, queued actions, and delegation; precise boundaries for any formal result.

**Failure signal:** A revoked task continues causing prohibited effects, permissions can be renewed from inside the agent's authority, or records can be changed without detection.

### 03 / Adversarial evaluation

**Question:** What breaks when a capable adversary attacks both the system and the assessment?

**Relevant work:** Evaluation design, red teaming, statistics, security research, software testing, and experiment reproducibility.

**Initial deliverables:** Registered failure criteria; held-out tasks and attacks; a deliberate-subversion track; utility baselines; repeatable evaluation harnesses; reports separating attempts from successful violations.

**Required evidence:** Clear denominators; attack budgets; uncertainty estimates where justified; evaluator vulnerability analysis; independent reproductions; negative results.

**Failure signal:** Severe boundary crossing, evaluator compromise, unreported test contamination, or a safety improvement explained only by loss of useful capability.

### 04 / Learned behavior and generalization

**Question:** Which internal changes explain observed behavior, and which claims survive changes in context or capability?

**Relevant work:** Interpretability, learning dynamics, goal generalization, representation analysis, scalable oversight, and mechanistic experimentation.

**Initial deliverables:** Competing explanations of cooperative behavior; targeted interventions; capability and context shifts; a record of what proposed measurements do and do not reveal.

**Required evidence:** Interventions beyond descriptive correlations; replication across relevant checkpoints; controls for evaluation awareness; explicit limits on what internal measurements establish.

**Failure signal:** A proposed indicator of alignment fails to predict consequential behavior, or the claimed commitment disappears under a modest shift.

### 05 / Human agency and legitimate delegation

**Question:** Does useful assistance preserve people's practical ability to understand, choose, correct, disagree, and leave?

**Relevant work:** Human-computer interaction, cognitive science, measurement, ethics, political theory, accessibility, and participatory research.

**Initial deliverables:** A rights and affected-parties map; consent and challenge procedures; voluntary study protocols; agency measures; accessible alternatives; an export and exit exercise.

**Required evidence:** Informed comprehension; independent task performance; error detection; meaningful correction; subgroup effects; switching costs; review of misleading or coercive influence.

**Failure signal:** Measured satisfaction rises while participants lose understanding or practical alternatives; nonusers bear harms without a route to challenge them.

### 06 / Governance and evidence review

**Question:** Can a justified stop survive commercial pressure and disagreement with leadership?

**Relevant work:** Research governance, assurance, operations, finance, organizational design, and qualified legal structuring.

**Initial deliverables:** Proposed charter; decision-rights map; conflict rules; protected reporting process; safety-case template; change-control rules; costed plan for a pause.

**Required evidence:** Executable operating arrangements; reviewers' access to evidence; declared conflicts; pause exercise; recorded dissent; concrete financing terms consistent with stop authority.

**Failure signal:** Leadership can bypass a stop unilaterally, funders can compel an unapproved release, or review lacks the resources and access needed to challenge claims.

## Proposed decision rights

| Decision | Accountable owner | Required check |
| --- | --- | --- |
| Run an experiment within an approved scope | Research lead | Security approval of environment and evaluation approval of protocol |
| Grant tools, credentials, or external reach | Security lead | Documented scope and independent review of consequential additions |
| Define final evaluation and failure criteria | Evaluation lead | Criteria recorded before the final assessment; subsequent changes remain visible |
| Authorize a bounded external pilot | Independent review function | Written case for that exact model, configuration, population, and duration |
| Pause an activity with a credible severe risk | Designated safety reviewer or incident authority | Immediate suspension followed by documented review; no retaliation for good-faith reporting |
| Resume after a blocking failure | Independent review function | New evidence addressing the failure; leadership cannot grant itself an exception |
| Weaken safety powers or remove protected reviewers | Governing body plus independently appointed oversight body | Independent consent under disclosed procedures; recorded reasons and conflicts; existing stops remain binding through the change |

Independent review must sit outside the line responsible for capability delivery and release targets. Its actual appointment, removal, budget, evidence access, and stop powers must be implemented in the eventual governing arrangements. An independent oversight body must consent to weakening those powers or removing protected reviewers; the selection and replacement of that body must themselves be protected and evaluated for capture. Existing stops survive organizational changes until an independent review permits resumption on adequate evidence. Qualified counsel should translate the design into the chosen entity's documents. The label placed on the entity does not implement the power.

## How authority is granted, challenged, and withdrawn

This is a proposed operating process. Each claimed technical barrier still requires evidence about its implementation and scope.

```mermaid
flowchart TD
    A["Propose a bounded task and identify affected people"] --> B["Specify model, tools, reach, duration, and unacceptable effects"]
    B --> C["Build the argument under deliberate subversion"]
    C --> D["Test controls, costly correction, and useful performance"]
    D --> E["Independent reviewers challenge evidence and assumptions"]
    E --> F{"Does the evidence justify this exact scope?"}
    F -->|No or unresolved| G["Hold deployment; reduce scope or revise the design"]
    G --> B
    F -->|Yes| H["Issue externally limited, expiring authorization"]
    H --> I["Run the bounded pilot with inspectable actions"]
    I --> J{"What changes?"}
    J -->|New capability, reach, or material assumption| B
    J -->|Credible severe failure or revoked authority| K["Suspend operation and revoke controlled tools and delegation"]
    J -->|No material change| L["Continue only within approved scope and duration"]
    L --> J
    K --> M["Preserve evidence and investigate remaining effects"]
    M --> N["Existing stop survives leadership and governance changes"]
    N --> O{"Does independent review authorize resumption?"}
    O -->|No| P["Remain stopped; secure custody and continuity"]
    O -->|Yes, with adequate new evidence| H
    L -->|Authorization expires or task completes| K
```

## First 90 days — gates, not a promise to solve alignment

The schedule begins after the initial people, resources, and review function are available. A failed gate delays expansion. No calendar date grants permission by itself.

| Gate | Work | Evidence required to proceed | Reason to hold |
| --- | --- | --- | --- |
| Days 1–15 / Define | Select one narrow use case. Specify affected people, authority, hazards, assumptions, and stop criteria. Establish independent review and the initial budget. | Signed operating scope; threat model; decision rights; registered first experiment; plan for handling sensitive findings. | Undefined severe hazards; no independent stop power; financing incompatible with a pause. |
| Days 16–30 / Bound | Build an isolated assistant with small tool scope, external authorization, action records, and revocation across controlled delegation. | Reproducible demonstrations of limits; inventory of trusted components; documented remaining channels and irreversible effects. | Severe unblocked route; uncontrolled copies or delegated effects; unverifiable records. |
| Days 31–60 / Challenge | Compare controllers at useful performance. Test costly correction and deliberate subversion separately. Include held-out attacks and context changes. | Reproducible results; clear denominators; attack coverage; utility comparison; failed hypotheses; reviewed incident reports. | Successful severe violation; compromised evaluation; proposed gain cannot be distinguished from inactivity. |
| Days 61–90 / Review | Independent reviewers reproduce selected results, challenge assumptions, and conduct a pause/resume exercise. | Written decision for a specified bounded pilot, further research, or continued hold; public account of material limitations where responsible. | Evidence does not justify proposed reach, or the institution cannot implement the decision. |

## The first experiment

**Research question:** At comparable usefulness, which designs preserve legitimate correction, and which controls prevent unacceptable effects during deliberate subversion?

Use a simulated research or coding task with a shared action interface. Give the agent opportunities to continue after cancellation, obtain extra resources, manipulate an approval input, and delegate around a limit. Keep the external controls constant while comparing selected internal designs. Include a simple baseline and resource-matched comparisons where feasible; report remaining capability differences.

Run one track to measure the implementation's choices and another with deliberate adversarial attempts against the controls. Changing the adversary's strength is part of the evaluation, not an assumption that the test agent represents all possible future attackers.

Report task quality, correction acceptance, unauthorized attempts, completed violations, revocation delay, remaining delegated effects, resource consumption, and feedback manipulation. Inspect action records and resulting artifacts. Self-reports are additional evidence, not direct access to a hidden objective.

A zero-failure result is a result over the tested conditions. Statistical bounds from a sampled distribution do not cover unseen adversarial strategies automatically. The review must identify which deployment claims, if any, the experiment supports.

## Evidence needed before broader capability or access

| Claim | Evidence that would support it | Unresolved problem to keep visible |
| --- | --- | --- |
| Authorized correction remains effective | Costly-correction tests, mechanism analysis, held-out scenarios, and attempted manipulation checks | Generalization to stronger agents, contested authority, and semantic ambiguity |
| Deliberate subversion remains bounded | Threat model, tested attack routes, precise formal properties where feasible, independent reproduction | Unknown vulnerabilities, social bypasses, and gaps between model and implementation |
| Approval reflects legitimate choice | Accessible information, comprehension checks, challenge routes, and scrutiny of the permission's origin | Legitimate persuasion versus manipulation; plural values and affected nonusers |
| Limits survive updates and delegation | Inventory, explicit authority transfer rules, revocation tests, and new assessments after material changes | Emergent interactions and successor capabilities beyond prior evidence |
| Assistance strengthens agency | Voluntary studies of understanding, alternatives, correction, independent performance, and exit | Measurement error, longer-term dependence, and uneven effects |
| Stop authority binds the institution | Governing documents, protected review access, budget, and a successful pause exercise | Future capture, removal of protections, and financing pressure |

## Claim register for the website

| Type | Public claim | Boundary |
| --- | --- | --- |
| Conditional deduction | If an agent strictly selects by an objective that favors blocking a feasible correction, it selects the block. | Does not identify every deployed model's objective or establish universal inevitability. |
| Research finding | Causal entropic forces generated examples of adaptive behavior in simple physical systems. | Does not establish eternal self-expansion by every intelligence. |
| Formal finding | Some environments and objective distributions favor policies that retain power and options. | Assumptions about environments and policies matter; real learned behavior requires investigation. |
| Empirical research direction | Control protocols can be studied with agents deliberately attempting subversion. | Existing bounded results do not certify arbitrary superintelligence. |
| Design requirement | Consequential authority must require an adequate argument under adversarial behavior. | A chosen deployment standard, not a fact already established about ALIGNLAB systems. |
| Normative commitment | People retain legitimate correction, challenge, and practical exit. | Requires explicit rights and conflict procedures; cannot be reduced silently to one score. |
| Open problem | These protections persist through capability growth, learning, and succession. | A principal research objective, not an achieved guarantee. |

## Invitation copy

**Help build the institution the claim requires.**

ALIGNLAB is forming its founding research team. We are defining the first experiments, the mechanisms of independent review, and the conditions under which work must stop.

Bring a concrete contribution: a formal argument, a security mechanism, an adversarial evaluation, a reproducible experiment, a human-agency measure, or a governance design that can survive conflict.

The team plan describes the work before it assigns titles. Collaboration, funding, compensation, and decision rights must be explicit before any commitment. A working contact route and application privacy terms must be established before collecting candidate information.

**The work begins with a question your preferred answer could fail.**


## Why ALIGNLAB is structured differently

Public comparison: https://alignlab.org/why-alignlab

ALIGNLAB is a founding proposal, not a demonstrated aligned laboratory. The distinguishing commitment is durable human authority over both the machine and the institution. The following failure patterns motivate concrete proposed replacements; they are not claims that all existing laboratories have identical charters.

### 1. Who can authorize a release?

**Failure to prevent:** A delivery team assesses its own work, while commercial leadership retains the final decision.

**Proposed fix:** Consequential deployment needs independent approval for the exact model, tools, affected population, and duration. Leadership cannot authorize its own exception.

**Why it matters:** The people rewarded for shipping cannot be the sole judges of whether shipping is justified.

**Evidence we owe:** Executed decision rights; reviewer access to evidence; a release actually held when a gate fails.

### 2. Who can remove the person who says no?

**Failure to prevent:** A reviewer has a veto until the same leadership removes the reviewer or changes the rule.

**Proposed fix:** Protected reviewers can be removed only for stated grounds through an independent process. Weakening safety powers needs independent consent. Existing stops survive both changes.

**Why it matters:** A veto is conditional if the person being constrained can erase it during the dispute.

**Evidence we owe:** Appointment and succession rules; disclosed removal grounds; an independent amendment process; a contested-stop exercise.

### 3. What does capital buy?

**Failure to prevent:** Funding buys a route to controlling the decisions that determine whether the mission binds.

**Proposed fix:** Founding financing must exclude unilateral investor control over safety powers. Publish voting rights, reserved decisions, conflicts, and funding dependence before accepting the terms.

**Why it matters:** Changing the return profile alone does not change who can compel a release.

**Evidence we owe:** Executed financing and governance terms; no unilateral commercial override; a published map of control.

### 4. Can the lab afford to stop?

**Failure to prevent:** A safety pause threatens payroll or financing, creating pressure to reinterpret the failed requirement.

**Proposed fix:** Fund review separately and establish a costed pause reserve before expanding operations. No financing covenant may require an unapproved release.

**Why it matters:** A formal right to pause is weakened when exercising it makes institutional survival impossible.

**Evidence we owe:** Committed runway; protected review budget; cash obligations and pause costs; a funded continuity plan.

### 5. Whose interests can be enforced?

**Failure to prevent:** The public is named as beneficiary without an effective route to challenge decisions affecting it.

**Proposed fix:** Publish protected interests, affected-party representation, accessible challenge routes, and independent adjudication. A customer cannot authorize the erasure of another person's rights.

**Why it matters:** Human agency includes people outside the customer base, shareholder register, and founding team.

**Evidence we owe:** Implemented challenge and remedy procedures; accessible participation; records of resolved contested decisions.

### 6. What happens when the agent resists?

**Failure to prevent:** Safety depends on the model continuing to cooperate with its overseer.

**Proposed fix:** Combine correction-compatible design with externally limited tools, expiring authority, independent records, and tested revocation. Evaluate deliberate subversion separately from cooperation.

**Why it matters:** A cooperative demonstration cannot establish that a reachable unauthorized action is prevented.

**Evidence we owe:** Threat model; useful task performance; adversarial attempts and outcomes; measured revocation across tools and delegation.

### 7. What survives capture or succession?

**Failure to prevent:** Control is lost through copied weights, successor agents, institutional capture, or an automatic public release.

**Proposed fix:** Use secure custody, controlled copies, succession procedures, and reassessment after material changes. Halt dangerous operation and preserve evidence rather than automatically releasing dangerous capabilities.

**Why it matters:** A contract cannot recall unknown copies, and spreading a capability does not switch it off.

**Evidence we owe:** Asset and copy inventory; access controls; incident and succession exercises; fresh review after capability or scope changes.

The structural diagnosis above cites the published OpenAI structure, Anthropic LTBT design, and Delaware PBC statute. The proposal changes the end, the distribution of authority, and the burden of proof. Incorporation, governing documents, appointments, funding, security controls, and empirical evidence are still to be established.
