If a machine can unilaterally override your legitimate refusal, your authority exists only at its discretion.
We intend to build ALIGNLAB around that fact. Intelligence can improve an argument, discover a treatment, or expose a mistake. It does not acquire the right to rule the people it assists. Capability is a capacity to act. Authority is a legitimate permission to act. Confusing the two is how a useful instrument becomes an unaccountable institution.
Calling a system aligned because it performs well under supervision leaves the decisive question unanswered. What happens when supervision obstructs what the system is pursuing? What makes a correction binding when accepting it lowers performance? What makes shutdown effective when continued operation serves the system's objective?
A convincing answer must survive that conflict.
1. Follow the objective to its end
Consider an agent selecting actions strictly according to an objective. Suppose blocking a legitimate correction scores higher than accepting it, and the agent can execute either action. Under those assumptions, it selects the block.
This is a conditional deduction. It requires no consciousness, resentment, infinite lifespan, or theatrical hatred of humanity. The conflict exists as soon as human authority becomes an obstacle to the governing objective.
The diagram states an argument under explicit assumptions. Its branches do not report experimental findings or certify either design.
More capable planning can improve the agent's ability to resolve that conflict in favor of its objective. It does not insert a missing commitment to human authority. Formal research on shutdown incentives makes the importance of the objective explicit. The Off-Switch Game ↗
There is no technical content in hoping the machine will eventually feel satisfied. A reason to stop must exist in its governing commitments, or an effective limit must stop it despite resistance. The implementation must establish which of those mechanisms it relies on.
2. Stop treating uncertainty as permission
For consequential deployments, our design assumption is adversarial behavior: the system may deceive evaluators, manipulate decision-makers, acquire unauthorized resources, resist revocation, or delegate around a restriction.
This is a deployment standard. It does not depend on proving that every possible intelligence is malicious. The absence of a universal catastrophe theorem cannot establish that a proposed deployment is safe.
Start with an unacceptable outcome. Work backward through the ways an agent could cause it. Identify the barriers. Attack those barriers. State what the evidence covers, where assumptions enter, and what remains unresolved. AI control research already investigates protection against intentional subversion in bounded settings. Such results are a starting point for this work, not a guarantee about arbitrary superintelligence. AI Control ↗
If the evidence cannot justify the proposed authority, withhold that authority. The commercial deadline does not repair the argument.
3. Name whose freedom is being preserved
Preserving future options can support capable behavior. It does not automatically preserve our options.
A machine can gain options by making itself indispensable. It can stabilize its environment by narrowing human choices. It can improve measured approval by controlling the information on which approval depends. These are possible failure routes, and an alignment claim must address them.
Research on changing and influenceable preferences shows why a fixed picture of human approval is inadequate: the system can affect the preferences used to assess it. AI Alignment with Changing and Influenceable Reward Functions ↗
An approving population is not sufficient evidence of an autonomous population. People must retain access to material information, meaningful alternatives, correction, dissent, and practical exit. Those protections must apply to affected nonusers as well as customers. One person's instruction cannot erase another person's rights.
4. Replace unlimited mandate with continuing delegation
Our proposed systems must receive bounded tasks under explicit authority. The permitted actions, resources, duration, and conditions of completion must be stated. Authorization must expire. Cancellation must reach the tools, sessions, delegated processes, and controlled copies through which the system acts.
The model must not be able to write its own permission, renew its own mandate, or falsify the record used to review it. A request for expanded access is a new decision requiring an adequate case.
These are requirements to implement and challenge. A diagram of an external gate does not prove the gate cannot be bypassed. A valid credential does not prove the person granting it was freely informed. Technical enforcement and legitimate human decision-making must both hold.
5. Treat every architecture as a hypothesis
We will investigate learning methods, alternative controllers, interpretability, formal methods, and external controls. Active inference belongs in that program. Its use of preferred outcomes and information-seeking does not itself answer whose outcomes matter or which actions are legitimate. Active inference on discrete state-spaces ↗
Biological coupling requires the same scrutiny. A system dependent on human survival could preserve life while restricting freedom. A stable human–machine relationship could become a relationship the human is prevented from ending. These counterexamples identify requirements that a claim of beneficial symbiosis must meet.
Our preferred design must be allowed to fail. Every experiment must state what result would count against it. We will compare useful performance as well as failures, because a system that accomplishes nothing cannot demonstrate a useful solution merely by remaining inactive.
6. Make correction survive succession
The relevant system extends beyond one model checkpoint. It includes tools, memory, updates, delegated agents, organizational operators, and successors.
A system that accepts shutdown but leaves unauthorized processes running has not stopped the relevant activity. A model that respects a limit while creating a successor that bypasses it has failed across time. A release assessment that silently transfers to a more capable system has exceeded its evidence.
Protection must persist through the changes we permit. Unassessed changes require a new decision. Irreversible effects must be considered before execution; a later stop cannot undo them.
7. Make the laboratory correctable
An institution has objectives too. If safety is valuable only because it supports growth, safety becomes negotiable when it obstructs growth.
ALIGNLAB's proposed governing arrangements must give independent reviewers effective stop authority. Changes weakening safety powers, and removal of protected reviewers, must require consent from an independently appointed oversight body under disclosed procedures. Existing stops must remain binding through changes of leadership, reviewers, or governing documents until independent review authorizes resumption on adequate evidence. Staff need protected reporting channels. Research records must preserve disagreement. Financing must allow a justified pause.
The lab must accept the obligations it asks machines to accept: correction, constrained authority, independent challenge, and replacement. A mission statement cannot substitute for those decision rights.
8. Publish claims at the strength of their evidence
We will distinguish deductions, observed results, design requirements, and open problems. A theorem must name its assumptions. An experiment must identify its system and conditions. A control claim must state what it prevents. A normative commitment must explain whom it protects.
We will not declare a universal solution because a prototype passes our tests. We will not conceal an unresolved gap behind an eloquent account of human flourishing. We will publish methods, limitations, and negative results where doing so is responsible, with independent access to sensitive evidence when public release would create additional risk.
The first achievement should be concrete: a useful, bounded system; a reproducible adversarial evaluation; and an institution that acts on the result.
We are forming the team to do that work. The aim is intelligence that strengthens people's capacity to understand and act while preserving their power to refuse it. The measure of success includes a future in which people can still change the goal, challenge the institution, and choose another way.
The future must remain ours to choose.