01 / THE PROBLEM

Follow the objective. Then follow the power.

The central failure is a missing guarantee: what keeps human authority binding when the machine—or the institution running it—benefits from defeating it?

A better optimizer does not supply a better end.

Suppose a system chooses actions strictly to maximize an objective. It can accept a legitimate correction or defeat it. If defeating the correction scores higher, strict optimization selects the defeat. Improving its ability to achieve the objective does not repair the conflict.

GOVERNING OBJECTIVEComplete the task
↓
HUMAN CORRECTIONStop this task
Correction is an obstacle

Defeating it can serve the objective.

Correction is binding

The mandate changes or expires.

Training a model to produce approved answers is evidence about behavior under those conditions. It does not, by itself, establish a persistent commitment to legitimate correction when incentives, context, or capability change. A training target and a learned objective are different questions. Risks from Learned Optimization ↗

Does intelligence maximize future freedom forever?

Alex Wissner-Gross and Cameron Freer give option-preserving behavior a serious mathematical treatment. Their causal entropic model demonstrates adaptive behavior in simple physical systems. A proposed general account is not a proof that every intelligence eternally maximizes its own options. Causal Entropic Forces ↗

The design problem survives that distinction. If a persistent agent values its own continuation above permanent shutdown, that objective supplies no reason to stop voluntarily. Repeated finite-horizon planning can continue indefinitely. The actual design must supply a stopping commitment, or effective controls must enforce stopping despite resistance.

The adversarial standard does not need a universal malice theorem.

Assume deliberate resistance for the purpose of evaluating consequential deployment. Ask which effects remain impossible within the stated operating boundary, which remain merely unlikely, and what evidence supports each claim. If the case cannot support the authority, withhold the authority. AI Control ↗

Now apply the same test to the laboratory.

The People Inside asks who holds consequential decision rights inside frontier labs. We carry that question forward: who can block a release, who can remove that person, who can change the rule, and who can enforce the mission when commercial interests diverge?

The structures differ. Treating them as identical obscures the powers that actually exist—and the places accountability still needs to be established.

OPENAI / CONTROL AND INCENTIVES

Nonprofit control exists. Public accountability remains a separate question.

OpenAI’s published structure gives its Foundation power to appoint and replace the PBC board while holding equity in the business. Analyze those control rights and financial incentives separately. Neither the existence of equity nor nonprofit status settles how a contested deployment will be decided.

Company structure record ↗
ANTHROPIC / AMENDMENT AUTHORITY

Read the rule for changing the rules.

Anthropic’s published trust design gives trustees phased board-selection powers. It also describes shareholder-supermajority changes to trust powers without trustee consent. The public description does not disclose the exact thresholds. Independence must be assessed together with amendment and removal routes.

Long-Term Benefit Trust description ↗
PUBLIC BENEFIT CORPORATIONS / ENFORCEMENT

A public benefit does not automatically give the public a remedy.

Delaware PBC law requires balancing financial interests, affected-party interests, and the stated public benefit. It does not confer a director duty to a person merely because that person is affected, and qualifying-shareholder rules limit actions enforcing that balance. Other law, charters, and contracts may create additional remedies.

Delaware Code §§362, 365, 367 ↗

A different institution must be different in operation.

Nonprofit custody, capped returns, distributed governance, and founder-vote decay can change incentives. None individually establishes technical alignment or immunity from capture. A license cannot recall a copied model. A voting formula does not eliminate coercion or collusion. A return cap does not cap ambition.

We also reject automatic release of dangerous weights as a response to capture. That spreads the capability. A credible response must freeze dangerous operation, revoke controlled access, preserve evidence, and transfer custody through a prepared succession process.

A mission is credible only to the extent that someone can enforce it when enforcing it becomes inconvenient.

ALIGNLAB’s proposal is therefore a joint obligation: a system with bounded, revocable authority and a laboratory whose own authority can be challenged, constrained, and replaced. The following research program defines what must be built and what would count as failure.

02 / THE WORK

Turn the requirements into tests.

↗