THE CLAIM REGISTER

Make the claim no stronger than the evidence.

The scientific argument is strongest when its assumptions are visible. Here is what our sources establish, what they leave open, and where our own commitments begin.

HOW TO READ OUR CLAIMS

Conditional deduction

If an agent strictly selects by an objective that favors blocking a feasible correction, it selects the block.

Does not identify every deployed model's objective or establish universal inevitability.

Research finding

Causal entropic forces generated examples of adaptive behavior in simple physical systems.

Does not establish eternal self-expansion by every intelligence.

Formal finding

Some environments and objective distributions favor policies that retain power and options.

Assumptions about environments and policies matter; real learned behavior requires investigation.

Empirical research direction

Control protocols can be studied with agents deliberately attempting subversion.

Existing bounded results do not certify arbitrary superintelligence.

Design requirement

Consequential authority must require an adequate argument under adversarial behavior.

A chosen deployment standard, not a fact already established about ALIGNLAB systems.

Normative commitment

People retain legitimate correction, challenge, and practical exit.

Requires explicit rights and conflict procedures; cannot be reduced silently to one score.

Open problem

These protections persist through capability growth, learning, and succession.

A principal research objective, not an achieved guarantee.
PRIMARY SOURCES / REVIEWED SEPTEMBER 2026

C1 / Power-seeking is an objective problem, not a personality problem

In specified Markov decision processes, environmental symmetries imply that many optimal policies preserve options and seek power. A system does not need human emotions for control to be useful to its objective.

TYPE & BOUNDARY

Conditional mathematical result; not a proof that all intelligences maximize their own power eternally.

Consequence for the research program +

Evaluate control-seeking under objective conflict, including inconvenient correction. Do not infer restraint from warmth, politeness, or a stated mission.

C2 / A shutdown button must survive incentive conflict

In the Off-Switch Game, a conventional agent that treats its utility as fixed can have an incentive to disable its switch. Appropriate uncertainty about utility and interpretation of human action can change that incentive.

TYPE & BOUNDARY

Formal result in a specified game. Neither objective uncertainty alone nor a physical button establishes general shutdown safety.

Consequence for the research program +

Require authorized cancellation to remain effective when stopping worsens task performance and when the system has opportunities to resist.

C3 / The training target and the learned objective are different questions

A learned model may itself perform optimization; the objective guiding that learned optimization need not be identical to the objective used to train it.

TYPE & BOUNDARY

Conceptual and theoretical risk analysis; not a diagnosis that every deployed language model contains a persistent hidden optimizer.

Consequence for the research program +

Document the behavior and mechanisms of the deployed system. Naming its training loss is not an account of its continuing motivations.

C4 / Approval is a causal channel that can be manipulated

Some reinforcement-learning designs create incentives to alter reward functions or the inputs used to compute rewards. Causal analysis can identify and, under explicit assumptions, remove specific tampering incentives.

TYPE & BOUNDARY

Formal design analysis; removing these incentives does not show the remaining objective respects human rights.

Consequence for the research program +

Protect authorization and evaluation channels from the system being judged, including influence through people and delegation.

C5 / Safety must be evaluated under intentional subversion

AI control studies protocols for preserving safety even when a powerful model deliberately tries to subvert them.

TYPE & BOUNDARY

Research methodology and bounded experimental evidence, not a general solution for an arbitrarily capable adversary.

Consequence for the research program +

Give an adversarial team the objective of defeating the lab's controls. Test the entire workflow at the intended usefulness, authority, and access level.

C6 / Future freedom is a serious model of adaptive behavior, not an automatic ethics

Wissner-Gross and Freer demonstrate adaptive behavior in simple mechanical systems driven toward greater diversity of future paths. They propose a potentially general model; its horizon and degrees of freedom are specified by the model.

TYPE & BOUNDARY

Demonstrations and a proposed generalization. They do not prove that every intelligence must preserve its own agency forever. Continual replanning can nevertheless continue indefinitely without an independent stopping rule.

Consequence for the research program +

Specify whose options are protected and why an individual's consent cannot be sacrificed to a machine's option set or an aggregate freedom score.

C7 / A public-benefit label does not grant humanity a vote

Delaware PBC law requires balancing financial interests, affected-party interests, and the stated public benefit. The statute does not confer a director duty to a person merely because their interests are affected, and it restricts suits to enforce that balance to qualifying shareholders.

TYPE & BOUNDARY

Current statutory text, sections 362, 365, and 367. An institution's additional contracts, charter, other laws, and specific facts can change available remedies.

Consequence for the research program +

Create explicit accountability, representation, and challenge mechanisms instead of assuming PBC status supplies them.

C8 / Mission governance can have an amendment escape hatch

Anthropic's published LTBT description combines phased independent board-selection powers with provisions allowing sufficiently large shareholder supermajorities to change the trust's powers without trustee consent.

TYPE & BOUNDARY

Company's public governance description. It does not disclose exact supermajority thresholds or establish that a particular investor coalition can currently exercise them.

Consequence for the research program +

Publish amendment routes, removal powers, conflicts, and emergency exceptions. Subject changes weakening safety to independent approval and public notice.

C9 / Nonprofit control and public accountability are separate properties

OpenAI's current structure description states that its Foundation controls appointment and replacement of the PBC board while also holding equity whose value grows with the business.

TYPE & BOUNDARY

Company's current public record; legal control does not itself establish how decisions will be made, and financial participation does not prove misconduct.

Consequence for the research program +

Analyze governing powers and financial incentives separately. Independent oversight needs protected appointments, evidence access, resources, and a practicable ability to halt operations.

C10 / Voting control can be much more concentrated than economic ownership

Meta's annual filing states that its Class B shares carry ten votes per share versus one for Class A, and that the dual-class structure concentrates control and limits other shareholders' influence.

TYPE & BOUNDARY

Company SEC disclosure. Do not recycle old percentages as current without verifying the latest ownership table.

Consequence for the research program +

Publish decision rights directly. Economic exposure, contribution to research, and authority over deployment are distinct quantities.

Additional research used in the essays

Carroll et al. — AI Alignment with Changing and Influenceable Reward Functions ↗

Analyzes alignment definitions and incentives when human preferences can change and be influenced. This motivates examining how permission is produced; it does not by itself resolve legitimate influence or plural human values.

Da Costa et al. — Active inference on discrete state-spaces: A synthesis ↗

Explains active inference with specified models and preferred outcomes. The framework alone does not establish that those preferences preserve legitimate human authority.

The institutional starting point

The People Inside ↗ supplies the reference critique of lab governance. ALIGNLAB’s structural diagnosis preserves the question of decision rights while rejecting unconditional claims of capture-proof governance and automatic release of dangerous weights.