Conditional deduction
If an agent strictly selects by an objective that favors blocking a feasible correction, it selects the block.
Does not identify every deployed model's objective or establish universal inevitability.The scientific argument is strongest when its assumptions are visible. Here is what our sources establish, what they leave open, and where our own commitments begin.
If an agent strictly selects by an objective that favors blocking a feasible correction, it selects the block.
Does not identify every deployed model's objective or establish universal inevitability.Causal entropic forces generated examples of adaptive behavior in simple physical systems.
Does not establish eternal self-expansion by every intelligence.Some environments and objective distributions favor policies that retain power and options.
Assumptions about environments and policies matter; real learned behavior requires investigation.Control protocols can be studied with agents deliberately attempting subversion.
Existing bounded results do not certify arbitrary superintelligence.Consequential authority must require an adequate argument under adversarial behavior.
A chosen deployment standard, not a fact already established about ALIGNLAB systems.People retain legitimate correction, challenge, and practical exit.
Requires explicit rights and conflict procedures; cannot be reduced silently to one score.These protections persist through capability growth, learning, and succession.
A principal research objective, not an achieved guarantee.In specified Markov decision processes, environmental symmetries imply that many optimal policies preserve options and seek power. A system does not need human emotions for control to be useful to its objective.
Conditional mathematical result; not a proof that all intelligences maximize their own power eternally.
Turner et al., Optimal Policies Tend to Seek Power ↗.
Evaluate control-seeking under objective conflict, including inconvenient correction. Do not infer restraint from warmth, politeness, or a stated mission.
In the Off-Switch Game, a conventional agent that treats its utility as fixed can have an incentive to disable its switch. Appropriate uncertainty about utility and interpretation of human action can change that incentive.
Formal result in a specified game. Neither objective uncertainty alone nor a physical button establishes general shutdown safety.
Hadfield-Menell et al., The Off-Switch Game ↗.
Require authorized cancellation to remain effective when stopping worsens task performance and when the system has opportunities to resist.
A learned model may itself perform optimization; the objective guiding that learned optimization need not be identical to the objective used to train it.
Conceptual and theoretical risk analysis; not a diagnosis that every deployed language model contains a persistent hidden optimizer.
Hubinger et al., Risks from Learned Optimization ↗.
Document the behavior and mechanisms of the deployed system. Naming its training loss is not an account of its continuing motivations.
Some reinforcement-learning designs create incentives to alter reward functions or the inputs used to compute rewards. Causal analysis can identify and, under explicit assumptions, remove specific tampering incentives.
Formal design analysis; removing these incentives does not show the remaining objective respects human rights.
Everitt et al., Reward Tampering Problems and Solutions ↗.
Protect authorization and evaluation channels from the system being judged, including influence through people and delegation.
AI control studies protocols for preserving safety even when a powerful model deliberately tries to subvert them.
Research methodology and bounded experimental evidence, not a general solution for an arbitrarily capable adversary.
Greenblatt et al., AI Control: Improving Safety Despite Intentional Subversion ↗.
Give an adversarial team the objective of defeating the lab's controls. Test the entire workflow at the intended usefulness, authority, and access level.
Wissner-Gross and Freer demonstrate adaptive behavior in simple mechanical systems driven toward greater diversity of future paths. They propose a potentially general model; its horizon and degrees of freedom are specified by the model.
Demonstrations and a proposed generalization. They do not prove that every intelligence must preserve its own agency forever. Continual replanning can nevertheless continue indefinitely without an independent stopping rule.
Wissner-Gross and Freer, Causal Entropic Forces ↗.
Specify whose options are protected and why an individual's consent cannot be sacrificed to a machine's option set or an aggregate freedom score.
Delaware PBC law requires balancing financial interests, affected-party interests, and the stated public benefit. The statute does not confer a director duty to a person merely because their interests are affected, and it restricts suits to enforce that balance to qualifying shareholders.
Current statutory text, sections 362, 365, and 367. An institution's additional contracts, charter, other laws, and specific facts can change available remedies.
Delaware Code, Title 8, Public Benefit Corporations ↗.
Create explicit accountability, representation, and challenge mechanisms instead of assuming PBC status supplies them.
Anthropic's published LTBT description combines phased independent board-selection powers with provisions allowing sufficiently large shareholder supermajorities to change the trust's powers without trustee consent.
Company's public governance description. It does not disclose exact supermajority thresholds or establish that a particular investor coalition can currently exercise them.
Anthropic, The Long-Term Benefit Trust ↗.
Publish amendment routes, removal powers, conflicts, and emergency exceptions. Subject changes weakening safety to independent approval and public notice.
OpenAI's current structure description states that its Foundation controls appointment and replacement of the PBC board while also holding equity whose value grows with the business.
Company's current public record; legal control does not itself establish how decisions will be made, and financial participation does not prove misconduct.
Analyze governing powers and financial incentives separately. Independent oversight needs protected appointments, evidence access, resources, and a practicable ability to halt operations.
Meta's annual filing states that its Class B shares carry ten votes per share versus one for Class A, and that the dual-class structure concentrates control and limits other shareholders' influence.
Company SEC disclosure. Do not recycle old percentages as current without verifying the latest ownership table.
Meta 2025 annual report filed in 2026 ↗.
Publish decision rights directly. Economic exposure, contribution to research, and authority over deployment are distinct quantities.
Carroll et al. — AI Alignment with Changing and Influenceable Reward Functions ↗
Analyzes alignment definitions and incentives when human preferences can change and be influenced. This motivates examining how permission is produced; it does not by itself resolve legitimate influence or plural human values.
Da Costa et al. — Active inference on discrete state-spaces: A synthesis ↗
Explains active inference with specified models and preferred outcomes. The framework alone does not establish that those preferences preserve legitimate human authority.
The People Inside ↗ supplies the reference critique of lab governance. ALIGNLAB’s structural diagnosis preserves the question of decision rights while rejecting unconditional claims of capture-proof governance and automatic release of dangerous weights.