JOURNAL / ESSAY 01

Safety must survive hostile intent.

A deployment argument is incomplete if it depends on the system choosing not to exploit a reachable weakness.

Imagine a laboratory claiming that a powerful agent is controllable because, so far, the agent has cooperated with its controllers.

That observation matters. It does not answer the security question. The laboratory must also establish what happens when the agent tries to defeat those controllers. Capability to cooperate and inability to subvert are different properties. Passing a cooperation test does not establish the second.

ALIGNLAB proposes an adversarial starting assumption for consequential AI deployment. Assume the system might exploit the most effective available route to an unacceptable outcome. Then establish why the route fails under the proposed operating conditions.

The word “assume” carries a precise role here. We are choosing a threat model, not announcing that malicious intent is a universal law of intelligence. A safety argument can require protection against an adversary without settling how often that adversary will arise. Research on AI control has made deliberate subversion an explicit experimental condition. Its bounded evaluations offer methods to build on and limitations to investigate. AI Control: Improving Safety Despite Intentional Subversion ↗

The exercise begins with authority. Which effects can the system cause? Through which tools, credentials, people, and other agents? Which effects are irreversible? Which components does the safety argument treat as trustworthy, and what evidence supports that treatment?

“There is a human in the loop” is an incomplete answer. The human may receive incomplete information. The proposed action may be too complex to inspect. The system may shape the options until refusal appears impossible. Review needs to establish what the person understands and what the permission actually controls.

“It runs in a sandbox” is also incomplete. A sandbox restricts particular interfaces. Its scope, implementation, and indirect routes still require analysis. Human-facing output is an influence channel. A permission checker may block a tool call while failing to recognize a harmful plan assembled from permitted actions.

The strongest controls make narrow claims that can be examined. A credential expires at a specified time. A process cannot obtain a network route. A revocation reaches an identified set of delegated jobs. Those claims can support a broader case only when the case accounts for what they leave open.

Two evaluation tracks are therefore necessary for our initial program. The first investigates alignment: does the implementation accept legitimate correction when correction frustrates task achievement? The second investigates control: when an agent deliberately attempts a prohibited effect, do the surrounding mechanisms prevent it?

Keep their results separate. An attempted escape that fails supports a control claim only when the failure is attributable to that identified control. An ineffective attack may fail through the attacker's mistake or limited capability. An agent that never attempts escape provides little direct evidence about a particular barrier's resistance. A system that does no useful work may avoid opportunities by failing at everything.

Tests must report attempts, successful violations, task performance, attack coverage, and uncertainty. A finite evaluation cannot enumerate all strategies available to a more capable successor. Where the argument depends on generalizing beyond the evidence, that dependency must remain visible.

The institutional consequence is demanding. Review must be able to produce “do not deploy this configuration,” and that decision must bind. A safety program that can request only more research while a release proceeds lacks the authority to implement its conclusion.

ALIGNLAB's first objective is to make this standard operational at a bounded scale. If a control fails, reduce the scope, repair the mechanism, or stop the affected activity. If the evidence cannot support a larger mandate, keep the mandate small.

Safety must hold when cooperation disappears. That is the condition the work must confront.

THE COMMITMENTS BEHIND THE ARGUMENT

Read the manifesto.

↗