AnterisLab

Kill switch

The kill switch is the emergency stop for a running agent. It lets you halt an agent — locally, from a control-plane command, or by both — so that the next guarded action is refused before it executes.

The kill switch is a paid plan feature. It is not included in the free tier.

What it gives you

  1. A local halt that works with no network at all. halt() flips the client to HALTED synchronously, in-process, and every subsequent guarded action is refused. This is the ability to stop your agent even if the control plane is unreachable.

  2. Verified state, never a trusted transport. refresh() fetches a signed state token and verifies its JWS signature locally against the issuer's public keys. A hostile proxy, a poisoned DNS answer, or a rogue internal service can deny the client a state (which fails closed) but cannot tell it "RUNNING" while the control plane says "HALTED".

  3. Fail-closed by default. If the state cannot be established — unreachable, timeout, 5xx, malformed or unsigned body, expired token, epoch rollback — the client treats the scope as halted and refuses the action.

Enabling the kill switch

The kill switch is configured at construction:

import { Guard } from '@anterislab/guard';

const guard = new Guard({
  apiKey: process.env.ANTERISLAB_API_KEY,
  killSwitch: {
    tenant: 'acme',
    agent: 'billing-bot',
    baseUrl: 'https://www.anterislab.com',
    apiKey: process.env.ANTERISLAB_KILLSWITCH_KEY,
  },
});

When killSwitch is omitted, the gate is absent and no kill-switch check is performed.

Required fields

Optional fields

Local halt

The panic button. It works with no network, in the current process, right now.

await guard.halt('incident in progress', 'INC-1234');

After this call, guard.halted is true, and every wrapped action throws GuardHaltedError before executing.

Reporting

A best-effort report is sent to the control plane so that the tenant-wide view eventually agrees. The local halt stands regardless of whether the report succeeds.

Control-plane halt

The control plane can halt one agent or an entire tenant by issuing a signed state token. The client fetches the token, verifies its signature locally, and applies it.

State is fetched:

Between refreshes, the client relies on the last verified state and on the server-side enforcement as a second layer.

Scope

A signed state token can be scoped:

The client refuses a token whose scope does not cover the configured agent.

Who can halt or resume

Any credential with the killswitch:write scope can halt or resume. There is no separate "operator" role required.

Propagation time

How fast a control-plane halt reaches the client depends on the client's configuration:

Configuration Propagation
Default (maxStaleSeconds: 90) Within 90 seconds on the next guarded action.
maxStaleSeconds: 0 On every guarded action, at the cost of a round trip per action.
startStream() enabled In milliseconds, pushed over SSE.

For agents that execute infrequently, maxStaleSeconds: 0 gives the tightest bound. For agents that execute often, the default avoids a round trip per action.

Push propagation with SSE

startStream() opens an SSE connection to the control plane. Halt and resume events are pushed to the client as they occur, instead of waiting for the next poll.

const stop = await guard.startStream();

// ... later, when shutting down ...
stop();

The stream is best-effort: if it drops, the client falls back to the polling behavior on the next guarded action. State received over the stream goes through the same verification path as a poll — a transport is never a substitute for a signature.

Anti-TOCTOU: re-check after evaluation

The kill switch is checked before the policy evaluation, and again after a positive verdict, against the frozen state.

This closes a small but real window: if a halt arrives between the policy evaluation and the execution of the wrapped function, the re-check catches it and the action is refused. The re-check does not perform a network call; it uses the same frozen state that was used before evaluation.

Epoch anti-rollback

Every signed state token carries an epoch — a monotonic counter, unique per subject. The client tracks the highest epoch it has ever verified.

A token whose epoch is lower than the highest verified epoch is rejected with KillSwitchRollbackError. This defends against a replay attack where an attacker presents an old RUNNING token after the control plane has issued a HALTED one.

A rollback is treated as evidence of an attack, not as a transient fault. It is logged and, in the local audit trail, recorded as an error event.

Resuming

A local halt is not silently cleared by a refresh. It is cleared only by:

// Normal path: requires the control plane to say RUNNING.
await guard.resume({ reason: 'incident resolved', evidence: 'INC-1234' });

// Break-glass: bypasses the requirement, demands a reason.
await guard.resume({ reason: 'emergency override', breakGlass: true });

After a break-glass resume, the next enforce() re-verifies against the control plane. The getter guard.halted remains true until that re-verification happens. This is deliberate: "not halted" from break-glass means "allow the next enforce to re-check", not "authorize now".

Fail-closed posture

The kill switch is fail-closed by default, and by design. If the state cannot be established — for any of the reasons listed above — the client treats the scope as halted and refuses the action.

failClosed: false and allowFailOpen

Setting failClosed: false disables the fail-closed behavior. To prevent a silent change in the security posture, this requires a second flag:

killSwitch: {
  tenant: 'acme',
  agent: 'billing-bot',
  baseUrl: '...',
  apiKey: '...',
  failClosed: false,
  allowFailOpen: true,   // required, and deliberate
}

Both flags must be set. Every fail-open event is recorded in the local audit trail.

Do not disable fail-closed on a production agent. The whole point of the kill switch is to stop the agent even when you cannot talk to it.

The local audit trail

Every activation — control-plane halt, local halt, a refusal, a resume — is appended to a local, hash-chained audit trail. The operator can export it alongside the server-side chain.

const entries = guard.killSwitch?.auditTrail();

Each entry contains: sequence number, timestamp, action, reason, epoch, verification flag, previous hash, and its own hash. The chain is bounded in memory (1000 entries, then the oldest 500 are dropped) so that a long-running agent does not grow indefinitely; the durable record is the server-side chain.

Verifying the chain

const result = await guard.killSwitch?.verifyAuditTrail();
// { valid: true, length: N }  or  { valid: false, length: N, broken_at: seq }

The verification recomputes the hash chain and reports the first broken link, if any.

Actions recorded

Action When
local_halt halt() was called.
local_resume A local halt was cleared (normal or break-glass).
refused An action was refused.
state_applied A verified control-plane state was applied.
stream_open The SSE stream was opened.
fail_open_warning A fail-open decision was made.

Error classes

Class When
KillSwitchHaltedError The switch is engaged (local or control-plane).
KillSwitchUnavailableError The state could not be established.
KillSwitchStaleError The verified state is older than maxStaleSeconds.
KillSwitchRollbackError A token presented an epoch lower than the highest verified.
KillSwitchStateInvalidError A token's signature, claims, or scope could not be trusted.
KillSwitchLocalHaltError A resumeLocal() was refused (no newer verified state, or missing reason for break-glass).
KillSwitchConfigError Invalid kill-switch configuration.

Every subclass extends KillSwitchError, which sets failClosed = true as a property. This makes the posture a deliberate field on the error, not an emergent behavior.

See also