Security

Security model

How Ant authenticates, authorizes, and limits what a peer can do.

How Ant authenticates, authorizes, and limits what a peer can do.

Identity

  • Every principal holds an ed25519 keypair. The public key is the NodeID, in base32 (no padding). Account keys live at ~/.ant/identity; a machine's key lives on the machine (~/.ant/worker/agent-identity, or wherever the worker's --identity points).
  • The transport binds an endpoint with the identity's seed and verifies the bound endpoint id equals the NodeID, so a peer cannot present a key it does not hold.

Transport

  • iroh = QUIC + TLS 1.3 with mutual authentication: both sides prove possession of their key. Dialing is by NodeID; no open ports, no host certificates.
  • A self-hosted relay, when configured, only forwards encrypted bytes.
  • Private keys are 0600. The account's portable credential bundle is passphrase-encrypted by default (PBKDF2-HMAC-SHA256, 210k iterations, AES-256-GCM), and load re-checks that the key matches the profile NodeID.

Authorization: the daemon is the authority

Mutual TLS authenticates the channel; it does not authorize the caller. The machine daemon resolves the caller's NodeID against its roster and enforces a method → capability table:

MethodRequires
ping, whoamipublic
invites.redeempublic, gated by a signed invite token
users.list, containers.list, containers.logs, images.list, volumes.list, metrics.get, caddy.status, tools.probe, swarm.status, swarm.nodes, swarm.services, swarm.service.tasks, swarm.service.logsread (viewer+)
users.add, users.setRole, users.remove, users.update, invites.revoke, state.export, audit.list, swarm.init, swarm.leave, swarm.node.update, swarm.node.rm, swarm.joinTokencolony manage, and the caller must be allowed to grant/remove/modify the target's role
state.importownership (owner only)
deploy.run, deploy.rollback, deploy.commit, image.load, build.run, compose.apply, compose.down, stack.apply, stack.down, route.apply, containers.action, containers.remove, volumes.remove, swarm.service.scale, swarm.service.restart, swarm.service.update, swarm.service.rm, handover.export, tunnel.opendeploy (deployer, ci, admin, owner)

Role capabilities: owner (all) · admin (all but ownership transfer) · deployer (read, deploy, osuser) · viewer (read) · ci (read, deploy). Ownership transfer (state.import) is owner-only.

Fail closed. With no roster and no boot owner, every gated method is refused. The first owner is seeded out of band:

  • ant-worker --owner <nodeid> (repeatable), the bootstrap/provisioning owner, emitted by ant nest bootstrap; or
  • writing the owner into the machine's state.json as root.

Invites

  • An invite is signed by the inviter's account key and carries the destination machine NodeID, the role, and an expiry.
  • Redemption is server-side (invites.redeem): the daemon verifies the signature, expiry, that the invite is for this machine, the inviter's current role, and the grant matrix, and that the nonce is unused.
  • Single-use is enforced by the daemon (state.json keeps consumed nonces until their expiry). The invitee need not be on the roster yet.
  • Revocable: invites.revoke (colony manage) records a nonce as revoked until its expiry, and redemption rejects it, so a leaked token can be killed.

Least privilege on the machine

  • The worker runs unprivileged and never invokes sudo. Privileged work on a machine is operator-run and executes exact argv, never a shell: ant nest reconcile --apply runs as root, and the SSH provisioning path (ant nest machines add --host, ant nest bootstrap --host) runs its privileged steps on the target through the SSH user's sudo.
  • Docker calls use exec.Command with argument slices, not sh -c. The image is passed after --, so a crafted reference cannot be read as a flag.
  • Server-side validation: container names, compose project names, and network names must match ^[a-zA-Z0-9][a-zA-Z0-9_.-]{0,62}$; image refs may not start with - or contain whitespace. This blocks path traversal (../…) and flag injection.
  • Container actions address the container by exact name only.

The deploy capability is root-equivalent on the machine

This is the most important thing to understand before inviting someone:

  • deployer and ci can run any image, with arbitrary environment and arbitrary -v mounts. A deployer can mount the worker's own state directory (~ant/.ant) into a container and read the machine's private identity, or mount the host filesystem. In group Docker mode the worker is in the docker group, which is host root; in rootless mode it is full control of the worker's unix user (including its NodeID key).
  • Treat deployer/ci as machine-level trust, not as a read-mostly role. Give the admin/owner roles to people you would trust with SSH root on that machine, and prefer viewer for read-only access.
  • run: compose service images can also declare arbitrary mounts and privileges; the compose file comes from a project config the deployer controls.

The role matrix (read, deploy, colony.manage, ownership.transfer, osuser) is enforced server-side, and the worker never runs privileged host commands, but container execution itself is the escape hatch, as with any Docker-based deploy tool.

Routes and ports

  • A published port binds loopback by default, so a deployed app is reachable through the machine's reverse proxy (and on the machine) but not from the network. deploy.publish: all opts into binding every interface.
  • route.apply accepts only valid hostnames (no wildcards, ports, schemes, or paths), only loopback upstreams (the machine's own containers), and rejects a host (or host+path) already routed by another app, so a deployer cannot shadow or hijack a teammate's domain.
  • The worker's Caddy admin API is bound to 127.0.0.1:2019 with an origin allowlist (origins 127.0.0.1:2019 localhost:2019), so a browser request carrying a foreign Origin (a page on another domain that resolves to loopback, i.e. DNS rebinding) cannot reprogram routing, and a request with a foreign Host is refused. On a multi-user host, prefer the distribution's caddy-only unix socket instead.
  • The dashboard is operator-local by design: one dashboard per laptop, reading that operator's own files, never hosted on a server. It binds loopback and has no authentication there (the trust boundary is the local machine). Binding it to a non-loopback address is an escape hatch for reaching your own dashboard from another device: it is refused unless a password is set (ant ui passwd), and is then wrapped in HTTP Basic auth and should sit behind TLS. A successful login issues a signed, HttpOnly session cookie whose key derives from the password hash, so the expensive hash runs at login rather than on every request, and changing the password invalidates existing sessions.
  • On a loopback bind the dashboard also applies a Host allow-list (rejecting a request whose Host is not loopback, DNS rebinding) and a same-site guard that rejects state-changing requests a browser marks as cross-site. JSON request bodies are capped at 8 MiB.
  • Deploys and compose files that reach the worker's state/identity directory, the docker socket, or the host root are refused, on the worker and on the client (so a local run and a remote run behave the same). The check is boundary-aware and covers ancestors of the state directory (e.g. /home/ant exposes /home/ant/.ant/worker/agent-identity). A project-relative mount such as . or ./data is allowed in a compose file (it is relative to the compose/project directory, not the worker's home) while a run: container deploy refuses relative bind sources outright, because the worker has no project directory to resolve them against. For compose this covers volumes, top-level secrets:/configs: file: sources, and it also rejects privileged: true and host devices:. A resolved compose document that still contains include: is refused on the worker: compose expands it at up time, after the guards, so its services could bypass every check.
  • Untrusted build contexts are extracted with an 8 GiB cap and reject absolute paths, .. traversal, hard links, and symlinks that leave the context. The transport caps a request frame at 2 MiB (a general frame at 8 MiB), a blob at 64 GiB, and bounds connections and streams per connection.
  • Every privileged RPC is written to a tamper-evident audit log (ant nest audit): each entry chains to the previous by hash and each rotated segment links to the one before it, so an edited entry or a deleted segment fails verification. (The newest lines can still be truncated by anyone who can write the state directory, and deleting the oldest retained segment is indistinguishable from normal rotation.) The log rotates at 8 MiB, and reading it is an admin action that does not itself append.
  • Secret values are never stored: a from: provider (1Password op://, a generic cmd:, or env:) is resolved at deploy time and only its warning is surfaced on failure.
  • Tunnels (ant nest tunnel, or the dashboard's Tunnel action) forward bytes to a published loopback port of a container ant manages and nothing else, so they add no reach beyond the deploy capability: a peer cannot use one to reach the docker socket, the Caddy admin API, or another local listener. See Tunnels.

Protocol version

Every RPC carries an API version, and both sides reject a mismatch with an actionable message (peer speaks protocol vN, this worker speaks v1; upgrade the older side) instead of failing as an unknown method. A mixed-version fleet therefore fails closed on the first call, which is why upgrading worker and CLI together matters (see the upgrade note in Known issues). Before the first tagged release the protocol is not stable and the version stays 1; it is bumped only once there are deployed workers that cannot simply be rebuilt.

At rest

  • ~/.ant/config.json, the machine state.json, and identity files are 0600; the machine state is written atomically.
  • Build-time secrets are resolved from the environment and never baked into images or the lockfile. Runtime secrets never appear in command argv or a persisted compose document: a container deploy (local or remote) passes them through the docker process environment (a bare docker run -e KEY, value from the invoking process's environment, so multi-line values survive), and a compose deploy ships ${KEY} references plus a 0600 secrets.env beside the compose file (docker compose --env-file, removed by compose down). Ant stores no secret values of its own: there is no secret store, no vault, and no encrypted field in the config or state files. Docker itself records a container's environment, so docker inspect still shows the values at runtime; that is the container engine's storage, not ant's.

Trust boundaries and gaps

  • The local dashboard trusts loopback. It performs real mutations (users, deploy, config). A non-loopback bind requires a shared password and should sit behind TLS; there is no per-user login or OIDC yet, so prefer a single-operator or otherwise trusted host.
  • The audit log is local and root-deletable. Privileged actions are hash-chained and ant nest audit detects an edited or removed line, but the log lives on the machine and is not shipped off-host; an operator with root can delete it. See Known issues.
  • Owner election is a modeled quorum. When the roster has zero owners and at least two admins, either admin may promote any member (including itself) to owner through the daemon; the second admin's consent is asserted by the caller, not verified, and the promotion is recorded in the audit log. Treat an admin account as equivalent to a potential owner.
  • No trust-on-first-use for a machine's NodeID beyond the value entered; invite tokens pin the machine NodeID, direct registration does not.
  • Supply chain: the iroh transport uses an unofficial, vendored Go binding (github.com/theinventorylib/iroh-go) with a static library; cosign verification exists but is opt-in (deploy.image_signing), and both the local CLI and the worker verify the signature before running when it is set.
Copyright © 2026