Security model
How Ant authenticates, authorizes, and limits what a peer can do.
Identity
- Every principal holds an ed25519 keypair. The public key is the NodeID, in
base32 (no padding). Account keys live at
~/.ant/identity; a machine's key lives on the machine (~/.ant/worker/agent-identity, or wherever the worker's--identitypoints). - The transport binds an endpoint with the identity's seed and verifies the bound endpoint id equals the NodeID, so a peer cannot present a key it does not hold.
Transport
- iroh = QUIC + TLS 1.3 with mutual authentication: both sides prove possession of their key. Dialing is by NodeID; no open ports, no host certificates.
- A self-hosted relay, when configured, only forwards encrypted bytes.
- Private keys are
0600. The account's portable credential bundle is passphrase-encrypted by default (PBKDF2-HMAC-SHA256, 210k iterations, AES-256-GCM), and load re-checks that the key matches the profile NodeID.
Authorization: the daemon is the authority
Mutual TLS authenticates the channel; it does not authorize the caller. The machine daemon resolves the caller's NodeID against its roster and enforces a method → capability table:
| Method | Requires |
|---|---|
ping, whoami | public |
invites.redeem | public, gated by a signed invite token |
users.list, containers.list, containers.logs, images.list, volumes.list, metrics.get, caddy.status, tools.probe, swarm.status, swarm.nodes, swarm.services, swarm.service.tasks, swarm.service.logs | read (viewer+) |
users.add, users.setRole, users.remove, users.update, invites.revoke, state.export, audit.list, swarm.init, swarm.leave, swarm.node.update, swarm.node.rm, swarm.joinToken | colony manage, and the caller must be allowed to grant/remove/modify the target's role |
state.import | ownership (owner only) |
deploy.run, deploy.rollback, deploy.commit, image.load, build.run, compose.apply, compose.down, stack.apply, stack.down, route.apply, containers.action, containers.remove, volumes.remove, swarm.service.scale, swarm.service.restart, swarm.service.update, swarm.service.rm, handover.export, tunnel.open | deploy (deployer, ci, admin, owner) |
Role capabilities: owner (all) · admin (all but ownership transfer) ·
deployer (read, deploy, osuser) · viewer (read) · ci (read, deploy).
Ownership transfer (state.import) is owner-only.
Fail closed. With no roster and no boot owner, every gated method is refused. The first owner is seeded out of band:
ant-worker --owner <nodeid>(repeatable), the bootstrap/provisioning owner, emitted byant nest bootstrap; or- writing the owner into the machine's
state.jsonas root.
Invites
- An invite is signed by the inviter's account key and carries the destination machine NodeID, the role, and an expiry.
- Redemption is server-side (
invites.redeem): the daemon verifies the signature, expiry, that the invite is for this machine, the inviter's current role, and the grant matrix, and that the nonce is unused. - Single-use is enforced by the daemon (
state.jsonkeeps consumed nonces until their expiry). The invitee need not be on the roster yet. - Revocable:
invites.revoke(colony manage) records a nonce as revoked until its expiry, and redemption rejects it, so a leaked token can be killed.
Least privilege on the machine
- The worker runs unprivileged and never invokes sudo. Privileged work on a
machine is operator-run and executes exact argv, never a shell:
ant nest reconcile --applyruns as root, and the SSH provisioning path (ant nest machines add --host,ant nest bootstrap --host) runs its privileged steps on the target through the SSH user'ssudo. - Docker calls use
exec.Commandwith argument slices, notsh -c. The image is passed after--, so a crafted reference cannot be read as a flag. - Server-side validation: container names, compose project names, and
network names must match
^[a-zA-Z0-9][a-zA-Z0-9_.-]{0,62}$; image refs may not start with-or contain whitespace. This blocks path traversal (../…) and flag injection. - Container actions address the container by exact name only.
The deploy capability is root-equivalent on the machine
This is the most important thing to understand before inviting someone:
deployerandcican run any image, with arbitrary environment and arbitrary-vmounts. A deployer can mount the worker's own state directory (~ant/.ant) into a container and read the machine's private identity, or mount the host filesystem. IngroupDocker mode the worker is in thedockergroup, which is host root; inrootlessmode it is full control of the worker's unix user (including its NodeID key).- Treat
deployer/cias machine-level trust, not as a read-mostly role. Give the admin/owner roles to people you would trust with SSH root on that machine, and preferviewerfor read-only access. run: composeservice images can also declare arbitrary mounts and privileges; the compose file comes from a project config the deployer controls.
The role matrix (read, deploy, colony.manage, ownership.transfer,
osuser) is enforced server-side, and the worker never runs privileged host
commands, but container execution itself is the escape hatch, as with any
Docker-based deploy tool.
Routes and ports
- A published port binds loopback by default, so a deployed app is reachable
through the machine's reverse proxy (and on the machine) but not from the
network.
deploy.publish: allopts into binding every interface. route.applyaccepts only valid hostnames (no wildcards, ports, schemes, or paths), only loopback upstreams (the machine's own containers), and rejects a host (or host+path) already routed by another app, so a deployer cannot shadow or hijack a teammate's domain.- The worker's Caddy admin API is bound to
127.0.0.1:2019with an origin allowlist (origins 127.0.0.1:2019 localhost:2019), so a browser request carrying a foreignOrigin(a page on another domain that resolves to loopback, i.e. DNS rebinding) cannot reprogram routing, and a request with a foreignHostis refused. On a multi-user host, prefer the distribution'scaddy-only unix socket instead. - The dashboard is operator-local by design: one dashboard per laptop,
reading that operator's own files, never hosted on a server. It binds
loopback and has no authentication there (the trust boundary is the local
machine). Binding it to a non-loopback address is an escape hatch for
reaching your own dashboard from another device: it is refused unless a
password is set (
ant ui passwd), and is then wrapped in HTTP Basic auth and should sit behind TLS. A successful login issues a signed, HttpOnly session cookie whose key derives from the password hash, so the expensive hash runs at login rather than on every request, and changing the password invalidates existing sessions. - On a loopback bind the dashboard also applies a Host allow-list (rejecting
a request whose
Hostis not loopback, DNS rebinding) and a same-site guard that rejects state-changing requests a browser marks as cross-site. JSON request bodies are capped at 8 MiB. - Deploys and compose files that reach the worker's state/identity directory,
the docker socket, or the host root are refused, on the worker and on the
client (so a local run and a remote run behave the same). The check is
boundary-aware and covers ancestors of the state directory (e.g.
/home/antexposes/home/ant/.ant/worker/agent-identity). A project-relative mount such as.or./datais allowed in a compose file (it is relative to the compose/project directory, not the worker's home) while arun: containerdeploy refuses relative bind sources outright, because the worker has no project directory to resolve them against. For compose this coversvolumes, top-levelsecrets:/configs:file:sources, and it also rejectsprivileged: trueand hostdevices:. A resolved compose document that still containsinclude:is refused on the worker: compose expands it atuptime, after the guards, so its services could bypass every check. - Untrusted build contexts are extracted with an 8 GiB cap and reject
absolute paths,
..traversal, hard links, and symlinks that leave the context. The transport caps a request frame at 2 MiB (a general frame at 8 MiB), a blob at 64 GiB, and bounds connections and streams per connection. - Every privileged RPC is written to a tamper-evident audit log
(
ant nest audit): each entry chains to the previous by hash and each rotated segment links to the one before it, so an edited entry or a deleted segment fails verification. (The newest lines can still be truncated by anyone who can write the state directory, and deleting the oldest retained segment is indistinguishable from normal rotation.) The log rotates at 8 MiB, and reading it is an admin action that does not itself append. - Secret values are never stored: a
from:provider (1Passwordop://, a genericcmd:, orenv:) is resolved at deploy time and only its warning is surfaced on failure. - Tunnels (
ant nest tunnel, or the dashboard's Tunnel action) forward bytes to a published loopback port of a container ant manages and nothing else, so they add no reach beyond the deploy capability: a peer cannot use one to reach the docker socket, the Caddy admin API, or another local listener. See Tunnels.
Protocol version
Every RPC carries an API version, and both sides reject a mismatch with an
actionable message (peer speaks protocol vN, this worker speaks v1; upgrade the older side) instead of failing as an unknown method. A mixed-version fleet
therefore fails closed on the first call, which is why upgrading worker and CLI
together matters (see the upgrade note in Known issues). Before the first
tagged release the protocol is not stable and the version stays 1; it is
bumped only once there are deployed workers that cannot simply be rebuilt.
At rest
~/.ant/config.json, the machinestate.json, and identity files are0600; the machine state is written atomically.- Build-time secrets are resolved from the environment and never baked into
images or the lockfile. Runtime secrets never appear in command argv or a
persisted compose document: a container deploy (local or remote) passes them
through the docker process environment (a bare
docker run -e KEY, value from the invoking process's environment, so multi-line values survive), and a compose deploy ships${KEY}references plus a0600secrets.envbeside the compose file (docker compose --env-file, removed bycompose down). Ant stores no secret values of its own: there is no secret store, no vault, and no encrypted field in the config or state files. Docker itself records a container's environment, sodocker inspectstill shows the values at runtime; that is the container engine's storage, not ant's.
Trust boundaries and gaps
- The local dashboard trusts loopback. It performs real mutations (users, deploy, config). A non-loopback bind requires a shared password and should sit behind TLS; there is no per-user login or OIDC yet, so prefer a single-operator or otherwise trusted host.
- The audit log is local and root-deletable. Privileged actions are
hash-chained and
ant nest auditdetects an edited or removed line, but the log lives on the machine and is not shipped off-host; an operator with root can delete it. See Known issues. - Owner election is a modeled quorum. When the roster has zero owners and at least two admins, either admin may promote any member (including itself) to owner through the daemon; the second admin's consent is asserted by the caller, not verified, and the promotion is recorded in the audit log. Treat an admin account as equivalent to a potential owner.
- No trust-on-first-use for a machine's NodeID beyond the value entered; invite tokens pin the machine NodeID, direct registration does not.
- Supply chain: the iroh transport uses an unofficial, vendored Go binding
(
github.com/theinventorylib/iroh-go) with a static library; cosign verification exists but is opt-in (deploy.image_signing), and both the local CLI and the worker verify the signature before running when it is set.