Known issues
Tracked gaps that are deliberately deferred. Each entry says what is missing, why, the impact, the path forward, and where the code lives.
1. Remote run: stack deploys; stack builds are open
Pillar: Deploy · Status: routing, zero-downtime, and destroy implemented
What works now. A run: container project with a machine: builds on the
machine: the client packages the build context (build.run, a gzipped tar with
the context's .dockerignore applied), the worker unpacks it into a temporary
directory, runs the same builder as the local path, and starts the container
with volumes, resources, restart policy, and log rotation. A run: compose
project with a machine: resolves the compose file (docker compose config,
all deploy.files layers, with deploy.profiles), builds every service that
declares build: on the machine (each service's context is uploaded and built
individually, then rewritten to the produced image:), and ships the file to
the worker, which runs docker compose up -d. Image-only services pull as
before. deploy.build_services: false is refused for a file that still
declares build:, the machine has no source tree to run it from.
A run: stack project with a machine: sends the stack file (single file;
deploy.files layering is refused) to the worker, which refuses any service
that declares build: (push the image first, for example with
deploy.registry), stores the file under the machine's project directory, and
runs docker stack deploy --with-registry-auth. Domains are routed through the
machine's Caddy at each stack service's published host port (the domain names
the service with service: and its container port with port:). Every service
gets ant's managed/app labels, so tasks show up in ant nest containers list
and ant trail status, and ant trail logs <deployment> [service] reads a
task's logs (name the service when the stack has several). The machine must be a
Swarm manager, ant nest swarm init --machine M sets up a single-node one;
ant nest swarm status and ant nest swarm leave --force manage it. Swarm is a
server feature: ant never initializes a swarm on this host.
zero_downtime patches a Swarm rolling-update policy (start-first, rollback on
failure, convergence wait); services that publish ports in host mode fall back
to stop-first with a warning, since two tasks cannot bind the same host port on
a single node.
What's missing.
- Stack builds: services that declare
build:are refused; prebuild and push, or userun: container/run: compose. - Stack rollback: refused, like compose: a stack's release is a file that names several images, not one recorded image. Pin the previous image in the stack file and redeploy (Swarm rolls it out).
- Cluster and service operations: the common ones are available
(
nest swarm nodes,node promote|demote|drain|activate|pause|rm,join-token,services,service ps|logs|scale|restart|update|rm). Anything beyond them (overlay/network administration, swarm configs and secrets, and the rest ofdocker service update) stays docker commands on the machine.
ant trail destroy --machine M stops and removes the deployment's workload on
the machine (container, compose project, or stack) and clears its Caddy routes.
The machine's images are left in place, so a redeploy reuses the build cache;
the local command deletes ant-built images instead.
- Compose build features:
build.target,build.secrets,build.ssh, andbuild.additional_contextsare refused (not silently ignored); prebuild and reference the image instead.
Note. The resolved compose file travels inside one RPC request frame,
capped at 2 MiB (response frames are capped at 8 MiB); a larger compose file
fails. Build logs from build.run are capped at 1 MiB (the tail is returned),
and the worker bounds context extraction to 8 GiB and rejects unsafe archive
entries.
2. Registry credentials travel with the deploy
Pillar: Deploy · Status: implemented
What works now. With deploy.registry set, a source-built remote deploy
builds on this machine, tags the image <host>/<repository>/<app>:<release>,
and pushes it (docker login --password-stdin when ANT_REGISTRY_PASSWORD is
set); the machine then pulls the qualified reference on deploy.run. The same
password travels with the deploy over the encrypted transport and is
materialized on the machine as a short-lived DOCKER_CONFIG (0700 directory,
0600 file) for every command that can pull: deploy.run, compose.apply, and
stack.apply (where --with-registry-auth forwards it to the nodes). The
directory is removed when the operation ends; ant never stores the credential
in a config or on disk. For an explicit --image ghcr.io/... pull without
deploy.registry, the host is parsed from the reference and the username comes
from ANT_REGISTRY_USERNAME (both with ANT_REGISTRY_PASSWORD).
Limits. The credential crosses the wire on each deploy, the transport
encrypts the frame, and the value is handled like any other runtime secret.
Anyone who may deploy can therefore pull the configured registry's private
images; deployer/ci are already root-equivalent by design. The local image
built for the push is not pruned by the local pruner yet.
3. The dashboard is local-only by design (no multi-user auth)
Pillar: Collab · Status: by design
Each operator runs their own dashboard on their own laptop, against their own
~/.ant, and it is never hosted on a server. Machine access is governed by
each machine's colony roster and roles, so a user's dashboard can only act on
the machines their account can reach. There is no shared-login model to
support: an operator editing their own files is the trust boundary. The
non-loopback password (ant ui passwd) exists only as a safety net for
reaching your own dashboard from another device, not as a hosted mode.
4. The audit log is local to the machine
Pillar: Security · Status: implemented; off-host shipping deferred
Privileged actions (deploy, route, roster, volume, tunnel) are appended to a
hash-chained log, and ant nest audit reports whether the chain verifies, so an
edited or removed line is detected. The log rotates at 8 MiB (the prior segment
is kept as audit.log.1 and is independently verifiable); reading the log
requires the admin role. The log lives on the machine and is not shipped
off-host: an operator with root can still delete the whole file. Forwarding a
copy to a collector is future work.
5. No trust-on-first-use for a machine NodeID
Pillar: Security · Status: partial
Invite tokens pin the machine NodeID; direct registration does not. Pinning on first contact is future work.
6. No secret store
Pillar: Security · Status: accepted
Ant has no vault of its own: build-time secrets are read from the operator's
environment and passed to the builder for one build, and runtime secrets are
injected as container environment variables. Nothing is encrypted at rest by
ant, and no secret is ever written to ant.yaml, the lockfile, or the machine
state. A secret may name a provider (from: op://… for 1Password, a generic
cmd:… for Vault/AWS/pass, or env:…), resolved through the operator's CLI
at deploy time so the value still never lands in ant. deploy.secrets entries
with no environment value are reported as warnings, not errors, so a typo is
visible but not fatal.
If you need rotation, audit, or per-environment access control for secrets, keep them in your existing secret manager and export them into the deploy environment
7. Worker upgrades are a re-provision
Pillar: Ops · Status: by design for now
ant-worker has no self-update path: it runs unprivileged and cannot write
/usr/local/bin or restart its own unit. Upgrading a machine re-runs the
bootstrap over SSH, which preserves the machine identity and state.json
(roster, invites):
ant nest bootstrap --machine prod --host 192.0.2.10 --ssh-key ~/.ssh/id_ed25519
The client and worker carry an RPC protocol version and refuse to operate against a mismatched peer, so upgrade both sides together (CLI first is fine: an old worker answers the new client with a version error until it is re-provisioned).
8. Remote zero_downtime: containers and stacks yes, compose later
Pillar: Deploy · Status: implemented; remote compose still staged
A run: container deploy with zero_downtime and at least one domain stages
the new container beside the running one on an ephemeral host port,
health-gates it, moves the app's Caddy routes to that port, and only then
parks-and-removes the old container and renames the new one into service. The
published port survives the rename, so the routes keep serving throughout. The
swap is one deploy.run call (routes travel in the payload), so a failure
before the route move leaves the old container and its routes untouched.
A run: stack deploy with zero_downtime patches a Swarm rolling-update
policy: order: start-first, failure_action: rollback, and a convergence
wait, so Swarm starts new tasks before stopping old ones while the published
port (held by the routing mesh) keeps serving. Services that publish ports in
mode: host cannot overlap two tasks on a single node, so those get stop-first
with a warning; services without a healthcheck warn that rollback relies on
task state. Routing warns too when a routed service publishes in mode: host
on a multi-node swarm: that port exists only on the node running the task,
while Caddy runs on the manager.
The container path falls back to the staged swap (with a warning) when there is
no domain to move, no running predecessor, or a writable mount that must not be
attached to a second container. Remote compose still uses the staged swap;
docker compose up -d has no native rolling update, so it needs the same
ephemeral-port + Caddy-repoint pattern per service.
Tracked on other branches
- SSH over iroh: the TCP-forwarding primitive now exists (
ant nest tunnel, a raw iroh stream), so forwarding a machine'ssshdis the next step; provisioning still uses the host's standard SSH (with--cloud-initas the no-SSH fallback). - Container-to-container across machines:
ant nest tunnellets an operator reach a peer machine's published ports over iroh. A worker dialing another worker would need the caller's machine NodeID on the peer's roster, a machine-to-machine trust model that does not exist yet (the roster holds account NodeIDs).