Remote deploy
Point a project at a machine and deploy:
# ant.yaml
machine: prod
ant trail deploy prod
For a container project, the client packages the build context (a gzipped tar
with the context's .dockerignore applied), uploads it over the transport, and
the worker builds the image on the machine with the same builder the local
path uses, docker build, nixpacks, pack, or railpack. The image never
crosses the wire, and the build sees the machine's architecture and cache. The
container then starts there and any domains route through the machine's Caddy.
ant trail deploy --machine M --image REF remains the low-level form for running
a pre-built image; --from-tar transfers a docker save archive, --pull
fetches from a registry.
Where the context fits
internal/docker/context is the seam that answers "where do build and run
execute?". Opening a context for the target and calling Deploy hides the
local/remote split:
- Local context (
dockerctx.Local()): build and run shell out to the local docker CLI, exactly as they always have. - Remote context (
dockerctx.Open(machine)):Buildpackages the build context and calls the worker'sbuild.run;Deploytranslates the resolved request intodeploy.run,compose.apply,stack.apply, androute.apply, and surfaces settings the target cannot apply as warnings. - Pre-built images: a
deploy.build: nonerequest skipsBuildentirely: the remote context runs the image and the machine pulls it (or receives it viaant trail deploy --image --from-tar). - Registry: a registry is not a context; it is the distribution channel
between contexts. With
deploy.registryset, a source-built remote deploy builds on the local context, tags and pushes the image (<host>/<repository>/<app>:<release>), then runs on the remote context with the machine pulling it. The password comes fromANT_REGISTRY_PASSWORDand travels with the deploy as a short-lived worker-side credential, so a private image pulls without its own login on the machine; without the environment variable (or for a machine that already logged in) nothing changes.
Two deploy shapes
run: | What happens |
|---|---|
container | The context is uploaded, built on the machine, and run. The common case. |
compose | The client resolves the compose file (docker compose config), builds every service that declares build: on the machine (rewriting it to the produced image), and ships the file; the worker runs docker compose up -d. Image-only services pull as before. |
stack | The client reads the stack file and the worker runs docker stack deploy --with-registry-auth on a Swarm manager, then routes domains at each service's published port. Tasks carry ant's managed/app labels, so nest containers list, trail status, and trail logs <deployment> [service] see them. Services need pullable images (build here and push with deploy.registry, or reference a registry image); build: services and deploy.files layering are refused, zero_downtime maps to a Swarm start-first rolling update, and rollback is refused (pin the previous image and redeploy). |
Build context packaging
- The context is the resolved
build.contextdirectory; the.dockerignorethere is applied before upload, sonode_modulesand friends do not travel. - The Dockerfile may live inside the context (sent by path, always included even when ignored) or outside it (read and shipped as content).
- The machine needs the chosen builder installed (
ant nest tools --machine Mchecks). Adockerfilebuild needs only Docker. - Contexts are unpacked into a temporary directory on the worker and removed
after the build. Absolute paths,
..traversal, hard links, and symlinks that leave the context are rejected.
Compose
Every service is handled: one that references an image pulls it on the
machine; one that declares build: has its context uploaded and built on
the machine, then the resolved file is rewritten to the produced image. The
compose file's own log rotation, resource limits, env, and restart policy are
preserved; ant's defaults are added where the file is silent.
services:
api:
build: . # built on the machine
ports: ["8080:8080"]
redis:
image: redis:7 # pulled on the machine
Remote compose build features that ant cannot reproduce yet are refused with a
clear message rather than silently ignored: build.target, build.secrets,
build.ssh, and build.additional_contexts.
Reliability on a server
- Staged replacement. When a container with the same name is already running, the new image, env, ports, and entrypoint are validated in a staging container first (on ephemeral host ports, with the deploy's read-only mounts). A bad image or crash loop leaves the running container serving. The staged container is removed and the canonical one replaced only after it is healthy. A deploy with a writable mount (a live data volume, or a bind the app needs to boot) cannot be staged safely and is replaced in place instead, then checked for exiting immediately.
- Zero-downtime rollouts. For a container deploy with at least one domain,
zero_downtimestarts the new container beside the running one on an ephemeral port, health-gates it, moves Caddy to that port, and only then replaces the old container, so no request hits a stopped upstream. It needs a read-only (or no) mount set; a writable mount falls back to the staged swap. A stack deploy instead gets a Swarm start-first rolling update (stop-first with a warning when services publish host-mode ports). Remote compose still uses the staged swap. - Serialized deploys. The worker runs one container/compose mutation at a time, so two deploys of the same app cannot fight over the staging name and the swap.
- Health gate. A configured
health_checkis probed over HTTP against the staging container before the swap. - Rollback.
ant trail rollback <deployment> [--to RELEASE] [--steps N]redeploys a previously released image without rebuilding (compose projects are refused). The image tag must still exist on the machine. - Restart policy. Remote deploys default to
restart: unless-stopped(ant trail deploy --image --restartoverrides it), and compose services that do not declare a policy get the same, so a reboot brings the app back. - Smart defaults. Log rotation (10m × 3) and resource limits apply on the machine even when the project does not set them.
- Mount guard. A deploy that reaches the worker's identity/state directory,
the docker socket, or the host root is refused, on the client and the worker.
For compose this includes
secrets:/configs:file:sources, andprivileged: trueand hostdevices:are rejected outright. - Loopback ports. A published port binds
127.0.0.1by default (deploy.publish: loopback), so a deployed service is reachable through the machine's Caddy and on the machine itself but not from the network. Setdeploy.publish: allfor a source deploy (orant trail deploy --image --publish-allfor the image path) when a service must be reached directly and brings its own TLS. - Image retention. After a source build, the machine's superseded image tags
are pruned to
deploy.keep_images(falling back torollback.keep_releases, then 10). - Logs.
ant trail logs <deployment>fetches the tail from the machine (--tail N,all); streaming (-f) is local-only for now. - Inspecting.
ant nest containers list --machine Mandant nest volume list --machine Mshow what is running and what is stored on the machine.
What is not supported yet
- Compose build extras:
build.target,build.secrets,build.ssh, andbuild.additional_contextsare refused; prebuild the image and reference it. run: stack:build:services androllbackare refused; the machine must be a Swarm manager (ant nest swarm init --machine M). Routing, zero-downtime rolling updates, andant trail destroyare implemented. See Known issues.
All of this is tracked in Known issues.
Running deploys from a CI pipeline? See CI: a ci-role account,
a credential-bundle secret, and ANT_MACHINE are all a job needs.
Workaround
For a compose build feature ant cannot reproduce yet (build.target,
build.secrets), prebuild the image and reference it with image:, or use a
run: container project.
Routes
The worker applies route changes (route.apply) so the deploy is reachable once
it lands, through the machine's Caddy, which is required on a remote nest.
If Caddy is not reachable the deploy is refused before anything starts. Route and
compose operations are authorized against the machine's roster like every other
gated RPC; see Security model.
Caddy and Swarm
For run: stack, Caddy stays what it is everywhere else: a host process on the
machine that runs the worker; for stacks, a Swarm manager. Routes target
each service's published host port on that machine, resolved at deploy time:
- Ingress mode (default) works at any node count: the routing mesh serves a
published port on every node, so
127.0.0.1:<port>on the manager reaches the task wherever it runs. Point the domains' DNS at the manager. - Host mode (
mode: host) binds the port only on the node running the task. On a single-node swarm that is still the manager; on a multi-node swarm the route can point at nothing, ant warns during the deploy when a routed service does this. Prefer ingress mode for routed services, or run Caddy on the node(s) hosting the task. - Worker nodes need only Docker.
ant nest swarm init --machine Mmakes the machine a manager; other nodes join with the token fromant nest swarm join-token --machine M.ant-workerand Caddy live on the manager.
The --handover bundle's Caddyfile snippet carries the same upstreams. If you
re-home Caddy onto the app's overlay network (as a Swarm service), switch the
upstreams to tasks.<project>_<service>:<container-port>, task DNS names
resolve inside the overlay.