Remote

Remote deploy

Build and run a project on a registered machine.

Point a project at a machine and deploy:

# ant.yaml
machine: prod
ant trail deploy prod

For a container project, the client packages the build context (a gzipped tar with the context's .dockerignore applied), uploads it over the transport, and the worker builds the image on the machine with the same builder the local path uses, docker build, nixpacks, pack, or railpack. The image never crosses the wire, and the build sees the machine's architecture and cache. The container then starts there and any domains route through the machine's Caddy.

ant trail deploy --machine M --image REF remains the low-level form for running a pre-built image; --from-tar transfers a docker save archive, --pull fetches from a registry.

Where the context fits

internal/docker/context is the seam that answers "where do build and run execute?". Opening a context for the target and calling Deploy hides the local/remote split:

  • Local context (dockerctx.Local()): build and run shell out to the local docker CLI, exactly as they always have.
  • Remote context (dockerctx.Open(machine)): Build packages the build context and calls the worker's build.run; Deploy translates the resolved request into deploy.run, compose.apply, stack.apply, and route.apply, and surfaces settings the target cannot apply as warnings.
  • Pre-built images: a deploy.build: none request skips Build entirely: the remote context runs the image and the machine pulls it (or receives it via ant trail deploy --image --from-tar).
  • Registry: a registry is not a context; it is the distribution channel between contexts. With deploy.registry set, a source-built remote deploy builds on the local context, tags and pushes the image (<host>/<repository>/<app>:<release>), then runs on the remote context with the machine pulling it. The password comes from ANT_REGISTRY_PASSWORD and travels with the deploy as a short-lived worker-side credential, so a private image pulls without its own login on the machine; without the environment variable (or for a machine that already logged in) nothing changes.

Two deploy shapes

run:What happens
containerThe context is uploaded, built on the machine, and run. The common case.
composeThe client resolves the compose file (docker compose config), builds every service that declares build: on the machine (rewriting it to the produced image), and ships the file; the worker runs docker compose up -d. Image-only services pull as before.
stackThe client reads the stack file and the worker runs docker stack deploy --with-registry-auth on a Swarm manager, then routes domains at each service's published port. Tasks carry ant's managed/app labels, so nest containers list, trail status, and trail logs <deployment> [service] see them. Services need pullable images (build here and push with deploy.registry, or reference a registry image); build: services and deploy.files layering are refused, zero_downtime maps to a Swarm start-first rolling update, and rollback is refused (pin the previous image and redeploy).

Build context packaging

  • The context is the resolved build.context directory; the .dockerignore there is applied before upload, so node_modules and friends do not travel.
  • The Dockerfile may live inside the context (sent by path, always included even when ignored) or outside it (read and shipped as content).
  • The machine needs the chosen builder installed (ant nest tools --machine M checks). A dockerfile build needs only Docker.
  • Contexts are unpacked into a temporary directory on the worker and removed after the build. Absolute paths, .. traversal, hard links, and symlinks that leave the context are rejected.

Compose

Every service is handled: one that references an image pulls it on the machine; one that declares build: has its context uploaded and built on the machine, then the resolved file is rewritten to the produced image. The compose file's own log rotation, resource limits, env, and restart policy are preserved; ant's defaults are added where the file is silent.

services:
  api:
    build: .                        # built on the machine
    ports: ["8080:8080"]
  redis:
    image: redis:7                  # pulled on the machine

Remote compose build features that ant cannot reproduce yet are refused with a clear message rather than silently ignored: build.target, build.secrets, build.ssh, and build.additional_contexts.

Reliability on a server

  • Staged replacement. When a container with the same name is already running, the new image, env, ports, and entrypoint are validated in a staging container first (on ephemeral host ports, with the deploy's read-only mounts). A bad image or crash loop leaves the running container serving. The staged container is removed and the canonical one replaced only after it is healthy. A deploy with a writable mount (a live data volume, or a bind the app needs to boot) cannot be staged safely and is replaced in place instead, then checked for exiting immediately.
  • Zero-downtime rollouts. For a container deploy with at least one domain, zero_downtime starts the new container beside the running one on an ephemeral port, health-gates it, moves Caddy to that port, and only then replaces the old container, so no request hits a stopped upstream. It needs a read-only (or no) mount set; a writable mount falls back to the staged swap. A stack deploy instead gets a Swarm start-first rolling update (stop-first with a warning when services publish host-mode ports). Remote compose still uses the staged swap.
  • Serialized deploys. The worker runs one container/compose mutation at a time, so two deploys of the same app cannot fight over the staging name and the swap.
  • Health gate. A configured health_check is probed over HTTP against the staging container before the swap.
  • Rollback. ant trail rollback <deployment> [--to RELEASE] [--steps N] redeploys a previously released image without rebuilding (compose projects are refused). The image tag must still exist on the machine.
  • Restart policy. Remote deploys default to restart: unless-stopped (ant trail deploy --image --restart overrides it), and compose services that do not declare a policy get the same, so a reboot brings the app back.
  • Smart defaults. Log rotation (10m × 3) and resource limits apply on the machine even when the project does not set them.
  • Mount guard. A deploy that reaches the worker's identity/state directory, the docker socket, or the host root is refused, on the client and the worker. For compose this includes secrets:/configs: file: sources, and privileged: true and host devices: are rejected outright.
  • Loopback ports. A published port binds 127.0.0.1 by default (deploy.publish: loopback), so a deployed service is reachable through the machine's Caddy and on the machine itself but not from the network. Set deploy.publish: all for a source deploy (or ant trail deploy --image --publish-all for the image path) when a service must be reached directly and brings its own TLS.
  • Image retention. After a source build, the machine's superseded image tags are pruned to deploy.keep_images (falling back to rollback.keep_releases, then 10).
  • Logs. ant trail logs <deployment> fetches the tail from the machine (--tail N, all); streaming (-f) is local-only for now.
  • Inspecting. ant nest containers list --machine M and ant nest volume list --machine M show what is running and what is stored on the machine.

What is not supported yet

  • Compose build extras: build.target, build.secrets, build.ssh, and build.additional_contexts are refused; prebuild the image and reference it.
  • run: stack: build: services and rollback are refused; the machine must be a Swarm manager (ant nest swarm init --machine M). Routing, zero-downtime rolling updates, and ant trail destroy are implemented. See Known issues.

All of this is tracked in Known issues.

Running deploys from a CI pipeline? See CI: a ci-role account, a credential-bundle secret, and ANT_MACHINE are all a job needs.

Workaround

For a compose build feature ant cannot reproduce yet (build.target, build.secrets), prebuild the image and reference it with image:, or use a run: container project.

Routes

The worker applies route changes (route.apply) so the deploy is reachable once it lands, through the machine's Caddy, which is required on a remote nest. If Caddy is not reachable the deploy is refused before anything starts. Route and compose operations are authorized against the machine's roster like every other gated RPC; see Security model.

Caddy and Swarm

For run: stack, Caddy stays what it is everywhere else: a host process on the machine that runs the worker; for stacks, a Swarm manager. Routes target each service's published host port on that machine, resolved at deploy time:

  • Ingress mode (default) works at any node count: the routing mesh serves a published port on every node, so 127.0.0.1:<port> on the manager reaches the task wherever it runs. Point the domains' DNS at the manager.
  • Host mode (mode: host) binds the port only on the node running the task. On a single-node swarm that is still the manager; on a multi-node swarm the route can point at nothing, ant warns during the deploy when a routed service does this. Prefer ingress mode for routed services, or run Caddy on the node(s) hosting the task.
  • Worker nodes need only Docker. ant nest swarm init --machine M makes the machine a manager; other nodes join with the token from ant nest swarm join-token --machine M. ant-worker and Caddy live on the manager.

The --handover bundle's Caddyfile snippet carries the same upstreams. If you re-home Caddy onto the app's overlay network (as a Swarm service), switch the upstreams to tasks.<project>_<service>:<container-port>, task DNS names resolve inside the overlay.

Copyright © 2026