Skip to content
DBDeependra Bhatta~/notes
Docker#docker · #networking · #virtualbox · #vagrant · #docker-swarm

Docker Swarm: Building a Cluster with Vagrant

Build a Docker Swarm cluster on Vagrant VMs, then initialise it, join nodes, run and scale services, and use overlay networks and volumes in swarm mode.

· updated · 10 min read
ON THIS PAGE

A single Docker host is a single point of failure: when it goes down, every container on it goes down too. This guide builds a Docker Swarm cluster on Vagrant VMs, joins the nodes, and runs, scales and networks services across them.

By the end, you can stand up a lab cluster and explain what manager and worker nodes do.

Prerequisites

  • Vagrant and VirtualBox installed on the host machine
  • Enough free memory for the VMs (each VM in the Vagrantfile gets 1 GB)
  • Working knowledge of the Docker CLI from part 2

Container orchestration

Container orchestration means automating where containers run, how many copies run, how they find each other and what happens when a host fails. You describe the desired state ("three copies of nginx on port 8080") and the orchestrator keeps the cluster in that state.

The two most common orchestrators are Docker Swarm and Kubernetes. Kubernetes is the most widely used and is supported by every major cloud, but it takes more work to set up (see Kubernetes architecture). Swarm is built into Docker Engine (since Docker 1.12), uses the same CLI and is quick to start with, but it has fewer advanced features such as autoscaling. That makes Swarm a practical first orchestrator.

Sketch of a YAML file sent through the Docker CLI to a swarm of managers that spread containers across Docker nodes

Swarm architecture

A swarm is a group of Docker hosts, called nodes. Each node is a manager or a worker, and every swarm needs at least one manager.

Manager nodes

  • Accept commands from the CLI or API, such as a service definition.
  • Store the cluster state and keep it consistent between managers using the Raft consensus algorithm. The state lives in a built-in store on each manager; there is no separate etcd to install.
  • Split each service into tasks (one task = one container) and assign them to nodes.
  • Watch the cluster and reschedule tasks if a node fails or a new node joins.
  • Run workloads too by default. You can drain a manager so it only does management work.

Worker nodes

  • Receive tasks from the managers and run the containers.
  • Report task status back to the managers.

For example, if a service requests 10 replicas of nginx and the swarm has 3 worker nodes, the manager spreads the 10 tasks across them as evenly as it can.

Key Swarm features

FeatureWhat it means in practice
Built into Docker EngineNo extra software; docker swarm init turns it on
Declarative servicesYou state the desired number of replicas, ports and networks
Desired-state reconciliationIf a container or node dies, the manager starts replacement tasks
Scalingdocker service scale adds or removes tasks
Overlay networkingServices on different hosts share a virtual network
Service discoveryEvery service gets a DNS name inside the swarm
Load balancingA published port is reachable on every node (the routing mesh) and spread across the tasks
Secure by defaultNodes use mutual TLS to authenticate and encrypt their traffic
Rolling updatesUpdate tasks in batches with a delay, and roll back if something breaks

The official Swarm mode overview covers each feature in detail.

Building the lab with Vagrant

Network requirements

Each VM needs Docker Engine, a fixed IP that the others can reach, and these ports open between the nodes:

PortProtocolUsed for
2377TCPCluster management, nodes talking to managers
7946TCP and UDPCommunication among nodes
4789UDPOverlay network traffic (VXLAN)

The Vagrantfile

This Vagrantfile uses VirtualBox to create two VMs, one manager and one worker. It installs Docker on each VM and forwards three ports from the first manager to the host machine:

RUBVagrantfile
$install_docker_script = <<SCRIPT
sudo apt-get update
curl -sSL https://get.docker.com/ | sh
sudo usermod -aG docker vagrant
SCRIPT
 
BOX_NAME = "generic/ubuntu2204"
MEMORY = "1024"
MANAGERS = 1
MANAGER_IP = "192.168.56.1"
WORKERS = 1
WORKER_IP = "192.168.56.10"
CPUS = 1
VAGRANTFILE_API_VERSION = "2"
 
Vagrant.configure(VAGRANTFILE_API_VERSION) do |config|
    #Common setup
    config.vm.box = BOX_NAME
    config.vm.synced_folder ".", "/vagrant"
    config.vm.provision "shell",inline: $install_docker_script, privileged: true
    config.vm.provider "virtualbox" do |vb|
      vb.memory = MEMORY
      vb.cpus = CPUS
    end
    #Setup Manager Nodes
    (1..MANAGERS).each do |i|
        config.vm.define "manager0#{i}" do |manager|
          manager.vm.network :private_network, ip: "#{MANAGER_IP}#{i}"
          manager.vm.hostname = "manager0#{i}"
          if i == 1
            #Only configure port to host for Manager01
            manager.vm.network :forwarded_port, guest: 8080, host: 8080
            manager.vm.network :forwarded_port, guest: 5001, host: 5001
            manager.vm.network :forwarded_port, guest: 9000, host: 9000
          end
        end
    end
    #Setup Worker Nodes
    (1..WORKERS).each do |i|
        config.vm.define "worker0#{i}" do |worker|
            worker.vm.network :private_network, ip: "#{WORKER_IP}#{i}"
            worker.vm.hostname = "worker0#{i}"
        end
    end
end

How the Vagrantfile works:

  • The IP is built by appending the node number to a prefix, so manager01 gets 192.168.56.11 and worker01 gets 192.168.56.101.
  • As written, the file creates two nodes: manager01 and worker01 (MANAGERS = 1, WORKERS = 1). Set WORKERS = 2 to get a three-node cluster with one manager and two workers (worker02 gets 192.168.56.102).
  • The provisioning script uses Docker's convenience script, which is acceptable for a lab but not for production (see part 2).

Start the VMs and connect to the manager:

terminal
$ vagrant up
$ vagrant ssh manager01

Initialise the swarm and join nodes

Check whether swarm mode is on. A fresh node shows Swarm: inactive:

terminal
$ docker info | grep -i swarm

On the manager, a plain docker swarm init fails on Vagrant VMs. Each VM has two network interfaces (the Vagrant NAT one and the private network), and Docker cannot choose which address to advertise. The error itself says to use --advertise-addr, so give it the private IP:

terminal
$ docker swarm init --advertise-addr 192.168.56.11

The output includes a ready-made join command. To print the worker or manager join command again later, run these on the manager:

terminal
$ docker swarm join-token worker
$ docker swarm join-token manager

On each worker, run the join command it prints:

terminal
$ docker swarm join --token <worker-token> 192.168.56.11:2377

Managing nodes

CommandWhat it does
docker node lsList nodes (run on a manager)
docker node inspect manager01 --prettyReadable details of a node
docker node promote worker01Turn a worker into a manager
docker node demote manager02Turn a manager into a worker
docker node update --availability drain worker01Stop scheduling tasks on a node and move its running tasks elsewhere
docker swarm leaveRun on a node to make it leave the swarm (--force on the last manager)
docker node rm -f worker01Remove a node from the node list (on a manager)

Reading the docker node ls output:

  • An asterisk next to the ID marks the node you are running the command on.
  • AVAILABILITY controls scheduling:
    • Active: the scheduler can assign tasks to the node.
    • Pause: no new tasks, but existing tasks keep running.
    • Drain: no new tasks, and existing tasks are shut down and rescheduled on other nodes.
  • MANAGER STATUS shows the node's role:
    • Leader: the manager currently making decisions for the swarm.
    • Reachable: another manager that takes part in Raft and can become leader.
    • Unavailable: a manager that other managers cannot reach.
    • Blank: a worker.

Running services

A service is the swarm version of docker run: you describe the container and how many replicas you want.

terminal
$ docker service create -d --name nginx_service -p 8080:80 --replicas 2 nginx:latest
$ docker service ls
$ docker service ps nginx_service
$ curl http://192.168.56.11:8080

Because of the routing mesh, port 8080 answers on every node, even one that runs no nginx task. With the Vagrant port forward, http://localhost:8080 on the host machine also reaches the service.

CommandWhat it does
docker service lsList services and their replica counts
docker service ps nginx_serviceList the tasks of a service and the node each runs on
docker service inspect nginx_service --prettyReadable details of a service
docker service logs nginx_serviceLogs from all tasks of the service
docker service scale nginx_service=3Change the number of replicas
docker service update --image nginx:1.27 nginx_serviceRolling update to a new image (docker service update --help lists every option)
docker service rm nginx_serviceRemove the service and its tasks

Overlay networks in swarm mode

An overlay network spans all nodes, so service tasks on different hosts can talk to each other by service name.

terminal
$ docker network create -d overlay my_overlay
$ docker network create -d overlay --opt encrypted encrypted_overlay
$ docker service create -d --name nginx_overlay --network my_overlay -p 8081:80 --replicas 2 nginx:latest

--opt encrypted encrypts the traffic between nodes on that network. Encryption adds some performance overhead.

CommandWhat it does
docker network lsList networks (overlay networks show swarm scope)
docker network inspect encrypted_overlayShow the network's settings and connected tasks
docker service update --network-add my_overlay nginx_serviceAttach an existing service to a network
docker service update --network-rm my_overlay nginx_serviceDetach it
docker network rm my_overlayRemove the network

Volumes and plugins in swarm mode

Create a volume with a specific driver:

terminal
$ docker volume create -d local portainer_data
$ docker volume ls

Docker plugins add drivers for volumes, networks and logging:

CommandWhat it does
docker plugin install splunk/docker-logging-pluginInstall a plugin (this one is a logging driver)
docker plugin lsList installed plugins
docker plugin disable <id>Disable a plugin
docker plugin rm <id>Remove a plugin

Troubleshooting

  • could not choose an IP address to advertise on docker swarm init: the VM has more than one interface. Use --advertise-addr <private-ip>.
  • Worker cannot join: check that TCP 2377 on the manager is reachable from the worker, and that the worker uses the manager's private IP, not the NAT one.
  • Service stuck at 0/2 replicas: run docker service ps <name> --no-trunc to see why tasks fail, for example an image that cannot be pulled.
  • Data missing after a task moved: the task landed on another node with its own local volume. See the note above.

Key takeaways

  • Swarm is built into Docker Engine; docker swarm init and docker swarm join are the only commands needed to form a cluster.
  • Managers hold the cluster state with Raft and schedule tasks; workers run them.
  • On multi-interface VMs, always set --advertise-addr to the private IP.
  • Services declare replicas and ports, and the routing mesh makes a published port reachable on every node.
  • Overlay networks connect tasks across hosts, but local volumes do not follow tasks between nodes.

This is the final part of the Docker from scratch series. For a more capable orchestrator, continue with Kubernetes architecture.