One network per operator
Instances accumulate. A ground station gets one instance for its radios to the aircraft, another for a phone that joined later, and suddenly two radios appear as four links, every probe runs twice, and the dashboard shows two topologies of one machine. When every node belongs to the same operator, fold them into one network. Nothing in the daemon changes — this is configuration, and it is how Atlas is meant to run. When the networks belong to other people, you want an interconnect instead.
1 — Choose the instance that survives
Keep the instance that is hardest to change, and move everything else into it:
- If phones are members, keep the phones' network. The mobile app's address space is fixed, so the network the phones already use must be the survivor.
- Otherwise keep the one with the most nodes, or the one other systems already address — the fewer things to re-point, the better.
- Keep its keys and addresses unchanged. Members that already trust the survivor's key need no change at all.
2 — Find everything that names the retiring instance
The daemon is the easy part; the dependants are what bite. Search each host for the retiring instance's tunnel addresses, interface name, config path and unit name:
on each hostgrep -rn -e '10.0.101.' -e 'atlas1' /etc /opt /usr/local/bin /etc/systemd 2>/dev/null
ls /tmp/atlas/ # snapshot files are named after the interface
Typical finds: applications bound to the old tunnel address, services that read the old interface's snapshot file, a second dashboard port, firewall rules, and scripts that restart the old unit by name.
3 — Give the survivor every link and every peer
Add the retiring instance's links to the survivor's config — each physical link once — and its peers, with their per-link endpoints. Where both instances used the same radio, the survivor already has that link: keep one session per physical link. Then validate before anything restarts:
atlasd config check -c /etc/atlas/survivor.toml
4 — Park the retiring instance; do not delete it
Stop and mask its unit, keep its config untouched, and leave a note beside the config that says how to bring it back. A parked instance costs nothing and turns a bad night into a two-minute rollback.
systemctl disable --now retiring.service
mv /etc/systemd/system/retiring.service /etc/systemd/system/retiring.service.parked
systemctl mask retiring.service
cat > /etc/atlas/RETIRING-PARKED.txt <<'EOF'
Parked on <date>: folded into the survivor instance.
Restore: systemctl unmask retiring.service
mv /etc/systemd/system/retiring.service.parked /etc/systemd/system/retiring.service
systemctl daemon-reload && systemctl enable --now retiring.service
EOF
5 — Re-point the dependants
Move each finding from step 2 to the survivor: applications to its tunnel addresses, readers to its snapshot file, dashboards to its port. Restart the survivor on every host in the same pass.
6 — Prove it
- One topology. The dashboard shows one panel for the network, and the number of links between two nodes equals the number of physical links between them.
- Failover still works. Kill each link in turn; traffic must keep flowing on the others with the scheduling strategy you chose.
- Every application still works at its new addresses — video, control, telemetry.
- A cold boot comes up clean. Power-cycle every host together and repeat the checks without touching anything: the parked unit stays parked, the survivor starts, the network forms.