Knowledge base
Documentation
How the MagMon fleet is built and kept running, in the order you actually do it: take a Raspberry Pi from the sealed box to an asset reporting on the dashboard, add environmental sensors and a UPS if the unit gets them, wire the site up for remote access, then keep it healthy. Read it top to bottom the first time; after that, jump to the part you need.
Part 1
Stand up a gateway
A gateway is one Raspberry Pi (or one shared server) running one collector per asset. These six steps take it from a sealed box to an asset showing online on the dashboard. Do them in order — each one depends on the last, and the two that people skip (join the tailnet, create the asset) are the two that force a second trip when they are left out.
Step 1
Flash the card and first boot
Every gateway starts as a headless Raspberry Pi running Raspberry Pi OS Lite (64-bit). Flash the card on your laptop with all the settings baked in, so the Pi comes up on the network ready to SSH into — no monitor or keyboard needed.
- Install Raspberry Pi Imager from raspberrypi.com/software and insert the microSD card.
- Choose device, then Raspberry Pi OS Lite (64-bit) under “Raspberry Pi OS (other)”, then the SD card as storage.
- Click Next → Edit Settings before writing, and set:
- Hostname —
<asset>-pi, e.g.site01-scanner. The asset code has to be in there: it is how the app matches this machine to the unit on the tailnet. - Enable SSH (password authentication is fine to start)
- Username + password — the account you’ll SSH in as and the one the collector runs as. The fleet uses
gateway; whatever you pick has to match the Service user on the asset in step 3. - Wi-Fi SSID, password and country — bake these in whenever the Pi will use Wi-Fi for its uplink, including the common case where its one Ethernet port goes straight to the MagMon. See how the Pi reaches the MagMon.
- Locale / timezone
- Hostname —
- Write the card, put it in the Pi, and power it on. Give it a minute on first boot.
- Find and connect to it by hostname:
ssh <user>@<hostname>.local
# e.g. ssh gateway@site01-scanner.localOnce you’re in, bring it fully up to date and reboot:
sudo apt update && sudo apt full-upgrade -y
sudo rebootThen the dependencies. These are per machine, not per asset — do them once and they cover every collector you install later. The first line is all a MagMon-only gateway needs; install the other two now anyway if this unit might ever get sensors or a UPS, so you are not doing apt work in a trailer later:
sudo apt-get install -y python3-requests # MagMon collector
sudo apt-get install -y python3-pymodbus nut-server nut-client # environmental collector
# if python3-pymodbus is not packaged for your release:
# sudo pip3 install --break-system-packages pymodbusAnd the serial groups, needed only if an RS-485 sensor bus is ever plugged in. Harmless on a unit without one, and easy to forget on a unit that gets one later:
sudo usermod -aG dialout,plugdev gateway
# log out and back in for it to take effectsudo raspi-config.Step 2
Join the tailnet (Tailscale)
Tailscale puts every Pi on one private mesh network so you can reach it by a stable 100.x.y.z address from anywhere, no port forwarding. All fleet devices live on the ops@example-imaging.com tailnet.
Do this before the collector, while the Pi is still on your bench and easy to reach. Everything after this step — copying scripts up, reading journals, fixing a bad address — is then the same whether the Pi is on your desk or in a trailer three hours away, and you never have to plan a second trip to finish an install.
- Install it on the Pi:
curl -fsSL https://tailscale.com/install.sh | sh- Bring it up — this prints a login URL:
sudo tailscale up- Open that URL and sign in with the ops@example-imaging.com account to join the device to that tailnet.
- Approve the machine in the Tailscale admin console. For an always-on gateway, open its menu and disable key expiry so it never drops off.
- Confirm its tailnet address:
tailscale ip -4
tailscale statustailscale status reading like a fleet list rather than a wall of IPs, and it is how the app tells “the Pi is offline” apart from “the Pi is up but not reporting” — a distinction that decides which of the two troubleshooting paths below you take.Step 3
Create the asset in Admin
The asset has to exist in the dashboard before there is a script to install: the generator bakes that asset’s gateway token, its device address and its service user into the file. Skipping ahead here is the usual cause of an install that runs cleanly and reports nothing.
Admin → + Add asset, then:
| Asset name | The unit code, e.g. MM-1002Goes in the filenames, the service names and the Pi hostname. Match it exactly. |
| Site name + address | The facility, e.g. a hospital and its street addressThe address is geocoded for outside weather on the card — a rough address is better than none, and it can be edited later. |
| Modality | MRI / PET/CT / Nuc MedMRI is the only one with a MagMon to scrape; the device fields appear only for it. Everything else is an environmental unit — see Part 2. |
| Stale threshold | Comfortably above the poll intervalA 5-minute poll wants about 30 minutes, so one missed report is not an alarm. |
| Service user | The login user on the target machineThe one you set when flashing. It becomes User= in the systemd unit; wrong here fails the service with 217/USER. |
| MagMon address + port | MRI only — the device’s address as the Pi sees itNot how you reach it from the office or over a VPN. See how the Pi reaches the MagMon. |
| MagMon credentials | Pre-filled with the fleet defaultChange them only if this site’s device was given its own login. |
mm-1002-pi and mm-1002pi both match, while an adjacent digit does not, so a shorter code cannot match a longer one. Get this wrong and everything still reports — you simply never get the “Pi offline” signal that tells a dead gateway apart from a dead device.Step 4
Generate the install script in Admin
Nothing is copied by hand. Each asset gets its own generated collector and its own systemd unit, carrying that asset’s gateway token and its MagMon’s address — so a script is only ever valid for the one asset it was generated for.
- In the web app: Admin → Existing assets → pick the asset → Get install script.
- Set Service user to the login user on the target machine — the account you created in step 1 (
pion a stock image). On the shared server it’s that box’s user instead. Runwhoamion the machine if you’re unsure; a wrong value here fails the service with217/USER. - Leave Collector on MagMon. The switch only appears on an MRI asset, because that is the only kind of unit that could run either collector — a magnet fitted with sensors or a UPS runs both, and you download each pair in turn. Part 2 covers that.
- Set the poll interval if it needs to differ from the default, then click Regenerate. Five minutes is the fleet default and it does not limit resolution: each cycle collects every new one-minute row the device has logged since the last one, so the stored history is minute-by-minute either way.
- Click Download script and Download systemd unit. You now have two files in
~/Downloads:magmon-gateway-<ASSET>.pymagmon-gateway-<ASSET>.service— already carries the rightUser=and paths, so there is nothing to edit
Repeat for every asset you’re deploying — each one is its own pair of files.
Step 5
Copy the files to the Pi with SCP
scp copies files over the same SSH connection. It runs from your laptop, not from the Pi.
Send the two files from step 4 up:
scp ~/Downloads/magmon-gateway-<ASSET>.py <user>@<host>:/home/<user>/
scp ~/Downloads/magmon-gateway-<ASSET>.service <user>@<host>:/home/<user>/Other shapes you’ll want:
| A whole folder (recursive) | scp -r ./scripts <user>@<host>:/home/<user>/ |
| Pull a file back down (e.g. a log) | scp <user>@<host>:/home/<user>/gateway.log ./ |
| Through a Device Manager tunnel port | scp -P <tunnel-port> ./file <user>@<tunnel-host>:/home/<user>/ |
scp uses -P (capital) for the port, while ssh uses -p (lowercase). Easy to mix up. Over Tailscale, just use the tailnet IP or MagicDNS name as <host> — no tunnel needed.Step 6
Install the collector and verify it reports
Now SSH in — everything below runs on the Pi. Three files: the collector script (the same on every Pi, installed once per host), this unit’s config (its token and MagMon address), and its service. Download all three from Admin → Assets → Get install script.
# once per host — skip if /opt/magmon-gateway.py is already there
sudo install -o root -g root -m 755 /home/<user>/magmon-gateway.py /opt/magmon-gateway.py
# per unit — the config holds its token and MagMon password, so 600
sudo install -d -m 755 /etc/nm-collector
sudo install -o <user> -g "$(id -gn <user>)" -m 600 \
/home/<user>/<ASSET>-magmon.json /etc/nm-collector/<ASSET>-magmon.json
sudo cp /home/<user>/magmon-gateway-<ASSET>.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now magmon-gateway-<ASSET>Verify — all four must pass before you call it done:
systemctl status magmon-gateway-<ASSET> --no-pager # active (running)
journalctl -u magmon-gateway-<ASSET> -n 30 --no-pager # "reported N sample(s)", no errors
pgrep -af <ASSET>-magmon.json # ONE python line, and you can see it
ls -l /var/tmp/magmon-gateway-<ASSET>.hwm # exists after the SECOND cyclessh host "… pgrep -c -f <ASSET>-magmon.json", the count comes back one too high: the remote shell’s own command line contains the text being searched for, so pgrep matches the question. Typed at a prompt it is fine, which is why this is easy to miss. -af prints what matched, so a real collector (/usr/bin/python3 /opt/magmon-gateway.py /etc/nm-collector/<ASSET>-magmon.json) is obvious next to a bash -c echoing your own command..hwm file. If the second cycle also says reported 1 sample(s) and no .hwm appears, the unit is storing one reading per poll instead of every minute: confirm minute resolution.…and the asset flips to online on the dashboard within a couple of minutes. If any check fails, go to troubleshooting. Once it’s green you can clean up the copies in the home directory:
rm -f /home/<user>/magmon-gateway*.py /home/<user>/magmon-gateway-<ASSET>.service /home/<user>/<ASSET>-magmon.jsonpgrep -c prints anything but 1, you have a duplicate./opt/magmon-gateway-<ASSET>.py with its token baked in. scripts/deploy-collector.sh migrates a host in one run once each unit’s config is staged beside it — see docs/pi-redeploy-guide.md. The new collector picks up the old one’s high-water mark, so nothing is re-sent or skipped.Part 2
Environmental hardware (sensors + UPS)
Temperature and humidity sensors on an RS-485 bus, and the UPS feeding the trailer. The hardware and the collector are the same wherever they go — what changes is whether they are the whole job or an addition to a magnet. Steps 1 to 3 are shared; Step 4 comes in two versions, one for each case. Everything about the Pi itself is Part 1, so do that first.
What it adds, and to what
This is an additive set of channels, not a different kind of unit. Whichever of these are fitted is what the dashboard draws:
| Zone temp / humidity | up to three XY-MD02 sensors on one RS-485 busOne is fine. Two or three is a wiring decision, never a code change — only the zones that report get drawn. |
| Mains / UPS power | on-battery state, battery %, input volts, via NUTThe fleet-wide outage rule starts paging for this unit the moment it reports. |
| Collector | env-gateway-<ASSET>.py |
| Extra packages | python3-pymodbus, nut-server, nut-client |
The two cases:
| On an MRIStep 4 · MRI | An addition. The asset stays modality MRI and keeps its MagMon address; its Pi runs BOTH collectors, reporting one asset. Helium, bay temperature and mains power land in the same row and show on one card. |
| On a PET/CT or Nuc Med unitStep 4 · PET/CT · Nuc Med | The whole job. There is no MagMon, so the asset is created with that modality, the device-address fields disappear, and one collector runs on its own. |
The collector names differ from the MagMon one on purpose — separate service, path, lock file and log tag — because on a magnet the two run side by side and would otherwise fight over each other’s files.
Step 1
Wire the sensors and the UPS
The three sensors share one RS-485 pair back to a USB adapter on the Pi. RS-485 is a bus: every sensor lands on the same two wires (A to A, B to B), daisy-chained from one to the next rather than home-run to the Pi.
- Mount one XY-MD02 in each zone: Section 1 Engineering, Section 2 Tech / Patient, Section 3 Equipment.
- Daisy-chain A and B between them and back to the USB–RS485 adapter. Keep the pair away from mains runs; it is a differential pair and will tolerate a lot, but not a compressor contactor.
- Power the sensors from their supply (they are not bus-powered).
- Plug the adapter into the Pi, and the UPS’s USB data cable into the Pi.
- The Pi itself plugs into the UPS — a gateway that dies with the mains cannot report the outage.
Confirm the Pi can see both:
ls -l /dev/ttyUSB* # the RS-485 adapter, usually /dev/ttyUSB0
lsusb # the UPS should appear by brandsudo usermod -aG dialout,plugdev gatewayls -l /dev/ttyUSB0 shows root dialout on some and root plugdev on others — both seen in this fleet. And a trailing + on those permissions is an ACL that logind grants the logged-in user — so a hand-run scan works over SSH while the collector, which has no seat, is refused. Testing by hand proves nothing about the service. Check the real thing with getfacl /dev/ttyUSB0 and groups.Step 2
Address each sensor: 1 = S1, 2 = S2, 3 = S3
This is the step that decides which zone is which, and the one worth slowing down for. The collector reads Modbus unit 1 into S1 Engineering, unit 2 into S2 Tech / Patient and unit 3 into S3 Equipment.
They all ship as address 1, so they must be addressed one at a time — two unaddressed sensors on the same bus both answer to 1 and you get garbage. Connect one, set it, unplug it, move to the next.
Set the address with the vendor’s Windows tool, or from the Pi:
# ONE sensor connected. Set NEW to 1, 2 or 3 for the zone it is going into.
python3 - <<'PY'
from pymodbus.client import ModbusSerialClient
NEW = 2 # 1 = Engineering, 2 = Tech / Patient, 3 = Equipment
c = ModbusSerialClient(port="/dev/ttyUSB0", baudrate=9600, bytesize=8,
parity="N", stopbits=1, timeout=2)
c.connect()
try:
r = c.write_register(address=0x0101, value=NEW, device_id=1)
except TypeError: # older pymodbus spells it slave=
r = c.write_register(address=0x0101, value=NEW, slave=1)
print("FAILED" if r.isError() else f"sensor is now address {NEW}")
c.close()
PYWith all three wired up, confirm the bus — this is also how you identify a sensor:
python3 - <<'PY'
from pymodbus.client import ModbusSerialClient
c = ModbusSerialClient(port="/dev/ttyUSB0", baudrate=9600, bytesize=8,
parity="N", stopbits=1, timeout=0.4)
print("port open:", c.connect())
found = 0
for uid in range(1, 8):
try:
# pymodbus renamed the keyword: >=3.7 wants device_id=, older wants slave=.
try:
rr = c.read_input_registers(address=1, count=2, device_id=uid)
except TypeError:
rr = c.read_input_registers(address=1, count=2, slave=uid)
except Exception as e:
# An address with no sensor on it RAISES rather than returning an error
# result. Catching it is what lets the scan reach address 2.
print(f"address {uid}: no answer ({type(e).__name__})")
continue
if rr and not rr.isError():
t, h = rr.registers[0] / 10.0, rr.registers[1] / 10.0
print(f"address {uid}: {t * 9 / 5 + 32:.1f} F {h:.1f} %RH <-- FOUND")
found += 1
print(f"{found} sensor(s) on the bus")
c.close()
PYExactly three lines marked FOUND, at addresses 1, 2 and 3 — a unit fitted with fewer sensors shows correspondingly fewer. Zero found, with the port opening, is a bus problem rather than an addressing one: check the sensor has its own 5–30 VDC supply (it is not bus-powered), swap A and B (the commonest RS-485 fault, and adapter silkscreens disagree with each other), confirm /dev/ttyUSB0 really is the RS-485 adapter with lsusb, and re-run at 4800 and 19200 baud before concluding a sensor is dead. To prove which is which, warm one sensor with your hand for a minute and re-run — the address that moves is that zone. Do this before you close the trailer up.
Step 3
Set up the UPS with NUT
The collector reads the UPS by shelling out to upsc, so NUT has to be answering locally before the collector will report anything about power.
sudo apt-get update
sudo apt-get install -y nut-server nut-client/etc/nut/ups.conf — the name in brackets matters:
[ups]
driver = usbhid-ups
port = auto
desc = "Trailer UPS"upsc ups. If you name the section anything else, every power field reports blank and the dashboard shows Power unknown — which reads as a broken link, not as a naming mistake. Either call it ups or change UPS_NAME at the top of the script.The remaining three files:
# /etc/nut/nut.conf
MODE=standalone
# /etc/nut/upsd.users
[monuser]
password = <pick-a-password>
upsmon master
# /etc/nut/upsmon.conf (append)
MONITOR ups@localhost 1 monuser <pick-a-password> mastersudo systemctl restart nut-server nut-monitor
sudo systemctl enable nut-server nut-monitor
upsc upsupsc ups should print a block of values. The three the collector uses are ups.status (OL on line, OB on battery), battery.charge and input.voltage. Pull the UPS’s mains plug for a few seconds and watch ups.status flip to OB — that is the whole alarm path proven at the source.
Step 4 · MRI
Install it alongside a MagMon collector
A UPS or a bay sensor on an MRI unit is an addition, not a different kind of unit. The asset stays on modality MRI with its MagMon address intact, and its Pi runs both collectors — two scripts, two services, two lock files, one asset.
Nothing about the asset changes — no new asset, no modality edit, no second tile on the dashboard.
- Wire and address only the sensors actually being fitted (Steps 1–3 above). One sensor gives one tile; the collector polls all three addresses regardless, so a second one added later is a wiring job with no redeploy.
- Admin → the asset → Get install script, then set Collector to Environmental. Set the poll interval to 1 minute — see below for why that is free on a magnet — and download the script, the config and the systemd unit.
- Install them exactly as in Step 6, substituting the environmental names: the shared environmental script (once per host), this unit’s
-env.jsonconfig, and its unit. Nothing about either install changes because the other one is there:sudo install -o root -g root -m 755 /home/<user>/env-gateway.py /opt/env-gateway.py sudo install -o <user> -g "$(id -gn <user>)" -m 600 \ /home/<user>/<ASSET>-env.json /etc/nm-collector/<ASSET>-env.json sudo cp /home/<user>/env-gateway-<ASSET>.service /etc/systemd/system/ sudo systemctl daemon-reload sudo systemctl enable --now env-gateway-<ASSET> - Watch a cycle. You want one line per fitted zone, one UPS line, and
reported N channel(s):journalctl -u env-gateway-<ASSET> -f
pgrep -c -f magmon-gateway counts MagMon collectors only, but a bare pgrep -c -f gateway prints 2 on a mixed unit. That is the healthy state, not a doubled collector. Check the specific names:pgrep -c -f <ASSET>-magmon.json # 1
pgrep -c -f <ASSET>-env.json # 1Step 4 · PET/CT · Nuc Med
Create the asset and install the collector
A unit with no MagMon — a PET/CT or Nuc Med trailer or lab — is an ordinary asset whose only channels are environmental. One collector, no device address, everything else identical to Part 1.
- Create the asset as in Step 3, setting Modality to PET/CT or Nuc Med. The MagMon address and credential fields disappear — there is no device to reach. A stale threshold of about 20 minutes suits a 5-minute poll.
- Get install script hands back the environmental collector automatically; there is no Collector switch, because this unit can only run the one.
- Set the poll interval. Unlike a magnet, every poll here is a new row — five minutes is the sensible default, and one minute is a five-fold increase in stored history for readings that move slowly.
- Download the script and the systemd unit.
On the Pi — dependencies first:
sudo apt-get install -y python3-requests python3-pymodbus nut-server nut-client
# if python3-pymodbus is not packaged for your release:
# sudo pip3 install --break-system-packages pymodbus
sudo usermod -aG dialout,plugdev <user> # then log out and back inCopy the three files up (same SCP as Part 1), then install:
# once per host -- skip if /opt/env-gateway.py is already there
sudo install -o root -g root -m 755 /home/<user>/env-gateway.py /opt/env-gateway.py
# per unit -- the config holds its gateway token, so 600
sudo install -d -m 755 /etc/nm-collector
sudo install -o <user> -g "$(id -gn <user>)" -m 600 \
/home/<user>/<ASSET>-env.json /etc/nm-collector/<ASSET>-env.json
sudo cp /home/<user>/env-gateway-<ASSET>.service \
/etc/systemd/system/env-gateway-<ASSET>.service
sudo systemctl daemon-reload
sudo systemctl enable --now env-gateway-<ASSET>Verify — all four before you call it done:
systemctl status env-gateway-<ASSET> --no-pager # active (running)
journalctl -u env-gateway-<ASSET> -n 30 --no-pager # three zones + a UPS line
pgrep -c -f <ASSET>-env.json # prints exactly 1
upsc ups | grep ups.status # OLThen check the dashboard: the asset shows a tile per fitted zone with readings, a green Wall power chip, and a battery percentage. If a zone you fitted reads an em-dash, go back to Step 2 and re-run the bus scan.
pgrep -c prints anything but 1, you have duplicates.Alert rules for temperature and power
Power is already covered fleet-wide. A rule of On UPS battery = 1 raises a critical POWER OUTAGE for any asset that reports the channel, so a new trailer is covered the moment it starts reporting — nothing to add per unit.
Temperature and humidity limits are deliberately unset. Add them in Admin → Alerts once the right numbers are known for the space:
- Pick the channel under the Environment group — each zone has its own temperature and humidity.
- Leave Scope on All assets for a fleet-wide limit, or pick one asset to override the fleet default for just that unit.
- Set the comparator and the threshold, and save.
Part 3
Site and network setup
Two different questions, and it is worth keeping them apart: how the Pi reaches the MagMon (a decision made at install time, and the one that decides the device address you put on the asset), and how you reach the Pi afterwards. Cellular sites come in behind an iR305; everything on the tailnet is reachable directly over Tailscale from anywhere.
How the Pi reaches the MagMon
The collector talks to the device over the LAN, so the address you enter as the MagMon address on the asset is whatever the Pi can reach — not what your laptop reaches, and not what a VPN server reaches. Two shapes cover the fleet.
A. Both on the router’s LAN — the simple one
The MagMon and the Pi each plug into the site router (or a small switch behind it) and get addresses on the same subnet. Nothing to configure on the Pi; the device address is its LAN IP. Use this whenever there is a spare LAN port. It also means the MagMon’s own web interface is reachable from anything else on that LAN.
B. Pi on Wi-Fi, MagMon on a private cable — when the device stays off the site network
The Pi uplinks over the router’s Wi-Fi and its single Ethernet port runs straight to the MagMon. The device is then reachable only from its own Pi, which is a real security gain and keeps a hospital network out of the picture entirely. Two things have to be true or it silently does not work:
- Bake the Wi-Fi credentials in when you flash the card (Part 1, Step 1). With Ethernet committed to the device, Wi-Fi is the only uplink — a Pi that cannot join it is a Pi you cannot reach.
- The Wi-Fi subnet and the device subnet must differ. Two interfaces on one subnet makes routing ambiguous. Check the router’s LAN range before you start.
A direct cable has no DHCP server, so give eth0 a static address on the device’s subnet. Keep the MagMon’s existing address if it has one — matching the Pi to the device is far easier than reconfiguring the device, and a device that kept its address through a router swap needs nothing done to it at all. ipv4.never-default is the load-bearing flag: it keeps the default route on Wi-Fi, so reporting and Tailscale are unaffected.
sudo nmcli con add type ethernet ifname eth0 con-name magmon-link \
ipv4.method manual ipv4.addresses 10.0.0.10/24,169.254.7.10/16 \
ipv4.never-default yes ipv6.method disabled connection.autoconnect-priority 100
sudo nmcli con up magmon-linkThe second, link-local address costs nothing and covers the case where the device was on DHCP from an old router and has fallen back to 169.254.x.x. Then confirm the routing and the device, in that order:
ip -4 addr show eth0 # the static address is on the interface
ip route | head -3 # default MUST be via wlan0, not eth0
ping -c3 -W2 <magmon-host>
curl -sS -m 8 -o /dev/null -w 'HTTP %{http_code}\n' http://<magmon-host>/Any HTTP status — 302 to the login page is the usual one — means the collector will be able to talk to it.
sudo apt-get install -y tcpdump
sudo tcpdump -ni eth0 -c 20 # ARP/DHCP chatter reveals its IP or subnet
ip neigh show dev eth0 # anything that has answerednmcli con add does not replace a profile of the same name, it adds another one — run it twice and a reboot picks one at random. If you have repeated it, clear them out and add one:for u in $(nmcli -t -f UUID,NAME con show | awk -F: '$2=="magmon-link"{print $1}'); do sudo nmcli con delete "$u"; doneRemote access via the iR305 (InHand Device Manager)
Cellular sites reach the internet through an InHand iR305 router. Its cloud console — the InHand IoT Device Manager — can open a secure tunnel to the Raspberry Pi sitting on the router’s LAN, handing you a temporary public host and port that forwards straight to the Pi’s SSH (port 22).
it@example-imaging.com.- Sign in and open the iR305 for the site; confirm it shows online.
- Open its remote-access / tunnel tool and create a tunnel to the Pi’s LAN address on port
22(the Pi’s IP on the router’s network — check the router’s connected-devices list if you’re unsure of it). - Copy the public host and port the tunnel issues.
- SSH through it from your laptop:
ssh <user>@<public-host> -p <public-port>ssh to it. Once a site is also on Tailscale (below), you can skip the tunnel entirely.SSH into Tailscale assets
From any machine signed in to the same tailnet, SSH straight to a device’s tailnet IP (or its MagicDNS name):
ssh <user>@100.x.y.z
# or, with MagicDNS:
ssh <user>@<hostname>The central Pi server is reachable two ways — on the operations LAN, or over Tailscale from anywhere:
| Pi server — LAN | ssh <user>@10.0.0.10 |
| Pi server — Tailscale | ssh <user>@100.100.10.20 |
List every machine on the tailnet and its IP with tailscale status from any joined device.
Open a unit's MagMon web interface
Do not browse the MagMon’s LAN address directly. That address is private and several sites reuse it, so over a tailnet subnet route it reaches whichever Pi currently owns that route — with nothing on screen to tell you which unit you got. This is not hypothetical: a week of minute data was read from one unit while believed to be another, and only the readings themselves gave it away.
Tunnel through a named Pi instead. The Pi is at exactly one site, so the tunnel can only land on one MagMon. Pick a different local port per unit and you can hold several open at once:
ssh -L 8080:<magmon-host>:80 <user>@<pi-host>
# then open http://localhost:8080Ready-made per unit. A shared address is one another unit also uses — for those the tunnel is the only safe route.
| MM-1001Example Regional | ssh -L 8080:192.168.1.50:80 gateway@mm1001-pi.example.ts.netthen http://localhost:8080 |
| MM-1002Example Communityshared address | ssh -L 8081:192.168.1.99:80 gateway@mm1002-pi.example.ts.netthen http://localhost:8081 |
| MM-1003Example Imagingshared address | ssh -L 8082:192.168.1.99:80 gateway@mm1003-pi.example.ts.netthen http://localhost:8082 |
To see which Pi currently owns a route from your own machine: tailscale status --json | jq -r '.Peer[] | select(.PrimaryRoutes) | "\(.HostName)\t\(.PrimaryRoutes | join(","))"'
Replace a collector script on a running Pi
Every host runs the same collector script; what makes a collector a particular unit’s is its config file in /etc/nm-collector. So a change to the collector code is one file, copied once to each of the 10 hosts — the Pi server’s five units pick it up together. The collector version shown against each asset in Admin tells you which are still behind. scripts/deploy-collector.sh does all of this per host; the steps below are the same thing by hand.
This is a script swap and a restart, not a reinstall. Configs and units do not change, so there is no daemon-reload and no re-enable. Do the units in maintenance first and live hospital sites last.
- In Admin → Assets, click Get install script on any asset and download the collector script. It is the same file for every unit and lands in
~/Downloadsasmagmon-gateway.py. - Copy it to the host’s home directory (one password prompt).
- Back up the running script, install the new one over it, and restart every collector service on that host.
- Verify three things: the service is
active, the journal shows the new version at startup, and within a few minutes the asset’s collector version in Admin matches.
Copy up, then install and restart:
scp ~/Downloads/magmon-gateway.py <user>@<pi-host>:/home/<user>/
ssh <user>@<pi-host> 'sudo cp /opt/magmon-gateway.py /opt/magmon-gateway.py.bak; \
sudo install -o root -g root -m 755 /home/<user>/magmon-gateway.py /opt/magmon-gateway.py && \
sudo systemctl restart "magmon-gateway-*" && sleep 20 && \
systemctl list-units "magmon-gateway-*" --no-pager && \
journalctl -u "magmon-gateway-*" -n 10 --no-pager'A healthy result is active followed by a startup line carrying the version, e.g. starting for asset ‘MM-1002’ v2026.09.24-2. Restarting makes the collector run a cycle immediately, so one extra reading lands out of step with the five-minute rhythm — that is the restart, not a fault.
If it misbehaves, roll back to the script that was running a moment ago:
ssh <user>@<pi-host> 'sudo cp /opt/magmon-gateway.py.bak /opt/magmon-gateway.py && \
sudo systemctl restart "magmon-gateway-*"'The <user>@<pi-host> for each unit. Units with no Tailscale node of their own run on the shared Pi server — use its address for those.
| MM-1001Example Regional | gateway@mm1001-pi.example.ts.net |
| MM-1002Example Community | gateway@mm1002-pi.example.ts.net |
| MM-1003Example Imaging | gateway@mm1003-pi.example.ts.net |
The shared Pi server (gw-server)
Some sites don’t get their own Pi — they’re reached over a VPN from one central box instead. The Pi server — hostname gw-server, 10.0.0.10 on the LAN or 100.100.10.20 over Tailscale — runs one collector per asset, each installed exactly as in Part 1. Every asset is a matching pair owned by the gateway user: a .py collector and its .service unit, both named magmon-gateway-<ASSET>.
Collector services
| magmon-gateway-MM-1001 | Reports telemetry for MM-1001 |
| magmon-gateway-MM-1002 | Reports telemetry for MM-1002 |
| magmon-gateway-MM-1003 | Reports telemetry for MM-1003 |
| magmon-gateway-MM-1004 | Reports telemetry for MM-1004 |
| magmon-gateway-MM-1005 | Reports telemetry for MM-1005 |
| magmon-gateway-MM-1006 | Reports telemetry for MM-1006 |
SSH in and manage a collector by its service name:
ssh gateway@100.100.10.20 # or 10.0.0.10 on the LAN
ls -la /etc/nm-collector # one config per unit
systemctl status magmon-gateway-MM-1002
journalctl -u magmon-gateway-MM-1002 -f # live logs
sudo systemctl restart magmon-gateway-MM-1002Check the whole box at once:
systemctl list-units 'magmon-gateway-*' --no-pager # one per asset, all running
pgrep -c -f magmon-gateway # should equal the asset countTake an installed collector out of service — stop it now and keep it from starting on boot (--now does both):
sudo systemctl disable --now magmon-gateway-<ASSET>
systemctl status magmon-gateway-<ASSET> # confirm: inactive (dead), disabled
pgrep -af <ASSET>-magmon.json # no output = nothing left runningRe-enable it later with sudo systemctl enable --now magmon-gateway-<ASSET>. The service is named magmon-gateway-<ASSET>, while every unit’s process is the same /opt/magmon-gateway.py with its own config — so match on <ASSET>-magmon.json for one unit, or magmon-gateway for all of them.
Part 4
Run and troubleshoot
What to do when an asset goes red. Work the triage table first — it splits the two failure families that look identical on the dashboard (the site is unreachable vs the site is reachable but nothing is arriving) before you start changing things.
Troubleshooting a gateway
“Offline” on the dashboard means no fresh telemetry — nothing more. A dead magnet, a dead Pi, a dead router and a Pi that simply can’t reach the internet all produce the same red chip. Start by finding out which one you have.
Triage — run these three, in this order:
# 1. Can you reach the Pi at all? (from your laptop, on the tailnet)
ssh <user>@<tailnet-ip>
# 2. Is the collector running, and is it saying anything?
systemctl status magmon-gateway-<ASSET> --no-pager
journalctl -u magmon-gateway-<ASSET> -n 50 --no-pager
# 3. Can the Pi reach the outside world, by name?
ping -c2 8.8.8.8 # raw-IP internet
getent hosts example-project.supabase.co # the DNS lookup that actually mattersWhat the answers mean
| SSH fails, Tailscale shows the node down | The Pi or the site link is down — power, network, or router. Start at the iR305. |
| SSH works, service inactive/failed | Collector problem. See the install-failure table below. |
| SSH works, service running, ping OK, getent FAILS | Pi-local DNS failure. This is the classic one — see below. |
| Everything green, journal says 'reported OK' | Telemetry is flowing. If values look wrong, judge the magnet itself. |
| Card shows 'MagMon not connected' or 'MagMon silent' | A unit with two collectors has one down. The other keeps it reporting, so the status chip stays green — check that specific service, not the asset. |
| Card shows 'Env silent' | Same, the other way round: the sensors or UPS link stopped while the MagMon side keeps reporting. |
| Every zone blank on a unit with sensors | Almost always serial permissions, not sensors — see the dialout/plugdev note in Part 2. |
curl to the MagMon all use a raw IP or Tailscale’s own path — none of them need DNS. So they all pass while the one DNS-dependent call, reporting to example-project.supabase.co, fails silently. The journal tell is socket.gaierror [Errno -3] Temporary failure in name resolution. Confirm with the getent line above; if ping 8.8.8.8 works and the lookup doesn’t, it’s DNS only. Fix by pointing the Pi at a public resolver, then restarting the collector:sudo nano /etc/resolv.conf # nameserver 1.1.1.1 / 8.8.8.8
sudo systemctl restart magmon-gateway-<ASSET>
# make it survive a reboot:
# add to /etc/dhcpcd.conf:
# static domain_name_servers=1.1.1.1 8.8.8.8pgrep -c -f magmon-gateway-<ASSET> prints 1 but journalctl shows nothing recent, the process is usually fine — Python is block-buffering its output. Current units set Environment=PYTHONUNBUFFERED=1; units deployed before that don’t. Trust systemctl status and CPU time over a live journalctl -f, and fix it by re-downloadingboth files from Admin and redoing step 6. Units must live in /etc/systemd/system/ — confirm which file is actually loaded with systemctl cat magmon-gateway-<ASSET> | grep FragmentPath.13-May-06. The collector only sends rows newer than the last one it sent, so every row now looks older than the stored high-water mark and the feed goes quiet while the service looks perfectly healthy. Clear the mark to recover:sudo systemctl stop magmon-gateway-<ASSET>
sudo rm -f /var/tmp/magmon-gateway-<ASSET>.hwm
sudo systemctl start magmon-gateway-<ASSET>Install-time failures (fresh deploys):
Symptom → fix
| ModuleNotFoundError: 'requests' | python3-requests was never installed. Install it, then restart the service. |
| Service state 217/USER | The Service user you set in Admin doesn't exist here. Check whoami, re-download, redo step 6. |
| Permission denied on the .py | Wrong owner or mode — re-run the sudo install line in step 6 exactly. |
| 'another copy already holds the lock' | A second copy really is running. pgrep -af, kill the extras, remove any cron entry. |
| 'Cannot reach MagMon at <ip>:<port>' | Wrong or unroutable MagMon address for this site. Fix it in Admin → Edit, re-download, redo step 6. |
| Still offline after a few minutes | Token mismatch — most likely rotated. Re-download the current script and redo step 6. |
magmon_collect_* scripts and their log files). They are unrelated to this app, but other reporting may still depend on them — don’t disable, delete, or clean up those cron entries as part of gateway work. If one is filling the disk, ask before touching it, and archive anything you remove: several of those scripts exist on the Pi and nowhere else.Confirm a collector is storing every minute
A collector can be perfectly healthy and still be storing a twelfth of what the device is serving. The MagMons log a row a minute; the collector fetches the last hour each cycle and files every row newer than the last one it sent. If it cannot read the device’s timestamps it degrades quietly instead of failing: it files just the newest row, stamped with the current time. The unit reports on schedule, the charts draw, and the resolution is gone. It ran that way fleet-wide from August 2026 until it was noticed on an install in September.
The tell, in the journal:
journalctl -u magmon-gateway-<ASSET> -n 20 --no-pager | grep 'reported'| Healthy | reported 5 sample(s) — on a 5-minute pollOne per minute elapsed. The first cycle after any restart reports 1 by design. |
| Degraded | reported 1 sample(s) — every cycle, foreverAnd no .hwm file, which is only written when a row’s timestamp parses. |
ls -l /var/tmp/magmon-gateway-<ASSET>.hwm # missing = nothing has ever datedThe same thing is visible from the dashboard side without SSH: a degraded unit stores exactly 60 ÷ poll-minutes rows an hour — a flat 12/hour on a five-minute poll, never more.
The fix is a script redeploy, not a device change:
- Check the asset’s collector version in Admin. Anything older than the current generator is suspect, and the panel marks it “behind”.
- Redeploy it — replace a collector script. Script swap and restart; the systemd unit does not change.
- Watch the second cycle after the restart. Five samples and a
.hwmfile means it is filing every minute.
curl -s --http0.9 --max-time 25 -u <user>:<pass> \
"http://<magmon-host>/goform/showMinutesCGI?num_hours=1&start_day=0&start_hour=1" \
| grep -oE '[0-9]{2}:[0-9]{2}' | tail -15Everyday Raspberry Pi commands
A quick cheat sheet for working on a Pi over SSH.
Getting around
| pwd | Print the folder you're currently in |
| ls -la | List everything, including hidden files, with details |
| cd /path/to/dir | Change directory |
| cd ~ | Go to your home directory |
| cd .. | Go up one level |
Files
| cat file | Print a file to the screen |
| less file | Scroll through a long file (q to quit) |
| nano file | Edit a file (Ctrl+O save, Ctrl+X exit) |
| cp a b | Copy a to b |
| mv a b | Move or rename a to b |
| rm file | Delete a file (no undo) |
| mkdir dir | Make a new folder |
| chmod +x script.py | Make a script executable |
| du -sh * | What's using the space in this folder |
System
| sudo apt update && sudo apt full-upgrade -y | Update all packages |
| df -h | Disk space, human-readable |
| free -h | Memory in use |
| htop | Live processes + CPU (q to quit) |
| uptime | How long it's been up, load average |
| date | The Pi's clock — worth checking during odd faults |
| sudo reboot | Restart the Pi |
| sudo shutdown now | Power it off |
Services (systemd)
| systemctl status <name> | Is the service running? |
| sudo systemctl restart <name> | Restart it |
| sudo systemctl enable --now <name> | Start now + on every boot |
| sudo systemctl disable --now <name> | Stop now + never start on boot |
| journalctl -u <name> -f | Follow its live logs |
| journalctl -u <name> -n 50 --no-pager | Last 50 lines, no pager |
| systemctl cat <name> | Show the unit file that's actually loaded |
Processes
| pgrep -af <pattern> | Every matching process, with its full command |
| pgrep -c -f <pattern> | Just the count — 1 is what you want per collector |
| sudo kill <pid> | Stop a stray process |
Scheduling (cron)
| crontab -e | Edit your scheduled jobs |
| crontab -l | List your scheduled jobs |
| * * * * * command | min hour day month weekday — then the command |
systemctl status magmon-gateway-<ASSET>, not crontab. Use cron for one-off periodic maintenance tasks.Networking
| hostname -I | This Pi's IP address(es) |
| ip a | All network interfaces in detail |
| ping 8.8.8.8 | Test raw-IP connectivity (Ctrl+C to stop) |
| getent hosts <hostname> | Test DNS specifically — resolves a name to an IP |
| curl -s http://<magmon-ip> | Read the MagMon directly, bypassing the collector |
| tailscale status | Every device on the tailnet + its IP |
| tailscale ip -4 | This device's tailnet IP |