Knowledge base

Documentation

How the MagMon fleet is built and kept running, in the order you actually do it: take a Raspberry Pi from the sealed box to an asset reporting on the dashboard, add environmental sensors and a UPS if the unit gets them, wire the site up for remote access, then keep it healthy. Read it top to bottom the first time; after that, jump to the part you need.

Part 1

Stand up a gateway

A gateway is one Raspberry Pi (or one shared server) running one collector per asset. These six steps take it from a sealed box to an asset showing online on the dashboard. Do them in order — each one depends on the last, and the two that people skip (join the tailnet, create the asset) are the two that force a second trip when they are left out.

Step 1

Flash the card and first boot

Every gateway starts as a headless Raspberry Pi running Raspberry Pi OS Lite (64-bit). Flash the card on your laptop with all the settings baked in, so the Pi comes up on the network ready to SSH into — no monitor or keyboard needed.

  1. Install Raspberry Pi Imager from raspberrypi.com/software and insert the microSD card.
  2. Choose device, then Raspberry Pi OS Lite (64-bit) under “Raspberry Pi OS (other)”, then the SD card as storage.
  3. Click Next → Edit Settings before writing, and set:
    • Hostname — <asset>-pi, e.g. site01-scanner. The asset code has to be in there: it is how the app matches this machine to the unit on the tailnet.
    • Enable SSH (password authentication is fine to start)
    • Username + password — the account you’ll SSH in as and the one the collector runs as. The fleet uses gateway; whatever you pick has to match the Service user on the asset in step 3.
    • Wi-Fi SSID, password and country — bake these in whenever the Pi will use Wi-Fi for its uplink, including the common case where its one Ethernet port goes straight to the MagMon. See how the Pi reaches the MagMon.
    • Locale / timezone
  4. Write the card, put it in the Pi, and power it on. Give it a minute on first boot.
  5. Find and connect to it by hostname:
ssh <user>@<hostname>.local
# e.g. ssh gateway@site01-scanner.local

Once you’re in, bring it fully up to date and reboot:

sudo apt update && sudo apt full-upgrade -y
sudo reboot

Then the dependencies. These are per machine, not per asset — do them once and they cover every collector you install later. The first line is all a MagMon-only gateway needs; install the other two now anyway if this unit might ever get sensors or a UPS, so you are not doing apt work in a trailer later:

sudo apt-get install -y python3-requests                    # MagMon collector
sudo apt-get install -y python3-pymodbus nut-server nut-client   # environmental collector
# if python3-pymodbus is not packaged for your release:
#   sudo pip3 install --break-system-packages pymodbus

And the serial groups, needed only if an RS-485 sensor bus is ever plugged in. Harmless on a unit without one, and easy to forget on a unit that gets one later:

sudo usermod -aG dialout,plugdev gateway
# log out and back in for it to take effect
Static-ish addressingFor a site Pi, reserve its IP in the router’s DHCP by MAC address rather than hand-configuring one on the Pi — it survives re-imaging and avoids collisions. Extra knobs live under sudo raspi-config.

Step 2

Join the tailnet (Tailscale)

Tailscale puts every Pi on one private mesh network so you can reach it by a stable 100.x.y.z address from anywhere, no port forwarding. All fleet devices live on the ops@example-imaging.com tailnet.

Do this before the collector, while the Pi is still on your bench and easy to reach. Everything after this step — copying scripts up, reading journals, fixing a bad address — is then the same whether the Pi is on your desk or in a trailer three hours away, and you never have to plan a second trip to finish an install.

  1. Install it on the Pi:
curl -fsSL https://tailscale.com/install.sh | sh
  1. Bring it up — this prints a login URL:
sudo tailscale up
  1. Open that URL and sign in with the ops@example-imaging.com account to join the device to that tailnet.
  2. Approve the machine in the Tailscale admin console. For an always-on gateway, open its menu and disable key expiry so it never drops off.
  3. Confirm its tailnet address:
tailscale ip -4
tailscale status
Name it clearly — the dashboard reads these namesSet the machine name in the admin console to match the asset (same as the hostname). It keeps tailscale status reading like a fleet list rather than a wall of IPs, and it is how the app tells “the Pi is offline” apart from “the Pi is up but not reporting” — a distinction that decides which of the two troubleshooting paths below you take.

Step 3

Create the asset in Admin

The asset has to exist in the dashboard before there is a script to install: the generator bakes that asset’s gateway token, its device address and its service user into the file. Skipping ahead here is the usual cause of an install that runs cleanly and reports nothing.

Admin → + Add asset, then:

Asset nameThe unit code, e.g. MM-1002Goes in the filenames, the service names and the Pi hostname. Match it exactly.
Site name + addressThe facility, e.g. a hospital and its street addressThe address is geocoded for outside weather on the card — a rough address is better than none, and it can be edited later.
ModalityMRI / PET/CT / Nuc MedMRI is the only one with a MagMon to scrape; the device fields appear only for it. Everything else is an environmental unit — see Part 2.
Stale thresholdComfortably above the poll intervalA 5-minute poll wants about 30 minutes, so one missed report is not an alarm.
Service userThe login user on the target machineThe one you set when flashing. It becomes User= in the systemd unit; wrong here fails the service with 217/USER.
MagMon address + portMRI only — the device’s address as the Pi sees itNot how you reach it from the office or over a VPN. See how the Pi reaches the MagMon.
MagMon credentialsPre-filled with the fleet defaultChange them only if this site’s device was given its own login.
The asset name has to appear in the Pi's hostnameThe app matches a Tailscale machine to an asset by finding the asset code inside the machine’s hostname — separator-insensitive, so mm-1002-pi and mm-1002pi both match, while an adjacent digit does not, so a shorter code cannot match a longer one. Get this wrong and everything still reports — you simply never get the “Pi offline” signal that tells a dead gateway apart from a dead device.
Nothing to fill in for the tokenThe gateway token is generated for you and only ever leaves the app inside a downloaded script. If one is ever rotated, every deployed copy of that asset’s script stops reporting until you re-download and redeploy it.

Step 4

Generate the install script in Admin

Nothing is copied by hand. Each asset gets its own generated collector and its own systemd unit, carrying that asset’s gateway token and its MagMon’s address — so a script is only ever valid for the one asset it was generated for.

  1. In the web app: Admin → Existing assets → pick the asset → Get install script.
  2. Set Service user to the login user on the target machine — the account you created in step 1 (pi on a stock image). On the shared server it’s that box’s user instead. Run whoami on the machine if you’re unsure; a wrong value here fails the service with 217/USER.
  3. Leave Collector on MagMon. The switch only appears on an MRI asset, because that is the only kind of unit that could run either collector — a magnet fitted with sensors or a UPS runs both, and you download each pair in turn. Part 2 covers that.
  4. Set the poll interval if it needs to differ from the default, then click Regenerate. Five minutes is the fleet default and it does not limit resolution: each cycle collects every new one-minute row the device has logged since the last one, so the stored history is minute-by-minute either way.
  5. Click Download script and Download systemd unit. You now have two files in ~/Downloads:
    • magmon-gateway-<ASSET>.py
    • magmon-gateway-<ASSET>.service — already carries the right User= and paths, so there is nothing to edit

Repeat for every asset you’re deploying — each one is its own pair of files.

One asset's files never go on another machineThe script embeds that asset’s unique gateway token and its MagMon’s IP/credentials. Installing the wrong pair reports one asset’s data under another name. The asset is in the filename to keep them straight — download fresh per asset rather than copying a file you already have. If a token is ever rotated, every deployed copy of that script is dead until you re-download it.

Step 5

Copy the files to the Pi with SCP

scp copies files over the same SSH connection. It runs from your laptop, not from the Pi.

Send the two files from step 4 up:

scp ~/Downloads/magmon-gateway-<ASSET>.py      <user>@<host>:/home/<user>/
scp ~/Downloads/magmon-gateway-<ASSET>.service <user>@<host>:/home/<user>/

Other shapes you’ll want:

A whole folder (recursive)scp -r ./scripts <user>@<host>:/home/<user>/
Pull a file back down (e.g. a log)scp <user>@<host>:/home/<user>/gateway.log ./
Through a Device Manager tunnel portscp -P <tunnel-port> ./file <user>@<tunnel-host>:/home/<user>/
It's -P, capital, for scpscp uses -P (capital) for the port, while ssh uses -p (lowercase). Easy to mix up. Over Tailscale, just use the tailnet IP or MagicDNS name as <host> — no tunnel needed.

Step 6

Install the collector and verify it reports

Now SSH in — everything below runs on the Pi. Three files: the collector script (the same on every Pi, installed once per host), this unit’s config (its token and MagMon address), and its service. Download all three from Admin → Assets → Get install script.

# once per host — skip if /opt/magmon-gateway.py is already there
sudo install -o root -g root -m 755 /home/<user>/magmon-gateway.py /opt/magmon-gateway.py

# per unit — the config holds its token and MagMon password, so 600
sudo install -d -m 755 /etc/nm-collector
sudo install -o <user> -g "$(id -gn <user>)" -m 600 \
  /home/<user>/<ASSET>-magmon.json /etc/nm-collector/<ASSET>-magmon.json
sudo cp /home/<user>/magmon-gateway-<ASSET>.service /etc/systemd/system/

sudo systemctl daemon-reload
sudo systemctl enable --now magmon-gateway-<ASSET>

Verify — all four must pass before you call it done:

systemctl status magmon-gateway-<ASSET> --no-pager      # active (running)
journalctl -u magmon-gateway-<ASSET> -n 30 --no-pager   # "reported N sample(s)", no errors
pgrep -af <ASSET>-magmon.json                                     # ONE python line, and you can see it
ls -l /var/tmp/magmon-gateway-<ASSET>.hwm                # exists after the SECOND cycle
Use pgrep -af, not pgrep -c, over SSHRun as ssh host "… pgrep -c -f <ASSET>-magmon.json", the count comes back one too high: the remote shell’s own command line contains the text being searched for, so pgrep matches the question. Typed at a prompt it is fine, which is why this is easy to miss. -af prints what matched, so a real collector (/usr/bin/python3 /opt/magmon-gateway.py /etc/nm-collector/<ASSET>-magmon.json) is obvious next to a bash -c echoing your own command.
The first cycle reports 1 sample. The second is the one to read.A fresh start has no high-water mark, so the collector deliberately sends only the newest row rather than dumping an hour of history. One poll interval later it should report one sample per minute elapsed — five on a five-minute poll — and write the .hwm file. If the second cycle also says reported 1 sample(s) and no .hwm appears, the unit is storing one reading per poll instead of every minute: confirm minute resolution.
'normal fetch failed … trying raw socket fallback' is normalSeveral MagMons answer with a malformed HTTP status line that Python refuses to parse, so the collector retries the same request over a raw socket and carries on. Seeing that line once per cycle is expected on those units and is not a fault.

…and the asset flips to online on the dashboard within a couple of minutes. If any check fails, go to troubleshooting. Once it’s green you can clean up the copies in the home directory:

rm -f /home/<user>/magmon-gateway*.py /home/<user>/magmon-gateway-<ASSET>.service /home/<user>/<ASSET>-magmon.json
systemd, never cronThe collector runs continuously and sleeps between polls on its own. A cron entry would launch an additional copy on every tick while the earlier ones keep running — which is what the lock file complaints in the journal mean. If pgrep -c prints anything but 1, you have a duplicate.
A host still on the old one-script-per-unit layoutBefore 2026-09-24 every unit had its own script at /opt/magmon-gateway-<ASSET>.py with its token baked in. scripts/deploy-collector.sh migrates a host in one run once each unit’s config is staged beside it — see docs/pi-redeploy-guide.md. The new collector picks up the old one’s high-water mark, so nothing is re-sent or skipped.

Part 2

Environmental hardware (sensors + UPS)

Temperature and humidity sensors on an RS-485 bus, and the UPS feeding the trailer. The hardware and the collector are the same wherever they go — what changes is whether they are the whole job or an addition to a magnet. Steps 1 to 3 are shared; Step 4 comes in two versions, one for each case. Everything about the Pi itself is Part 1, so do that first.

What it adds, and to what

This is an additive set of channels, not a different kind of unit. Whichever of these are fitted is what the dashboard draws:

Zone temp / humidityup to three XY-MD02 sensors on one RS-485 busOne is fine. Two or three is a wiring decision, never a code change — only the zones that report get drawn.
Mains / UPS poweron-battery state, battery %, input volts, via NUTThe fleet-wide outage rule starts paging for this unit the moment it reports.
Collectorenv-gateway-<ASSET>.py
Extra packagespython3-pymodbus, nut-server, nut-client

The two cases:

On an MRIStep 4 · MRIAn addition. The asset stays modality MRI and keeps its MagMon address; its Pi runs BOTH collectors, reporting one asset. Helium, bay temperature and mains power land in the same row and show on one card.
On a PET/CT or Nuc Med unitStep 4 · PET/CT · Nuc MedThe whole job. There is no MagMon, so the asset is created with that modality, the device-address fields disappear, and one collector runs on its own.

The collector names differ from the MagMon one on purpose — separate service, path, lock file and log tag — because on a magnet the two run side by side and would otherwise fight over each other’s files.

A blank channel is never a zeroA sensor that fails to answer is reported as no reading, not as 0. That is what lets a dead UPS link show up as a fault instead of quietly reading “wall power is fine” forever. If a zone shows an em-dash on the dashboard, the sensor is not answering — it is not a cold room.

Step 1

Wire the sensors and the UPS

The three sensors share one RS-485 pair back to a USB adapter on the Pi. RS-485 is a bus: every sensor lands on the same two wires (A to A, B to B), daisy-chained from one to the next rather than home-run to the Pi.

  1. Mount one XY-MD02 in each zone: Section 1 Engineering, Section 2 Tech / Patient, Section 3 Equipment.
  2. Daisy-chain A and B between them and back to the USB–RS485 adapter. Keep the pair away from mains runs; it is a differential pair and will tolerate a lot, but not a compressor contactor.
  3. Power the sensors from their supply (they are not bus-powered).
  4. Plug the adapter into the Pi, and the UPS’s USB data cable into the Pi.
  5. The Pi itself plugs into the UPS — a gateway that dies with the mains cannot report the outage.

Confirm the Pi can see both:

ls -l /dev/ttyUSB*        # the RS-485 adapter, usually /dev/ttyUSB0
lsusb                     # the UPS should appear by brand
The service user needs the serial port groupsWithout this every zone reads blank and nothing in the log says why — it is a permission error on the serial port, not a sensor fault. Log out and back in (or reboot) afterwards for it to take effect:
sudo usermod -aG dialout,plugdev gateway
Both groups, because which one owns the device depends on the adapter: ls -l /dev/ttyUSB0 shows root dialout on some and root plugdev on others — both seen in this fleet. And a trailing + on those permissions is an ACL that logind grants the logged-in user — so a hand-run scan works over SSH while the collector, which has no seat, is refused. Testing by hand proves nothing about the service. Check the real thing with getfacl /dev/ttyUSB0 and groups.

Step 2

Address each sensor: 1 = S1, 2 = S2, 3 = S3

This is the step that decides which zone is which, and the one worth slowing down for. The collector reads Modbus unit 1 into S1 Engineering, unit 2 into S2 Tech / Patient and unit 3 into S3 Equipment.

The address is the only identity a sensor hasRS-485 is a multidrop bus, not a chain the Pi can walk. Nothing in the protocol reports position, order or location, so no software can work out where a sensor is. The mapping is correct only if the sensor you install in a zone is the one addressed with that zone’s number. Write the address on each sensor with a marker before it goes on the wall.

They all ship as address 1, so they must be addressed one at a time — two unaddressed sensors on the same bus both answer to 1 and you get garbage. Connect one, set it, unplug it, move to the next.

Set the address with the vendor’s Windows tool, or from the Pi:

# ONE sensor connected. Set NEW to 1, 2 or 3 for the zone it is going into.
python3 - <<'PY'
from pymodbus.client import ModbusSerialClient
NEW = 2   # 1 = Engineering, 2 = Tech / Patient, 3 = Equipment
c = ModbusSerialClient(port="/dev/ttyUSB0", baudrate=9600, bytesize=8,
                       parity="N", stopbits=1, timeout=2)
c.connect()
try:
    r = c.write_register(address=0x0101, value=NEW, device_id=1)
except TypeError:                      # older pymodbus spells it slave=
    r = c.write_register(address=0x0101, value=NEW, slave=1)
print("FAILED" if r.isError() else f"sensor is now address {NEW}")
c.close()
PY
Check the register against your datasheet0x0101 is the device-address register on the common XY-MD02. Variants exist. If the write fails, use the vendor tool rather than guessing at registers — a wrong write can change the baud rate and take the sensor off the bus entirely.

With all three wired up, confirm the bus — this is also how you identify a sensor:

python3 - <<'PY'
from pymodbus.client import ModbusSerialClient
c = ModbusSerialClient(port="/dev/ttyUSB0", baudrate=9600, bytesize=8,
                       parity="N", stopbits=1, timeout=0.4)
print("port open:", c.connect())
found = 0
for uid in range(1, 8):
    try:
        # pymodbus renamed the keyword: >=3.7 wants device_id=, older wants slave=.
        try:
            rr = c.read_input_registers(address=1, count=2, device_id=uid)
        except TypeError:
            rr = c.read_input_registers(address=1, count=2, slave=uid)
    except Exception as e:
        # An address with no sensor on it RAISES rather than returning an error
        # result. Catching it is what lets the scan reach address 2.
        print(f"address {uid}: no answer ({type(e).__name__})")
        continue
    if rr and not rr.isError():
        t, h = rr.registers[0] / 10.0, rr.registers[1] / 10.0
        print(f"address {uid}: {t * 9 / 5 + 32:.1f} F   {h:.1f} %RH   <-- FOUND")
        found += 1
print(f"{found} sensor(s) on the bus")
c.close()
PY

Exactly three lines marked FOUND, at addresses 1, 2 and 3 — a unit fitted with fewer sensors shows correspondingly fewer. Zero found, with the port opening, is a bus problem rather than an addressing one: check the sensor has its own 5–30 VDC supply (it is not bus-powered), swap A and B (the commonest RS-485 fault, and adapter silkscreens disagree with each other), confirm /dev/ttyUSB0 really is the RS-485 adapter with lsusb, and re-run at 4800 and 19200 baud before concluding a sensor is dead. To prove which is which, warm one sensor with your hand for a minute and re-run — the address that moves is that zone. Do this before you close the trailer up.

Step 3

Set up the UPS with NUT

The collector reads the UPS by shelling out to upsc, so NUT has to be answering locally before the collector will report anything about power.

sudo apt-get update
sudo apt-get install -y nut-server nut-client

/etc/nut/ups.conf — the name in brackets matters:

[ups]
    driver = usbhid-ups
    port = auto
    desc = "Trailer UPS"
It must be called upsThe generated collector runs upsc ups. If you name the section anything else, every power field reports blank and the dashboard shows Power unknown — which reads as a broken link, not as a naming mistake. Either call it ups or change UPS_NAME at the top of the script.

The remaining three files:

# /etc/nut/nut.conf
MODE=standalone

# /etc/nut/upsd.users
[monuser]
    password = <pick-a-password>
    upsmon master

# /etc/nut/upsmon.conf  (append)
MONITOR ups@localhost 1 monuser <pick-a-password> master
sudo systemctl restart nut-server nut-monitor
sudo systemctl enable nut-server nut-monitor
upsc ups

upsc ups should print a block of values. The three the collector uses are ups.status (OL on line, OB on battery), battery.charge and input.voltage. Pull the UPS’s mains plug for a few seconds and watch ups.status flip to OB — that is the whole alarm path proven at the source.

Step 4 · MRI

Install it alongside a MagMon collector

A UPS or a bay sensor on an MRI unit is an addition, not a different kind of unit. The asset stays on modality MRI with its MagMon address intact, and its Pi runs both collectors — two scripts, two services, two lock files, one asset.

Nothing about the asset changes — no new asset, no modality edit, no second tile on the dashboard.

  1. Wire and address only the sensors actually being fitted (Steps 1–3 above). One sensor gives one tile; the collector polls all three addresses regardless, so a second one added later is a wiring job with no redeploy.
  2. Admin → the asset → Get install script, then set Collector to Environmental. Set the poll interval to 1 minute — see below for why that is free on a magnet — and download the script, the config and the systemd unit.
  3. Install them exactly as in Step 6, substituting the environmental names: the shared environmental script (once per host), this unit’s -env.json config, and its unit. Nothing about either install changes because the other one is there:
    sudo install -o root -g root -m 755 /home/<user>/env-gateway.py /opt/env-gateway.py
    sudo install -o <user> -g "$(id -gn <user>)" -m 600 \
      /home/<user>/<ASSET>-env.json /etc/nm-collector/<ASSET>-env.json
    sudo cp /home/<user>/env-gateway-<ASSET>.service /etc/systemd/system/
    sudo systemctl daemon-reload
    sudo systemctl enable --now env-gateway-<ASSET>
  4. Watch a cycle. You want one line per fitted zone, one UPS line, and reported N channel(s):
    journalctl -u env-gateway-<ASSET> -f
Poll environmental channels every minute on a magnetIt costs nothing here. The MagMon collector already writes a row for every minute, and the environmental readings merge into those same rows — so a one-minute poll adds no rows at all, it only fills columns. And it matters for power: the alert crons run every minute, so a five-minute poll can burn a third of a UPS’s runtime before anyone is paged. On a standalone unit the trade is different — there, every poll is a new row.
"No sensor answered" on unfitted zones is expectedAddresses with nothing on them log one line per cycle for the first three cycles, then drop to being retried every tenth cycle so they stop costing a Modbus timeout on every poll. A zone you did fit showing that line is the fault — check its address first.
Two collectors on this box is correctThe Part 1 check pgrep -c -f magmon-gateway counts MagMon collectors only, but a bare pgrep -c -f gateway prints 2 on a mixed unit. That is the healthy state, not a doubled collector. Check the specific names:
pgrep -c -f <ASSET>-magmon.json      # 1
pgrep -c -f <ASSET>-env.json         # 1
Both collectors write the same minuteEach one merges its own channels into that minute’s row and leaves the other’s alone, so the card shows helium, bay temperature and mains power together rather than one blanking the other. Nothing to configure.
Half a unit can go quiet, and the app says which halfBecause either collector keeps the asset reporting, “online” on a mixed unit does not mean both halves are alive. Each side keeps its own clock, so a dark one gets a red chip on the card — MagMon not connected if it has never reported, MagMon silent if it was reporting and stopped, Env silent for the reverse — and raises one alert naming the collector, rather than six sensor-fault events for the channels it carried. Expect the “not connected” chip in the window between installing the environmental collector and the MagMon one; it clears itself the moment magnet telemetry lands.

Step 4 · PET/CT · Nuc Med

Create the asset and install the collector

A unit with no MagMon — a PET/CT or Nuc Med trailer or lab — is an ordinary asset whose only channels are environmental. One collector, no device address, everything else identical to Part 1.

  1. Create the asset as in Step 3, setting Modality to PET/CT or Nuc Med. The MagMon address and credential fields disappear — there is no device to reach. A stale threshold of about 20 minutes suits a 5-minute poll.
  2. Get install script hands back the environmental collector automatically; there is no Collector switch, because this unit can only run the one.
  3. Set the poll interval. Unlike a magnet, every poll here is a new row — five minutes is the sensible default, and one minute is a five-fold increase in stored history for readings that move slowly.
  4. Download the script and the systemd unit.

On the Pi — dependencies first:

sudo apt-get install -y python3-requests python3-pymodbus nut-server nut-client
# if python3-pymodbus is not packaged for your release:
#   sudo pip3 install --break-system-packages pymodbus
sudo usermod -aG dialout,plugdev <user>   # then log out and back in

Copy the three files up (same SCP as Part 1), then install:

# once per host -- skip if /opt/env-gateway.py is already there
sudo install -o root -g root -m 755 /home/<user>/env-gateway.py /opt/env-gateway.py

# per unit -- the config holds its gateway token, so 600
sudo install -d -m 755 /etc/nm-collector
sudo install -o <user> -g "$(id -gn <user>)" -m 600 \
  /home/<user>/<ASSET>-env.json /etc/nm-collector/<ASSET>-env.json

sudo cp /home/<user>/env-gateway-<ASSET>.service \
        /etc/systemd/system/env-gateway-<ASSET>.service

sudo systemctl daemon-reload
sudo systemctl enable --now env-gateway-<ASSET>

Verify — all four before you call it done:

systemctl status env-gateway-<ASSET> --no-pager   # active (running)
journalctl -u env-gateway-<ASSET> -n 30 --no-pager # three zones + a UPS line
pgrep -c -f <ASSET>-env.json                                     # prints exactly 1
upsc ups | grep ups.status                                       # OL

Then check the dashboard: the asset shows a tile per fitted zone with readings, a green Wall power chip, and a battery percentage. If a zone you fitted reads an em-dash, go back to Step 2 and re-run the bus scan.

systemd, never cronSame rule as the MagMon collector, same reason: this script never exits, so a cron entry starts a fresh copy on every tick while the old ones keep running. If pgrep -c prints anything but 1, you have duplicates.

Alert rules for temperature and power

Power is already covered fleet-wide. A rule of On UPS battery = 1 raises a critical POWER OUTAGE for any asset that reports the channel, so a new trailer is covered the moment it starts reporting — nothing to add per unit.

Temperature and humidity limits are deliberately unset. Add them in Admin → Alerts once the right numbers are known for the space:

  1. Pick the channel under the Environment group — each zone has its own temperature and humidity.
  2. Leave Scope on All assets for a fleet-wide limit, or pick one asset to override the fleet default for just that unit.
  3. Set the comparator and the threshold, and save.
A fleet-wide environmental rule is safe on a magnetA rule only applies where the reading exists. A channel an asset does not report is skipped rather than treated as zero, so adding a zone-temperature rule across all assets does nothing to units without sensors — and starts working by itself on the day they are fitted.
Blank for an hour is its own alarmA channel that reads blank on every sample for an hour, while the unit is otherwise reporting, raises a sensor fault naming that channel. That covers a sensor that drops off the bus and the UPS link going down — no rule needed, and it is why a missing reading must never be filled in with a zero.

Part 3

Site and network setup

Two different questions, and it is worth keeping them apart: how the Pi reaches the MagMon (a decision made at install time, and the one that decides the device address you put on the asset), and how you reach the Pi afterwards. Cellular sites come in behind an iR305; everything on the tailnet is reachable directly over Tailscale from anywhere.

How the Pi reaches the MagMon

The collector talks to the device over the LAN, so the address you enter as the MagMon address on the asset is whatever the Pi can reach — not what your laptop reaches, and not what a VPN server reaches. Two shapes cover the fleet.

A. Both on the router’s LAN — the simple one

The MagMon and the Pi each plug into the site router (or a small switch behind it) and get addresses on the same subnet. Nothing to configure on the Pi; the device address is its LAN IP. Use this whenever there is a spare LAN port. It also means the MagMon’s own web interface is reachable from anything else on that LAN.

B. Pi on Wi-Fi, MagMon on a private cable — when the device stays off the site network

The Pi uplinks over the router’s Wi-Fi and its single Ethernet port runs straight to the MagMon. The device is then reachable only from its own Pi, which is a real security gain and keeps a hospital network out of the picture entirely. Two things have to be true or it silently does not work:

  1. Bake the Wi-Fi credentials in when you flash the card (Part 1, Step 1). With Ethernet committed to the device, Wi-Fi is the only uplink — a Pi that cannot join it is a Pi you cannot reach.
  2. The Wi-Fi subnet and the device subnet must differ. Two interfaces on one subnet makes routing ambiguous. Check the router’s LAN range before you start.

A direct cable has no DHCP server, so give eth0 a static address on the device’s subnet. Keep the MagMon’s existing address if it has one — matching the Pi to the device is far easier than reconfiguring the device, and a device that kept its address through a router swap needs nothing done to it at all. ipv4.never-default is the load-bearing flag: it keeps the default route on Wi-Fi, so reporting and Tailscale are unaffected.

sudo nmcli con add type ethernet ifname eth0 con-name magmon-link \
  ipv4.method manual ipv4.addresses 10.0.0.10/24,169.254.7.10/16 \
  ipv4.never-default yes ipv6.method disabled connection.autoconnect-priority 100

sudo nmcli con up magmon-link

The second, link-local address costs nothing and covers the case where the device was on DHCP from an old router and has fallen back to 169.254.x.x. Then confirm the routing and the device, in that order:

ip -4 addr show eth0        # the static address is on the interface
ip route | head -3          # default MUST be via wlan0, not eth0
ping -c3 -W2 <magmon-host>
curl -sS -m 8 -o /dev/null -w 'HTTP %{http_code}\n' http://<magmon-host>/

Any HTTP status — 302 to the login page is the usual one — means the collector will be able to talk to it.

If you don't know the device's addressListen to what it says on the wire rather than guessing subnets. With the cable in and the interface up, its own broadcasts give it away:
sudo apt-get install -y tcpdump
sudo tcpdump -ni eth0 -c 20      # ARP/DHCP chatter reveals its IP or subnet
ip neigh show dev eth0           # anything that has answered
The device’s front panel usually shows it too.
One profile, not threenmcli con add does not replace a profile of the same name, it adds another one — run it twice and a reboot picks one at random. If you have repeated it, clear them out and add one:
for u in $(nmcli -t -f UUID,NAME con show | awk -F: '$2=="magmon-link"{print $1}'); do sudo nmcli con delete "$u"; done
Reaching the device yourself, on shape BOnly the Pi can see it, so browse it through an SSH tunnel to that named Pi — see open a unit’s MagMon. That is the recommended route on shape A as well, for a different reason: several sites reuse the same private address.

Remote access via the iR305 (InHand Device Manager)

Cellular sites reach the internet through an InHand iR305 router. Its cloud console — the InHand IoT Device Manager — can open a secure tunnel to the Raspberry Pi sitting on the router’s LAN, handing you a temporary public host and port that forwards straight to the Pi’s SSH (port 22).

Console loginInHand Device Manager: iot.inhandnetworks.com/dashboard — sign in as it@example-imaging.com.
  1. Sign in and open the iR305 for the site; confirm it shows online.
  2. Open its remote-access / tunnel tool and create a tunnel to the Pi’s LAN address on port 22 (the Pi’s IP on the router’s network — check the router’s connected-devices list if you’re unsure of it).
  3. Copy the public host and port the tunnel issues.
  4. SSH through it from your laptop:
ssh <user>@<public-host> -p <public-port>
Tunnel to the Pi, not the routerPoint the tunnel at the Pi’s LAN IP:22 — not the router itself. Menu wording can vary slightly by firmware, but the mechanism is always the same: the tunnel gives you a host:port, and you ssh to it. Once a site is also on Tailscale (below), you can skip the tunnel entirely.
An offline router is not an offline magnetIf the iR305 itself drops, the dashboard says so separately from the asset. That means the site’s connectivity is down and telemetry has nowhere to go — the magnet underneath may be perfectly fine. Restore the router first, then re-read the asset.

SSH into Tailscale assets

From any machine signed in to the same tailnet, SSH straight to a device’s tailnet IP (or its MagicDNS name):

ssh <user>@100.x.y.z
# or, with MagicDNS:
ssh <user>@<hostname>

The central Pi server is reachable two ways — on the operations LAN, or over Tailscale from anywhere:

Pi server — LANssh <user>@10.0.0.10
Pi server — Tailscalessh <user>@100.100.10.20

List every machine on the tailnet and its IP with tailscale status from any joined device.

Open a unit's MagMon web interface

Do not browse the MagMon’s LAN address directly. That address is private and several sites reuse it, so over a tailnet subnet route it reaches whichever Pi currently owns that route — with nothing on screen to tell you which unit you got. This is not hypothetical: a week of minute data was read from one unit while believed to be another, and only the readings themselves gave it away.

Tunnel through a named Pi instead. The Pi is at exactly one site, so the tunnel can only land on one MagMon. Pick a different local port per unit and you can hold several open at once:

ssh -L 8080:<magmon-host>:80 <user>@<pi-host>
# then open http://localhost:8080

Ready-made per unit. A shared address is one another unit also uses — for those the tunnel is the only safe route.

MM-1001Example Regionalssh -L 8080:192.168.1.50:80 gateway@mm1001-pi.example.ts.netthen http://localhost:8080
MM-1002Example Communityshared addressssh -L 8081:192.168.1.99:80 gateway@mm1002-pi.example.ts.netthen http://localhost:8081
MM-1003Example Imagingshared addressssh -L 8082:192.168.1.99:80 gateway@mm1003-pi.example.ts.netthen http://localhost:8082

To see which Pi currently owns a route from your own machine: tailscale status --json | jq -r '.Peer[] | select(.PrimaryRoutes) | "\(.HostName)\t\(.PrimaryRoutes | join(","))"'

Replace a collector script on a running Pi

Every host runs the same collector script; what makes a collector a particular unit’s is its config file in /etc/nm-collector. So a change to the collector code is one file, copied once to each of the 10 hosts — the Pi server’s five units pick it up together. The collector version shown against each asset in Admin tells you which are still behind. scripts/deploy-collector.sh does all of this per host; the steps below are the same thing by hand.

This is a script swap and a restart, not a reinstall. Configs and units do not change, so there is no daemon-reload and no re-enable. Do the units in maintenance first and live hospital sites last.

  1. In Admin → Assets, click Get install script on any asset and download the collector script. It is the same file for every unit and lands in ~/Downloads as magmon-gateway.py.
  2. Copy it to the host’s home directory (one password prompt).
  3. Back up the running script, install the new one over it, and restart every collector service on that host.
  4. Verify three things: the service is active, the journal shows the new version at startup, and within a few minutes the asset’s collector version in Admin matches.

Copy up, then install and restart:

scp ~/Downloads/magmon-gateway.py <user>@<pi-host>:/home/<user>/

ssh <user>@<pi-host> 'sudo cp /opt/magmon-gateway.py /opt/magmon-gateway.py.bak; \
  sudo install -o root -g root -m 755 /home/<user>/magmon-gateway.py /opt/magmon-gateway.py && \
  sudo systemctl restart "magmon-gateway-*" && sleep 20 && \
  systemctl list-units "magmon-gateway-*" --no-pager && \
  journalctl -u "magmon-gateway-*" -n 10 --no-pager'

A healthy result is active followed by a startup line carrying the version, e.g. starting for asset ‘MM-1002’ v2026.09.24-2. Restarting makes the collector run a cycle immediately, so one extra reading lands out of step with the five-minute rhythm — that is the restart, not a fault.

If it misbehaves, roll back to the script that was running a moment ago:

ssh <user>@<pi-host> 'sudo cp /opt/magmon-gateway.py.bak /opt/magmon-gateway.py && \
  sudo systemctl restart "magmon-gateway-*"'

The <user>@<pi-host> for each unit. Units with no Tailscale node of their own run on the shared Pi server — use its address for those.

MM-1001Example Regionalgateway@mm1001-pi.example.ts.net
MM-1002Example Communitygateway@mm1002-pi.example.ts.net
MM-1003Example Imaginggateway@mm1003-pi.example.ts.net

The shared Pi server (gw-server)

Some sites don’t get their own Pi — they’re reached over a VPN from one central box instead. The Pi server — hostname gw-server, 10.0.0.10 on the LAN or 100.100.10.20 over Tailscale — runs one collector per asset, each installed exactly as in Part 1. Every asset is a matching pair owned by the gateway user: a .py collector and its .service unit, both named magmon-gateway-<ASSET>.

Collector services

magmon-gateway-MM-1001Reports telemetry for MM-1001
magmon-gateway-MM-1002Reports telemetry for MM-1002
magmon-gateway-MM-1003Reports telemetry for MM-1003
magmon-gateway-MM-1004Reports telemetry for MM-1004
magmon-gateway-MM-1005Reports telemetry for MM-1005
magmon-gateway-MM-1006Reports telemetry for MM-1006

SSH in and manage a collector by its service name:

ssh gateway@100.100.10.20              # or 10.0.0.10 on the LAN
ls -la /etc/nm-collector              # one config per unit
systemctl status magmon-gateway-MM-1002
journalctl -u magmon-gateway-MM-1002 -f     # live logs
sudo systemctl restart magmon-gateway-MM-1002

Check the whole box at once:

systemctl list-units 'magmon-gateway-*' --no-pager   # one per asset, all running
pgrep -c -f magmon-gateway                            # should equal the asset count

Take an installed collector out of service — stop it now and keep it from starting on boot (--now does both):

sudo systemctl disable --now magmon-gateway-<ASSET>
systemctl status magmon-gateway-<ASSET>     # confirm: inactive (dead), disabled
pgrep -af <ASSET>-magmon.json                     # no output = nothing left running

Re-enable it later with sudo systemctl enable --now magmon-gateway-<ASSET>. The service is named magmon-gateway-<ASSET>, while every unit’s process is the same /opt/magmon-gateway.py with its own config — so match on <ASSET>-magmon.json for one unit, or magmon-gateway for all of them.

Confirm routing before adding sites to one boxEach MagMon must be reachable from this server at a distinct address (the monitor host set on the asset). If two sites reuse the same private subnet behind the tunnel they will collide — give them distinct routes or NAT first.

Part 4

Run and troubleshoot

What to do when an asset goes red. Work the triage table first — it splits the two failure families that look identical on the dashboard (the site is unreachable vs the site is reachable but nothing is arriving) before you start changing things.

Troubleshooting a gateway

“Offline” on the dashboard means no fresh telemetry — nothing more. A dead magnet, a dead Pi, a dead router and a Pi that simply can’t reach the internet all produce the same red chip. Start by finding out which one you have.

Triage — run these three, in this order:

# 1. Can you reach the Pi at all?  (from your laptop, on the tailnet)
ssh <user>@<tailnet-ip>

# 2. Is the collector running, and is it saying anything?
systemctl status magmon-gateway-<ASSET> --no-pager
journalctl -u magmon-gateway-<ASSET> -n 50 --no-pager

# 3. Can the Pi reach the outside world, by name?
ping -c2 8.8.8.8                       # raw-IP internet
getent hosts example-project.supabase.co   # the DNS lookup that actually matters

What the answers mean

SSH fails, Tailscale shows the node downThe Pi or the site link is down — power, network, or router. Start at the iR305.
SSH works, service inactive/failedCollector problem. See the install-failure table below.
SSH works, service running, ping OK, getent FAILSPi-local DNS failure. This is the classic one — see below.
Everything green, journal says 'reported OK'Telemetry is flowing. If values look wrong, judge the magnet itself.
Card shows 'MagMon not connected' or 'MagMon silent'A unit with two collectors has one down. The other keeps it reporting, so the status chip stays green — check that specific service, not the asset.
Card shows 'Env silent'Same, the other way round: the sensors or UPS link stopped while the MagMon side keeps reporting.
Every zone blank on a unit with sensorsAlmost always serial permissions, not sensors — see the dialout/plugdev note in Part 2.
Offline while the Pi looks perfectly healthy = suspect DNSSSH, Tailscale and a direct curl to the MagMon all use a raw IP or Tailscale’s own path — none of them need DNS. So they all pass while the one DNS-dependent call, reporting to example-project.supabase.co, fails silently. The journal tell is socket.gaierror [Errno -3] Temporary failure in name resolution. Confirm with the getent line above; if ping 8.8.8.8 works and the lookup doesn’t, it’s DNS only. Fix by pointing the Pi at a public resolver, then restarting the collector:
sudo nano /etc/resolv.conf        # nameserver 1.1.1.1 / 8.8.8.8
sudo systemctl restart magmon-gateway-<ASSET>

# make it survive a reboot:
# add to /etc/dhcpcd.conf:
#   static domain_name_servers=1.1.1.1 8.8.8.8
If only one asset gapped while the others kept reporting, it is that Pi — not the backend.
Running but the journal is empty — that's buffering, not a hangIf pgrep -c -f magmon-gateway-<ASSET> prints 1 but journalctl shows nothing recent, the process is usually fine — Python is block-buffering its output. Current units set Environment=PYTHONUNBUFFERED=1; units deployed before that don’t. Trust systemctl status and CPU time over a live journalctl -f, and fix it by re-downloadingboth files from Admin and redoing step 6. Units must live in /etc/systemd/system/ — confirm which file is actually loaded with systemctl cat magmon-gateway-<ASSET> | grep FragmentPath.
Reporting, but the data never moves onThese controllers have no battery-backed clock, so a power blip reboots the MagMon and snaps its clock back to 13-May-06. The collector only sends rows newer than the last one it sent, so every row now looks older than the stored high-water mark and the feed goes quiet while the service looks perfectly healthy. Clear the mark to recover:
sudo systemctl stop magmon-gateway-<ASSET>
sudo rm -f /var/tmp/magmon-gateway-<ASSET>.hwm
sudo systemctl start magmon-gateway-<ASSET>
Current scripts detect the reset and re-anchor themselves, so this is only needed on a Pi that hasn’t been redeployed since. A 2006 date on its own is normal for these units and is not evidence of anything — rule out DNS first.

Install-time failures (fresh deploys):

Symptom → fix

ModuleNotFoundError: 'requests'python3-requests was never installed. Install it, then restart the service.
Service state 217/USERThe Service user you set in Admin doesn't exist here. Check whoami, re-download, redo step 6.
Permission denied on the .pyWrong owner or mode — re-run the sudo install line in step 6 exactly.
'another copy already holds the lock'A second copy really is running. pgrep -af, kill the extras, remove any cron entry.
'Cannot reach MagMon at <ip>:<port>'Wrong or unroutable MagMon address for this site. Fix it in Admin → Edit, re-download, redo step 6.
Still offline after a few minutesToken mismatch — most likely rotated. Re-download the current script and redo step 6.
"Compressor powered off or cable disconnected" — the coldhead tells you whichThat status code covers two different jobs, and the cryogenics decide between them. A compressor genuinely stopped for weeks cannot leave a coldhead at 4 K, so a cold coldhead beside that alarm means the MagMon has lost its 24 V sense line and the cryocooler is still running — check the cable before you order a compressor. A coldhead up at 50 K or 300 K beside the same code means it really has stopped. The alert now says which of the two it is; this is the reasoning behind it.
Falling helium with a cold coldhead is not boil-offSame idea, the other way round. If the coldhead is at 4 K the cryocooler is working, so helium leaving the vessel is going out somewhere else — a relief valve or a seal, not a cryogenic failure. A warming coldhead alongside falling helium is the boil-off case. Worth knowing how small the normal signal is: across this fleet a healthy magnet moves less than 0.1 % of helium in three weeks, so any sustained daily loss is a large departure long before the level looks low.
Telling a real magnet fault from a collection bugWhen the numbers look alarming, cross-check the two independent read paths — the HTTP scrape and the FTP file use completely different parsers. If both agree to the decimal, the numbers are the device’s true state, not misalignment. Then judge health from the cryo diodes, not the helium level: a cold, running magnet reads coldhead and recon around 4 K with the shield near 40 K, while a warm or powered-down one reads them all at 300 K+ with helium at 0 and no water flow. An all-zero reading from a genuinely warm unit looks exactly like a parser bug — this is how you tell.
Leave legacy collectors aloneSome older Pis still run pre-app collectors under cron alongside the gateway service (older magmon_collect_* scripts and their log files). They are unrelated to this app, but other reporting may still depend on them — don’t disable, delete, or clean up those cron entries as part of gateway work. If one is filling the disk, ask before touching it, and archive anything you remove: several of those scripts exist on the Pi and nowhere else.

Confirm a collector is storing every minute

A collector can be perfectly healthy and still be storing a twelfth of what the device is serving. The MagMons log a row a minute; the collector fetches the last hour each cycle and files every row newer than the last one it sent. If it cannot read the device’s timestamps it degrades quietly instead of failing: it files just the newest row, stamped with the current time. The unit reports on schedule, the charts draw, and the resolution is gone. It ran that way fleet-wide from August 2026 until it was noticed on an install in September.

The tell, in the journal:

journalctl -u magmon-gateway-<ASSET> -n 20 --no-pager | grep 'reported'
Healthyreported 5 sample(s) — on a 5-minute pollOne per minute elapsed. The first cycle after any restart reports 1 by design.
Degradedreported 1 sample(s) — every cycle, foreverAnd no .hwm file, which is only written when a row’s timestamp parses.
ls -l /var/tmp/magmon-gateway-<ASSET>.hwm   # missing = nothing has ever dated

The same thing is visible from the dashboard side without SSH: a degraded unit stores exactly 60 ÷ poll-minutes rows an hour — a flat 12/hour on a five-minute poll, never more.

The fix is a script redeploy, not a device change:

  1. Check the asset’s collector version in Admin. Anything older than the current generator is suspect, and the panel marks it “behind”.
  2. Redeploy it — replace a collector script. Script swap and restart; the systemd unit does not change.
  3. Watch the second cycle after the restart. Five samples and a .hwm file means it is filing every minute.
Confirming the device really is logging every minutePull the minute table by hand and look at the spacing. One-minute steps mean the data is there and the collector is the problem; five-minute steps mean the device itself is configured to log at that interval, and no collector change will help:
curl -s --http0.9 --max-time 25 -u <user>:<pass> \
  "http://<magmon-host>/goform/showMinutesCGI?num_hours=1&start_day=0&start_hour=1" \
  | grep -oE '[0-9]{2}:[0-9]{2}' | tail -15
Expect roughly five times the rows per unitRaw samples are purged on a rolling window, so the database settles at a new steady state a week after a unit is converted rather than growing without bound — but convert the fleet deliberately and watch the size as you go. Daily rollups carry the long history, so shortening the raw window is the first lever if it gets tight.

Everyday Raspberry Pi commands

A quick cheat sheet for working on a Pi over SSH.

Getting around

pwdPrint the folder you're currently in
ls -laList everything, including hidden files, with details
cd /path/to/dirChange directory
cd ~Go to your home directory
cd ..Go up one level

Files

cat filePrint a file to the screen
less fileScroll through a long file (q to quit)
nano fileEdit a file (Ctrl+O save, Ctrl+X exit)
cp a bCopy a to b
mv a bMove or rename a to b
rm fileDelete a file (no undo)
mkdir dirMake a new folder
chmod +x script.pyMake a script executable
du -sh *What's using the space in this folder

System

sudo apt update && sudo apt full-upgrade -yUpdate all packages
df -hDisk space, human-readable
free -hMemory in use
htopLive processes + CPU (q to quit)
uptimeHow long it's been up, load average
dateThe Pi's clock — worth checking during odd faults
sudo rebootRestart the Pi
sudo shutdown nowPower it off

Services (systemd)

systemctl status <name>Is the service running?
sudo systemctl restart <name>Restart it
sudo systemctl enable --now <name>Start now + on every boot
sudo systemctl disable --now <name>Stop now + never start on boot
journalctl -u <name> -fFollow its live logs
journalctl -u <name> -n 50 --no-pagerLast 50 lines, no pager
systemctl cat <name>Show the unit file that's actually loaded

Processes

pgrep -af <pattern>Every matching process, with its full command
pgrep -c -f <pattern>Just the count — 1 is what you want per collector
sudo kill <pid>Stop a stray process

Scheduling (cron)

crontab -eEdit your scheduled jobs
crontab -lList your scheduled jobs
* * * * * commandmin hour day month weekday — then the command
MagMon runs under systemd, not cronThe gateway collector is a long-running service managed by systemd — check it with systemctl status magmon-gateway-<ASSET>, not crontab. Use cron for one-off periodic maintenance tasks.

Networking

hostname -IThis Pi's IP address(es)
ip aAll network interfaces in detail
ping 8.8.8.8Test raw-IP connectivity (Ctrl+C to stop)
getent hosts <hostname>Test DNS specifically — resolves a name to an IP
curl -s http://<magmon-ip>Read the MagMon directly, bypassing the collector
tailscale statusEvery device on the tailnet + its IP
tailscale ip -4This device's tailnet IP