How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
The Oversight Linux agent is a single small program that runs on a Linux host, watches it continuously and decides for itself whether anything is wrong. It serves that verdict over HTTP: nothing wrong, or what is wrong and how badly. Oversight reads it with a Linux agent sensor.
The agent does the judging on the host because it knows the host far better than a threshold in a template can: whether a filesystem was writable an hour ago, which network ports have ever had a link, whether a service is quietly crash-looping. So a sensor on it needs two rules, not two hundred. It is read only by design, runs as an unprivileged user, has no way to run commands, and nothing it serves is a secret or a process list.
It is for bare-metal servers and ordinary virtual machines. Proxmox hosts, guests and containers are better watched through their own APIs.
Installing it on the host
As root on the host to be watched:
curl -fsSL https://oversight.im/linuxsensor | bash
or, from a user with sudo, curl -fsSL https://oversight.im/linuxsensor | sudo bash. The script picks the build for the machine (x86_64, arm64 or 32-bit ARM), checks it against its published checksum, and installs /opt/oversight/agent, an oversight system user, /etc/oversight/agent.json and the oversight-agent service.
It then asks one question, and that is the whole setup:
Address for the agent to listen on, which the probe connects to [0.0.0.0:29731]:
Press Enter for every interface, or give the address the probe will reach it on, such as 10.1.1.100. An address with no port takes the default one. It generates a bearer token, prints it, and starts the service:
/etc/oversight/agent.json: listen 0.0.0.0:29731, 1 token(s), from any address
token: 7pQz3XfKdR2mVn8sLtYwJhB4
oversight-agent enabled and running, listening on 0.0.0.0:29731
Keep that token. It goes in a Linux agent credential in Oversight, below. It is in /etc/oversight/agent.json as well, so it can be read back off the host later.
Running the same line again is how it is upgraded. An upgrade asks nothing and changes nothing it was not told to: the config, and a service file already there, are left as they are.
Where a script installs it rather than a person, the answers go on the command line and nothing is asked:
curl -fsSL https://oversight.im/linuxsensor | sudo bash -s -- --listen 10.1.1.100:29731 --token a-long-random-token --allow 10.1.1.25
Check it from the host before going near Oversight:
curl -s -H "Authorization: Bearer a-long-random-token" http://10.1.1.100:29731/v1/status
The host's firewall needs a rule. Nothing in the install opens the port: firewalld, ufw, nftables or whatever this host runs has to allow TCP 29731 in from the probe. Until it does, the curl above works from the host itself while the sensor sees nothing.
It speaks plain HTTP. Where it must be reached across the internet, put a reverse proxy such as nginx in front of it for TLS, set allow to the proxy's address and let the proxy decide who reaches it.
Changing it afterwards
Everything the agent is told is in /etc/oversight/agent.json, which only root and the oversight user can read, because it holds the token:
{
"listen": "10.1.1.100:29731",
"tokens": ["a-long-random-token"],
"allow": ["10.1.1.25"],
"checks": {},
"ignore": []
}
| Key | What it does |
|---|---|
listen | The one address and port it answers on. 0.0.0.0:29731 means every interface. |
tokens | Bearer tokens it accepts. The installer generates the first one. Two may be listed at once, so a token can be changed without a gap. |
allow | The addresses, or ranges such as 10.1.1.0/24, allowed to connect. Put your probe's address here. |
checks | Turns a check off, or changes the limits it alarms at, on this host only. See Turning a check off, or changing its limits, below. |
ignore | Things no check alarms on: a mount point, disk, interface, service or port. Also below. |
None of it has to be edited by hand. Each of these adds to the file and restarts the service:
/opt/oversight/agent --listen 10.1.1.100:29731
/opt/oversight/agent --token another-long-random-token
/opt/oversight/agent --allow 10.1.1.25,10.1.2.0/24
A token already listed is not added twice, and none is ever taken away, so a token can be changed without a gap. The agent refuses to listen beyond the host with neither tokens nor allow set, and says so, which is why the installer generates a token; setting allow as well is better still. After editing the file by hand, systemctl restart oversight-agent.
Adding the sensor
On the device that is this host, add a sensor of type Linux agent. If the agent has tokens set, add a credential of type Linux agent on the Credentials panel of that device, or of a group above it to share one token across many hosts, with the token in the Bearer token field. The sensor sends it for you; there is no header to write.
| Setting | What to put |
|---|---|
| Endpoint | What to read. Status is the one that raises alarms; the others are figures for charts. See below. |
| Agent address | http and 29731, unless it was installed with another port, or https and the proxy's port where one is in front of it. |
| Interval | 60 seconds for Status. The agent samples every 10 seconds whether anyone asks or not, so polling faster gains nothing. 300 seconds is plenty for a figures endpoint. |
A new Linux agent sensor is ready to use as soon as it is saved. It starts on the Status endpoint with the extraction rows and rules below already filled in. Changing the endpoint afterwards does not change the extraction rows, because each endpoint carries different fields: replace them with the ones listed for that endpoint.
Every endpoint also fills SV20, HTTP status, by itself. It is 200 when the agent answered, 401 for a missing or wrong token and 403 for a probe outside allow. Because it always reads, a sensor on any endpoint stays readable even when the fields it extracts are absent.
Status: the verdict
/v1/status is quiet while all is well:
{"alarm":"green","alarms":0,"status":"","storage":"","filesystem":"", ... ,"unavailable":[]}
and when something is wrong, says what, area by area, worst first:
{"alarm":"red","alarms":2,
"status":"md0 degraded, 1 of 2 members missing; php-fpm.service restarted 7 times in 15 min",
"storage":"md0 degraded, 1 of 2 members missing",
"services":"php-fpm.service restarted 7 times in 15 min", ...}
These are the rows a new sensor starts with:
| Slot | Name | Method | Expression | Holds |
|---|---|---|---|---|
| SD01 | Alarm | JSON | alarm | green, amber or red |
| SD02 | Status | JSON | status | Every current problem, for the alert text |
| SD03 | Storage | JSON | storage | Arrays, disks, iSCSI, multipath, NFS, LVM |
| SD04 | Filesystem | JSON | filesystem | Missing, read-only, unresponsive, full, filling |
| SD05 | Memory | JSON | memory | OOM kills, pressure, swapping, low memory |
| SD06 | Processes | JSON | processes | Stuck (D state), zombies, running out of PIDs |
| SD07 | Services | JSON | services | Failed units, crash loops, ports that stopped listening. A user's own login session (user@1000.service) is not counted: on a desktop install it often fails its own stop on logout, which is not a service down |
| SD08 | Network | JSON | network | Links down, bonds degraded, errors, retransmissions |
| SD09 | CPU | JSON | cpu | Sustained load, steal |
| SD10 | Host | JSON | host_ | Clock, file handles, temperature, a recent reboot |
| SV01 | Alarms | JSON | alarms | How many problems there are now, charted |
| SV02 | Uptime | JSON | uptime_s | Seconds since the host booted |
| SV03 to SV05 | Load 1, 5 and 15 min | JSON | load1, load5, load15 | Load averages, charted together |
| SV06 | Agent uptime | JSON | agent_uptime_s | Seconds since the agent started |
An area with nothing wrong is an empty string. Note the trailing underscore on host_: plain host is the host's name.
And the rules, which are the whole of the judging Oversight needs to do:
| Slot | Rule | State | Why |
|---|---|---|---|
| SV20, HTTP status | not equal to 200 | CRIT | The agent refused the probe. A token or allow problem should not look like a healthy host. |
| SD01, Alarm | equal to red | CRIT | Something is broken now. |
| SD01, Alarm | equal to amber | WARN | Something is heading that way, or needs a look. |
Actions then work as for any other sensor: bind a notification rule to the device or a group above it. The alert's detail is the agent's own words, such as php-fpm.service failed since 14:02, so the message says what is wrong, not just that the alarm slot read red.
Treating one area differently. A rule can be written against any area slot. To make any storage problem CRIT whatever its colour, add SD03 matches /./ as CRIT: an area is empty when it has nothing to say, so anything in it at all matches. To send storage problems to a different team from everything else, give the host a second Linux agent sensor on Status with only that rule, and bind the storage team's notification rule to it. Rules set a state; who is told follows the sensor.
Checks that cannot run on this host are listed in unavailable, for example the pressure checks on a RHEL or AlmaLinux kernel booted without psi=1. That is not an alarm, but it can be counted: a spare slot with method JSONCOUNT and expression unavailable reads how many there are.
Turning a check off, or changing its limits
The defaults are meant to be right for nearly every host, but not every host is ordinary. A backup server spends hours with its disks flat out, and the agent will rightly say so. The fix belongs on that host, in /etc/oversight/agent.json, rather than in a rule in Oversight that hides a whole area. It can only be done on the host: nothing in Oversight changes the agent's settings.
First find the check. The alert's detail is the agent's own words, and the table below gives the check each message comes from: every task stalled on I/O 10% of the last minute, for 22 min is io_stall. Then, under checks, either switch it off:
"checks": {"io_stall": false}
or keep it and raise its limits, so that a normal night's work passes and a genuine hang still alarms:
"checks": {"io_stall": {"amber": 50, "red": 80}}
Only the limits named change; the others keep their defaults. Several checks share the one object: {"io_stall": false, "load": {"amber": 4}}.
To exclude one thing rather than a whole check, name it under ignore and no check alarms on it, however it goes wrong: a USB backup disk that is usually unplugged, a spare interface, a service that fails harmlessly, a port that is only sometimes open. Use the name as it appears in the alert.
"ignore": ["/mnt/usbbackup", "sdf", "eth3", "certbot.service", "tcp/8080"]
Then restart the agent: systemctl restart oversight-agent, or initctl restart oversight-agent on a Synology. A misspelt check or limit name is refused at start with a message listing the right ones, so check systemctl status oversight-agent afterwards: an agent that did not start is a sensor that stops answering. To see the verdict the edited file gives before restarting, /opt/oversight/agent --dump status samples for a minute and prints it.
| Area | Check | What the alert says | Limits, with their defaults |
|---|---|---|---|
| storage | md_failed | md2 degraded, 1 of 8 members missing, md2 inactive | none |
md_rebuild | md2 rebuilding, 40% | none | |
disk_offline | sdc offline since 14:02 | none | |
disk_slow | sdc slow: w_await 12x its md2 siblings for 14 min | none | |
disk_latency | sdc w_await 320ms for 12 min | amber 50, red 200 (ms, SSD and NVMe); amber_rotational 250, red_rotational 1000 (ms, spinning disks) | |
iscsi | iSCSI iqn.2026-01.uk.example:lun1 offline since 14:02 | none | |
multipath | mpatha 1 of 2 paths down | none | |
nfs_slow | /mnt/nfs write 800ms for 6 min | amber 500, red 2000 (ms) | |
nfs_timeouts | /mnt/nfs 12 major timeouts in 10 min | none | |
io_stall | every task stalled on I/O 10% of the last minute, for 22 min | amber 10, red 30 (%) | |
lv_suspended | vg0-data suspended since 14:02 | none | |
| filesystem | fs_missing | /srv not mounted, /srv has /dev/sdb1 mounted, fstab expects /dev/sdc1, swap /dev/sda2 not active | none |
fs_readonly | /srv read-only since 14:02 | none | |
fs_unresponsive | /mnt/nfs not responding since 14:02 | none | |
fs_full | /srv 92% full, /srv inodes 91% used | amber 90, red 95 (%) | |
fs_filling | /srv full in about 9 h at the last hour's rate | hours 12 | |
| memory | oom | OOM killer killed 2 processes in the last hour, last at 14:02 | none |
mem_pressure | every task stalled on memory 6% of the last minute, for 7 min | amber 5, red 20 (%) | |
swapping | swapping in 2.4 MB/s for 6 min | amber 1 (MB/s) | |
mem_low | memory 3.2% available (1.0 GB of 31.3 GB) for 6 min | amber 5 (% available) | |
| processes | dstate | 14 processes in uninterruptible sleep (D) for 6 min | amber 10, red 50 (processes) |
zombies | 34 zombies, up 12 in 30 min | amber 20 (zombies) | |
pids | threads at 85% of pid_max (27853 of 32768) | amber 80, red 95 (%) | |
| services | unit_failed | php-fpm.service failed since 14:02 | none |
unit_restarting | php-fpm.service restarted 7 times in 15 min | amber 3 (restarts in 15 minutes) | |
port_lost | tcp/3306 not listening since 14:02 | none | |
| network | link_down | eth1 link down since 14:02 | none |
bond_degraded | bond0 1 of 2 members down | none | |
link_speed | eth0 at 100 Mb/s, was 1 Gb/s, eth0 at half duplex | none | |
net_errors | eth0 3.2 errors a second for 6 min | amber 1 (errors a second) | |
tcp_retrans | TCP retransmitting 6.3% of segments for 11 min | amber 5 (%) | |
conntrack | conntrack table 84% full (220200 of 262144) | amber 80, red 95 (%) | |
| cpu | load | load 2.4 per CPU (8 CPUs) for 16 min, which adds largely processes waiting on I/O, see io_stall when that is the cause | amber 2, red 4 (per CPU) |
steal | CPU steal 24% for 16 min, the hypervisor is oversubscribed | amber 20 (%) | |
| host_ | time | clock not synchronised for 11 min, clock 2.3 s out for 11 min | amber 1 (seconds) |
fds | open files at 85% of the system limit (891289 of 1048576) | amber 80, red 95 (%) | |
temperature | Package id 0 92C, maximum 90C | none, the hardware's own limits are used | |
rebooted | rebooted at 14:02 | none | |
gpu_lost | gpu0 (0000:01:00.0) gone from the driver since 14:02 | none |
One check can bring on another. A storage stall also leaves processes waiting (dstate) and raises the load (load), so a backup server may need those eased as well as io_stall. Change the one the alert names first, and the others only if they raise too.
For the first 5 minutes after the host boots, nothing raises except the amber rebooted, so mounts and services that take a minute to arrive after a restart do not page anyone. Anything still wrong when the 5 minutes are up raises at once. Most checks must also hold for a while before they raise, and be clear as long again before they drop, so a backup window or a burst does not flap the sensor. Some compare with what the agent has seen since it started, such as a filesystem that was writable or a link that was up; restarting the agent starts that memory afresh.
Host: CPU, load, processes and the clock
/v1/host is figures, for charting. Nothing on the figure endpoints raises an alarm by itself; Status has already judged them. Useful rows:
| Slot | Name | Method | Expression | Unit |
|---|---|---|---|---|
| SV01 | CPU busy | JSON | cpu_busy_pct | % |
| SV02 | CPU iowait | JSON | cpu_iowait_pct | % |
| SV03 | CPU steal | JSON | cpu_steal_pct | % |
| SV04 | Load per CPU | JSON | load_per_cpu | |
| SV05 | Processes | JSON | procs_total | count |
| SV06 | Stuck processes | JSON | procs_dstate | count |
| SV07 | Zombies | JSON | procs_zombie | count |
| SV08 | File handles | JSON | fds_used_pct | % |
| SV09 | Clock offset | JSON | time_offset_ms | ms |
| SV10 | Temperature | JSON | temp_c | C |
| SD01 | Kernel | JSON | kernel | |
| SD02 | Boot id | JSON | boot_id |
temp_c is the hottest temperature sensor the hardware reports, with temp_sensor naming it, which on a NAS or similar small box is as near as it comes to the box's own temperature. A virtual machine has none, and it reads null. Also available: uptime_s, load1, load5, load15, cpu_user_pct, cpu_system_pct, psi_cpu_some, psi_io_some, psi_io_full, procs_running and threads_used_pct.
Two rules worth having here, because they say something Status does not: Boot id changed as WARN catches every reboot, however quick, and Kernel changed as WARN shows when an update has taken effect. Figures the host cannot provide, such as the pressure figures without PSI, read null rather than zero.
Memory
/v1/memory:
| Slot | Name | Method | Expression | Unit |
|---|---|---|---|---|
| SV01 | Available | JSON | available_pct | % |
| SV02 | Swap used | JSON | swap_used_pct | % |
| SV03 | Swap in | JSON | swapin_kbps | KB/s |
| SV04 | Major faults | JSON | majfault_ps | /s |
| SV05 | Memory stall | JSON | psi_mem_full | % |
| SV06 | OOM kills, last hour | JSON | oom_kills_hour | count |
Also: total, available, used_pct, cached, dirty, writeback, swap_total (sizes in bytes), swapout_kbps, scan_direct_ps, psi_mem_some, and on a ZFS host zfs_arc and zfs_arc_min. Available memory counts the ZFS cache above its minimum as available, as it really is.
Disk: latency per physical disk
/v1/disk covers physical disks only, keyed by device name. The path is disks., the device, then the figure:
| Slot | Name | Method | Expression | Unit |
|---|---|---|---|---|
| SV01 | sda read latency | JSON | disks.sda.r_await_ms | ms |
| SV02 | sda write latency | JSON | disks.sda.w_await_ms | ms |
| SV03 | sda queue | JSON | disks.sda.queue_depth | |
| SV04 | Worst write latency | JSON | max_w_await_ms | ms |
| SD01 | Worst disk | JSON | max_w_await_disk |
Per disk there are also reads_ps, writes_ps, read_bytes_ps, write_bytes_ps and util_pct; across the host, max_r_await_ms, max_queue_depth and max_util_pct, each with its _disk. The max_ figures are the ones to chart on a host with many disks, and they survive a disk being renamed after a reboot, which sd names can be. On SSD and NVMe, util_pct reaching 100 is not saturation; latency and queue are what matter there. All are rates over the last minute, and read null in the agent's first minute.
Storage: only what is in distress
/v1/storage always lists every md array, healthy or not, as /proc/mdstat shows it. LVM volumes, filesystems and NFS mounts appear only while a check is raised against them. Each carries a problem, empty for an array that is well:
{"alarms":1,
"md":{"md2":{"problem":"degraded, 1 of 8 members missing","level":"raid5","state":"clean",
"sync_action":"idle","members":8,"active":7,"degraded":1,"map":"UU_UUUUU",
"devices":{"sda5":{"slot":0,"state":"in_sync"},"sdb5":{"slot":1,"state":"in_sync"}, ...}}},
"fs":{}}
map is mdstat's own picture of the array, one character per slot, U for a member in sync and _ for one missing, so UU_UUUUU says the third disk has gone. The member names under devices say which partition, and so which disk, to go and look at.
| Slot | Name | Method | Expression |
|---|---|---|---|
| SV01 | Storage alarms | JSON | alarms |
| SD01 | md2 | JSON | md.md2.map |
| SV02 | md2 active | JSON | md.md2.active |
| SV05 | md2 rebuilt | JSON | md.md2.sync_pct |
| SD03 | sdc5 in md2 | JSON | md.md2.devices.sdc5.state |
| SD02 | /srv | JSON | fs./srv.problem |
| SV03 | /srv used | JSON | fs./srv.used_pct |
| SV04 | NFS write time | JSON | nfs./mnt/pve/nfs1.write_exe_ms |
Status already raises a degraded array, so none of this is needed to be told. It is here for a sensor that watches one array by itself: md2 map not matching /^U+$/ as CRIT says the array is not whole, and the map in the alert says which slot. Mount points go into the path as they are, slashes and all; the root filesystem is fs./.problem. A filesystem that is not in distress is simply not there, so its slots read empty and no rule on them matches, which is what you want. alarms is always present, so the sensor is never left with nothing to read.
Network
/v1/network, per interface under interfaces., then the name:
| Slot | Name | Method | Expression | Unit |
|---|---|---|---|---|
| SV01 | eth0 in | JSON | interfaces.eth0.rx_bytes_ps | bytes/s |
| SV02 | eth0 out | JSON | interfaces.eth0.tx_bytes_ps | bytes/s |
| SV03 | eth0 errors | JSON | interfaces.eth0.errors_ps | /s |
| SV04 | eth0 speed | JSON | interfaces.eth0.speed_mbps | Mb/s |
| SV05 | Established | JSON | tcp_established | count |
| SV06 | Retransmitted | JSON | tcp_retrans_pct | % |
| SV07 | Conntrack | JSON | conntrack_pct | % |
Also per interface carrier (1 or 0) and drops_ps, and overall tcp_time_wait. A dot in an interface name becomes an underscore, because the dot separates the path: VLAN eth0.100 is interfaces.eth0_100.rx_bytes_ps. The real name is in its name field. Interfaces belonging to guests (veth, tap, fwbr and the like) come and go with them and are not listed.
GPU
/v1/gpu lists NVIDIA GPUs, from what the driver publishes without its own library. It is empty on a host without one.
| Slot | Name | Method | Expression |
|---|---|---|---|
| SV01 | GPUs | JSON | count |
| SD01 | GPU present | JSON | gpus.gpu0.present |
| SD02 | Driver | JSON | driver |
| SD03 | GPU power state | JSON | gpus.gpu0.power_state |
| SV02 | PCIe fatal errors | JSON | gpus.gpu0.aer_fatal |
Also per GPU: bus, model, uuid, firmware, link_speed, link_width, aer_nonfatal and aer_correctable. A GPU that has gone stays listed with present false. Status already raises the one thing that matters, gpu_lost in the Host area, red when a GPU drops out of the driver or reports a PCIe fatal error. A GPU is never judged on how busy it is: one running models sits at 100% or 0%, and both are normal. Utilisation, temperature and memory use are not here, because they need NVIDIA's own library, which the agent does not load. Driver changed as WARN is a useful rule, since a driver update is the usual reason a GPU host behaves differently the next day.
On a Synology
The agent knows when it is on a Synology, and /v1/host says so in platform and model, for example Synology DSM 6.2.4-25556 on a DS1812+. The install line is the same, run as root: log in over SSH as an administrator, then sudo -i. On DSM it differs in a few ways, all handled for you:
- It runs as
nobody, since DSM has no way to add a system user from the shell, and is started by an upstart job,/etc/init/oversight-agent.conf, which restarts it if it stops.initctl status oversight-agentshows it;initctl restart oversight-agentafter editing the config. - A DSM update can remove the upstart job. If the sensor stops answering after an update, run the install line again: it puts the job and binary back and leaves the config alone.
- The system arrays are judged by their members, not their slots. DSM mirrors its system (md0) and swap (md1) partitions across every bay, so on a unit with empty bays those arrays always look short of members. On DSM they alarm when a member is faulty or out of sync, or when a disk drops out of them; the data arrays (md2 upwards) are judged normally.
- Services are unavailable. DSM 6 does not use systemd, so failed units and crash loops cannot be seen and
unavailablesays so. Ports that stop listening are still watched. - Older DSM kernels have no pressure figures and no OOM counter, so
io_stall,mem_pressureandoomshow underunavailabletoo. Available memory is worked out from free memory and the caches the kernel can drop, since those kernels do not report it.
An eSATA or USB disk plugged into the unit is watched like any other disk. The small USB boot module inside the unit (synoboot) is not a disk that holds anything and is left out.
If it does not work
| What you see | What it is |
|---|---|
| Connection refused | The agent is not listening where the probe is looking: on 127.0.0.1, or on another port. Set listen as above; ss -ltn | grep 29731 on the host shows what it is bound to. |
| Nothing at all, but curl works on the host | The host's firewall is not letting the probe in. Allow TCP 29731 from the probe's address in firewalld, ufw, nftables or whichever this host runs. |
| HTTP status 401 | No token was sent, or it does not match. Check there is a Linux agent credential on the device or above it, holding a token listed in tokens. |
| HTTP status 403 | The probe's address is not in allow. Behind NAT or a proxy, it is the address the agent sees that counts, not the probe's own. |
| The service will not start | journalctl -u oversight-agent says why: usually a listen beyond the host with neither tokens nor allow, or a misspelt key in the config. |
| Figure slots read empty for the first minute | Rates are worked out over the last minute, so a freshly started agent has none yet. |
Something expected under unavailable | That check cannot see on this host. no /proc/pressure means the kernel was booted without PSI: add psi=1 to the kernel command line to enable the pressure checks. |