Articles Linux agent: a verdict from the host itself setup, LINUX, updated 2026-09-23

The Oversight Linux agent is a single small program that runs on a Linux host, watches it continuously and decides for itself whether anything is wrong. It serves that verdict over HTTP: nothing wrong, or what is wrong and how badly. Oversight reads it with a Linux agent sensor.

The agent does the judging on the host because it knows the host far better than a threshold in a template can: whether a filesystem was writable an hour ago, which network ports have ever had a link, whether a service is quietly crash-looping. So a sensor on it needs two rules, not two hundred. It is read only by design, runs as an unprivileged user, has no way to run commands, and nothing it serves is a secret or a process list.

It is for bare-metal servers and ordinary virtual machines. Proxmox hosts, guests and containers are better watched through their own APIs.

Installing it on the host

As root on the host to be watched:

curl -fsSL https://oversight.im/linuxsensor | bash

or, from a user with sudo, curl -fsSL https://oversight.im/linuxsensor | sudo bash. The script picks the build for the machine (x86_64, arm64 or 32-bit ARM), checks it against its published checksum, and installs /opt/oversight/agent, an oversight system user, /etc/oversight/agent.json and the oversight-agent service.

It then asks one question, and that is the whole setup:

Address for the agent to listen on, which the probe connects to [0.0.0.0:29731]:

Press Enter for every interface, or give the address the probe will reach it on, such as 10.1.1.100. An address with no port takes the default one. It generates a bearer token, prints it, and starts the service:

/etc/oversight/agent.json: listen 0.0.0.0:29731, 1 token(s), from any address

  token: 7pQz3XfKdR2mVn8sLtYwJhB4

oversight-agent enabled and running, listening on 0.0.0.0:29731

Keep that token. It goes in a Linux agent credential in Oversight, below. It is in /etc/oversight/agent.json as well, so it can be read back off the host later.

Running the same line again is how it is upgraded. An upgrade asks nothing and changes nothing it was not told to: the config, and a service file already there, are left as they are.

Where a script installs it rather than a person, the answers go on the command line and nothing is asked:

curl -fsSL https://oversight.im/linuxsensor | sudo bash -s -- --listen 10.1.1.100:29731 --token a-long-random-token --allow 10.1.1.25

Check it from the host before going near Oversight:

curl -s -H "Authorization: Bearer a-long-random-token" http://10.1.1.100:29731/v1/status

The host's firewall needs a rule. Nothing in the install opens the port: firewalld, ufw, nftables or whatever this host runs has to allow TCP 29731 in from the probe. Until it does, the curl above works from the host itself while the sensor sees nothing.

It speaks plain HTTP. Where it must be reached across the internet, put a reverse proxy such as nginx in front of it for TLS, set allow to the proxy's address and let the proxy decide who reaches it.

Changing it afterwards

Everything the agent is told is in /etc/oversight/agent.json, which only root and the oversight user can read, because it holds the token:

{
  "listen": "10.1.1.100:29731",
  "tokens": ["a-long-random-token"],
  "allow": ["10.1.1.25"],
  "checks": {},
  "ignore": []
}
KeyWhat it does
listenThe one address and port it answers on. 0.0.0.0:29731 means every interface.
tokensBearer tokens it accepts. The installer generates the first one. Two may be listed at once, so a token can be changed without a gap.
allowThe addresses, or ranges such as 10.1.1.0/24, allowed to connect. Put your probe's address here.
checksTurns a check off, or changes the limits it alarms at, on this host only. See Turning a check off, or changing its limits, below.
ignoreThings no check alarms on: a mount point, disk, interface, service or port. Also below.

None of it has to be edited by hand. Each of these adds to the file and restarts the service:

/opt/oversight/agent --listen 10.1.1.100:29731
/opt/oversight/agent --token another-long-random-token
/opt/oversight/agent --allow 10.1.1.25,10.1.2.0/24

A token already listed is not added twice, and none is ever taken away, so a token can be changed without a gap. The agent refuses to listen beyond the host with neither tokens nor allow set, and says so, which is why the installer generates a token; setting allow as well is better still. After editing the file by hand, systemctl restart oversight-agent.

Adding the sensor

On the device that is this host, add a sensor of type Linux agent. If the agent has tokens set, add a credential of type Linux agent on the Credentials panel of that device, or of a group above it to share one token across many hosts, with the token in the Bearer token field. The sensor sends it for you; there is no header to write.

SettingWhat to put
EndpointWhat to read. Status is the one that raises alarms; the others are figures for charts. See below.
Agent addresshttp and 29731, unless it was installed with another port, or https and the proxy's port where one is in front of it.
Interval60 seconds for Status. The agent samples every 10 seconds whether anyone asks or not, so polling faster gains nothing. 300 seconds is plenty for a figures endpoint.

A new Linux agent sensor is ready to use as soon as it is saved. It starts on the Status endpoint with the extraction rows and rules below already filled in. Changing the endpoint afterwards does not change the extraction rows, because each endpoint carries different fields: replace them with the ones listed for that endpoint.

Every endpoint also fills SV20, HTTP status, by itself. It is 200 when the agent answered, 401 for a missing or wrong token and 403 for a probe outside allow. Because it always reads, a sensor on any endpoint stays readable even when the fields it extracts are absent.

Status: the verdict

/v1/status is quiet while all is well:

{"alarm":"green","alarms":0,"status":"","storage":"","filesystem":"", ... ,"unavailable":[]}

and when something is wrong, says what, area by area, worst first:

{"alarm":"red","alarms":2,
 "status":"md0 degraded, 1 of 2 members missing; php-fpm.service restarted 7 times in 15 min",
 "storage":"md0 degraded, 1 of 2 members missing",
 "services":"php-fpm.service restarted 7 times in 15 min", ...}

These are the rows a new sensor starts with:

SlotNameMethodExpressionHolds
SD01AlarmJSONalarmgreen, amber or red
SD02StatusJSONstatusEvery current problem, for the alert text
SD03StorageJSONstorageArrays, disks, iSCSI, multipath, NFS, LVM
SD04FilesystemJSONfilesystemMissing, read-only, unresponsive, full, filling
SD05MemoryJSONmemoryOOM kills, pressure, swapping, low memory
SD06ProcessesJSONprocessesStuck (D state), zombies, running out of PIDs
SD07ServicesJSONservicesFailed units, crash loops, ports that stopped listening. A user's own login session (user@1000.service) is not counted: on a desktop install it often fails its own stop on logout, which is not a service down
SD08NetworkJSONnetworkLinks down, bonds degraded, errors, retransmissions
SD09CPUJSONcpuSustained load, steal
SD10HostJSONhost_Clock, file handles, temperature, a recent reboot
SV01AlarmsJSONalarmsHow many problems there are now, charted
SV02UptimeJSONuptime_sSeconds since the host booted
SV03 to SV05Load 1, 5 and 15 minJSONload1, load5, load15Load averages, charted together
SV06Agent uptimeJSONagent_uptime_sSeconds since the agent started

An area with nothing wrong is an empty string. Note the trailing underscore on host_: plain host is the host's name.

And the rules, which are the whole of the judging Oversight needs to do:

SlotRuleStateWhy
SV20, HTTP statusnot equal to 200CRITThe agent refused the probe. A token or allow problem should not look like a healthy host.
SD01, Alarmequal to redCRITSomething is broken now.
SD01, Alarmequal to amberWARNSomething is heading that way, or needs a look.

Actions then work as for any other sensor: bind a notification rule to the device or a group above it. The alert's detail is the agent's own words, such as php-fpm.service failed since 14:02, so the message says what is wrong, not just that the alarm slot read red.

Treating one area differently. A rule can be written against any area slot. To make any storage problem CRIT whatever its colour, add SD03 matches /./ as CRIT: an area is empty when it has nothing to say, so anything in it at all matches. To send storage problems to a different team from everything else, give the host a second Linux agent sensor on Status with only that rule, and bind the storage team's notification rule to it. Rules set a state; who is told follows the sensor.

Checks that cannot run on this host are listed in unavailable, for example the pressure checks on a RHEL or AlmaLinux kernel booted without psi=1. That is not an alarm, but it can be counted: a spare slot with method JSONCOUNT and expression unavailable reads how many there are.

Turning a check off, or changing its limits

The defaults are meant to be right for nearly every host, but not every host is ordinary. A backup server spends hours with its disks flat out, and the agent will rightly say so. The fix belongs on that host, in /etc/oversight/agent.json, rather than in a rule in Oversight that hides a whole area. It can only be done on the host: nothing in Oversight changes the agent's settings.

First find the check. The alert's detail is the agent's own words, and the table below gives the check each message comes from: every task stalled on I/O 10% of the last minute, for 22 min is io_stall. Then, under checks, either switch it off:

"checks": {"io_stall": false}

or keep it and raise its limits, so that a normal night's work passes and a genuine hang still alarms:

"checks": {"io_stall": {"amber": 50, "red": 80}}

Only the limits named change; the others keep their defaults. Several checks share the one object: {"io_stall": false, "load": {"amber": 4}}.

To exclude one thing rather than a whole check, name it under ignore and no check alarms on it, however it goes wrong: a USB backup disk that is usually unplugged, a spare interface, a service that fails harmlessly, a port that is only sometimes open. Use the name as it appears in the alert.

"ignore": ["/mnt/usbbackup", "sdf", "eth3", "certbot.service", "tcp/8080"]

Then restart the agent: systemctl restart oversight-agent, or initctl restart oversight-agent on a Synology. A misspelt check or limit name is refused at start with a message listing the right ones, so check systemctl status oversight-agent afterwards: an agent that did not start is a sensor that stops answering. To see the verdict the edited file gives before restarting, /opt/oversight/agent --dump status samples for a minute and prints it.

AreaCheckWhat the alert saysLimits, with their defaults
storagemd_failedmd2 degraded, 1 of 8 members missing, md2 inactivenone
md_rebuildmd2 rebuilding, 40%none
disk_offlinesdc offline since 14:02none
disk_slowsdc slow: w_await 12x its md2 siblings for 14 minnone
disk_latencysdc w_await 320ms for 12 minamber 50, red 200 (ms, SSD and NVMe); amber_rotational 250, red_rotational 1000 (ms, spinning disks)
iscsiiSCSI iqn.2026-01.uk.example:lun1 offline since 14:02none
multipathmpatha 1 of 2 paths downnone
nfs_slow/mnt/nfs write 800ms for 6 minamber 500, red 2000 (ms)
nfs_timeouts/mnt/nfs 12 major timeouts in 10 minnone
io_stallevery task stalled on I/O 10% of the last minute, for 22 minamber 10, red 30 (%)
lv_suspendedvg0-data suspended since 14:02none
filesystemfs_missing/srv not mounted, /srv has /dev/sdb1 mounted, fstab expects /dev/sdc1, swap /dev/sda2 not activenone
fs_readonly/srv read-only since 14:02none
fs_unresponsive/mnt/nfs not responding since 14:02none
fs_full/srv 92% full, /srv inodes 91% usedamber 90, red 95 (%)
fs_filling/srv full in about 9 h at the last hour's ratehours 12
memoryoomOOM killer killed 2 processes in the last hour, last at 14:02none
mem_pressureevery task stalled on memory 6% of the last minute, for 7 minamber 5, red 20 (%)
swappingswapping in 2.4 MB/s for 6 minamber 1 (MB/s)
mem_lowmemory 3.2% available (1.0 GB of 31.3 GB) for 6 minamber 5 (% available)
processesdstate14 processes in uninterruptible sleep (D) for 6 minamber 10, red 50 (processes)
zombies34 zombies, up 12 in 30 minamber 20 (zombies)
pidsthreads at 85% of pid_max (27853 of 32768)amber 80, red 95 (%)
servicesunit_failedphp-fpm.service failed since 14:02none
unit_restartingphp-fpm.service restarted 7 times in 15 minamber 3 (restarts in 15 minutes)
port_losttcp/3306 not listening since 14:02none
networklink_downeth1 link down since 14:02none
bond_degradedbond0 1 of 2 members downnone
link_speedeth0 at 100 Mb/s, was 1 Gb/s, eth0 at half duplexnone
net_errorseth0 3.2 errors a second for 6 minamber 1 (errors a second)
tcp_retransTCP retransmitting 6.3% of segments for 11 minamber 5 (%)
conntrackconntrack table 84% full (220200 of 262144)amber 80, red 95 (%)
cpuloadload 2.4 per CPU (8 CPUs) for 16 min, which adds largely processes waiting on I/O, see io_stall when that is the causeamber 2, red 4 (per CPU)
stealCPU steal 24% for 16 min, the hypervisor is oversubscribedamber 20 (%)
host_timeclock not synchronised for 11 min, clock 2.3 s out for 11 minamber 1 (seconds)
fdsopen files at 85% of the system limit (891289 of 1048576)amber 80, red 95 (%)
temperaturePackage id 0 92C, maximum 90Cnone, the hardware's own limits are used
rebootedrebooted at 14:02none
gpu_lostgpu0 (0000:01:00.0) gone from the driver since 14:02none

One check can bring on another. A storage stall also leaves processes waiting (dstate) and raises the load (load), so a backup server may need those eased as well as io_stall. Change the one the alert names first, and the others only if they raise too.

For the first 5 minutes after the host boots, nothing raises except the amber rebooted, so mounts and services that take a minute to arrive after a restart do not page anyone. Anything still wrong when the 5 minutes are up raises at once. Most checks must also hold for a while before they raise, and be clear as long again before they drop, so a backup window or a burst does not flap the sensor. Some compare with what the agent has seen since it started, such as a filesystem that was writable or a link that was up; restarting the agent starts that memory afresh.

Host: CPU, load, processes and the clock

/v1/host is figures, for charting. Nothing on the figure endpoints raises an alarm by itself; Status has already judged them. Useful rows:

SlotNameMethodExpressionUnit
SV01CPU busyJSONcpu_busy_pct%
SV02CPU iowaitJSONcpu_iowait_pct%
SV03CPU stealJSONcpu_steal_pct%
SV04Load per CPUJSONload_per_cpu
SV05ProcessesJSONprocs_totalcount
SV06Stuck processesJSONprocs_dstatecount
SV07ZombiesJSONprocs_zombiecount
SV08File handlesJSONfds_used_pct%
SV09Clock offsetJSONtime_offset_msms
SV10TemperatureJSONtemp_cC
SD01KernelJSONkernel
SD02Boot idJSONboot_id

temp_c is the hottest temperature sensor the hardware reports, with temp_sensor naming it, which on a NAS or similar small box is as near as it comes to the box's own temperature. A virtual machine has none, and it reads null. Also available: uptime_s, load1, load5, load15, cpu_user_pct, cpu_system_pct, psi_cpu_some, psi_io_some, psi_io_full, procs_running and threads_used_pct.

Two rules worth having here, because they say something Status does not: Boot id changed as WARN catches every reboot, however quick, and Kernel changed as WARN shows when an update has taken effect. Figures the host cannot provide, such as the pressure figures without PSI, read null rather than zero.

Memory

/v1/memory:

SlotNameMethodExpressionUnit
SV01AvailableJSONavailable_pct%
SV02Swap usedJSONswap_used_pct%
SV03Swap inJSONswapin_kbpsKB/s
SV04Major faultsJSONmajfault_ps/s
SV05Memory stallJSONpsi_mem_full%
SV06OOM kills, last hourJSONoom_kills_hourcount

Also: total, available, used_pct, cached, dirty, writeback, swap_total (sizes in bytes), swapout_kbps, scan_direct_ps, psi_mem_some, and on a ZFS host zfs_arc and zfs_arc_min. Available memory counts the ZFS cache above its minimum as available, as it really is.

Disk: latency per physical disk

/v1/disk covers physical disks only, keyed by device name. The path is disks., the device, then the figure:

SlotNameMethodExpressionUnit
SV01sda read latencyJSONdisks.sda.r_await_msms
SV02sda write latencyJSONdisks.sda.w_await_msms
SV03sda queueJSONdisks.sda.queue_depth
SV04Worst write latencyJSONmax_w_await_msms
SD01Worst diskJSONmax_w_await_disk

Per disk there are also reads_ps, writes_ps, read_bytes_ps, write_bytes_ps and util_pct; across the host, max_r_await_ms, max_queue_depth and max_util_pct, each with its _disk. The max_ figures are the ones to chart on a host with many disks, and they survive a disk being renamed after a reboot, which sd names can be. On SSD and NVMe, util_pct reaching 100 is not saturation; latency and queue are what matter there. All are rates over the last minute, and read null in the agent's first minute.

Storage: only what is in distress

/v1/storage always lists every md array, healthy or not, as /proc/mdstat shows it. LVM volumes, filesystems and NFS mounts appear only while a check is raised against them. Each carries a problem, empty for an array that is well:

{"alarms":1,
 "md":{"md2":{"problem":"degraded, 1 of 8 members missing","level":"raid5","state":"clean",
              "sync_action":"idle","members":8,"active":7,"degraded":1,"map":"UU_UUUUU",
              "devices":{"sda5":{"slot":0,"state":"in_sync"},"sdb5":{"slot":1,"state":"in_sync"}, ...}}},
 "fs":{}}

map is mdstat's own picture of the array, one character per slot, U for a member in sync and _ for one missing, so UU_UUUUU says the third disk has gone. The member names under devices say which partition, and so which disk, to go and look at.

SlotNameMethodExpression
SV01Storage alarmsJSONalarms
SD01md2JSONmd.md2.map
SV02md2 activeJSONmd.md2.active
SV05md2 rebuiltJSONmd.md2.sync_pct
SD03sdc5 in md2JSONmd.md2.devices.sdc5.state
SD02/srvJSONfs./srv.problem
SV03/srv usedJSONfs./srv.used_pct
SV04NFS write timeJSONnfs./mnt/pve/nfs1.write_exe_ms

Status already raises a degraded array, so none of this is needed to be told. It is here for a sensor that watches one array by itself: md2 map not matching /^U+$/ as CRIT says the array is not whole, and the map in the alert says which slot. Mount points go into the path as they are, slashes and all; the root filesystem is fs./.problem. A filesystem that is not in distress is simply not there, so its slots read empty and no rule on them matches, which is what you want. alarms is always present, so the sensor is never left with nothing to read.

Network

/v1/network, per interface under interfaces., then the name:

SlotNameMethodExpressionUnit
SV01eth0 inJSONinterfaces.eth0.rx_bytes_psbytes/s
SV02eth0 outJSONinterfaces.eth0.tx_bytes_psbytes/s
SV03eth0 errorsJSONinterfaces.eth0.errors_ps/s
SV04eth0 speedJSONinterfaces.eth0.speed_mbpsMb/s
SV05EstablishedJSONtcp_establishedcount
SV06RetransmittedJSONtcp_retrans_pct%
SV07ConntrackJSONconntrack_pct%

Also per interface carrier (1 or 0) and drops_ps, and overall tcp_time_wait. A dot in an interface name becomes an underscore, because the dot separates the path: VLAN eth0.100 is interfaces.eth0_100.rx_bytes_ps. The real name is in its name field. Interfaces belonging to guests (veth, tap, fwbr and the like) come and go with them and are not listed.

GPU

/v1/gpu lists NVIDIA GPUs, from what the driver publishes without its own library. It is empty on a host without one.

SlotNameMethodExpression
SV01GPUsJSONcount
SD01GPU presentJSONgpus.gpu0.present
SD02DriverJSONdriver
SD03GPU power stateJSONgpus.gpu0.power_state
SV02PCIe fatal errorsJSONgpus.gpu0.aer_fatal

Also per GPU: bus, model, uuid, firmware, link_speed, link_width, aer_nonfatal and aer_correctable. A GPU that has gone stays listed with present false. Status already raises the one thing that matters, gpu_lost in the Host area, red when a GPU drops out of the driver or reports a PCIe fatal error. A GPU is never judged on how busy it is: one running models sits at 100% or 0%, and both are normal. Utilisation, temperature and memory use are not here, because they need NVIDIA's own library, which the agent does not load. Driver changed as WARN is a useful rule, since a driver update is the usual reason a GPU host behaves differently the next day.

On a Synology

The agent knows when it is on a Synology, and /v1/host says so in platform and model, for example Synology DSM 6.2.4-25556 on a DS1812+. The install line is the same, run as root: log in over SSH as an administrator, then sudo -i. On DSM it differs in a few ways, all handled for you:

  • It runs as nobody, since DSM has no way to add a system user from the shell, and is started by an upstart job, /etc/init/oversight-agent.conf, which restarts it if it stops. initctl status oversight-agent shows it; initctl restart oversight-agent after editing the config.
  • A DSM update can remove the upstart job. If the sensor stops answering after an update, run the install line again: it puts the job and binary back and leaves the config alone.
  • The system arrays are judged by their members, not their slots. DSM mirrors its system (md0) and swap (md1) partitions across every bay, so on a unit with empty bays those arrays always look short of members. On DSM they alarm when a member is faulty or out of sync, or when a disk drops out of them; the data arrays (md2 upwards) are judged normally.
  • Services are unavailable. DSM 6 does not use systemd, so failed units and crash loops cannot be seen and unavailable says so. Ports that stop listening are still watched.
  • Older DSM kernels have no pressure figures and no OOM counter, so io_stall, mem_pressure and oom show under unavailable too. Available memory is worked out from free memory and the caches the kernel can drop, since those kernels do not report it.

An eSATA or USB disk plugged into the unit is watched like any other disk. The small USB boot module inside the unit (synoboot) is not a disk that holds anything and is left out.

If it does not work

What you seeWhat it is
Connection refusedThe agent is not listening where the probe is looking: on 127.0.0.1, or on another port. Set listen as above; ss -ltn | grep 29731 on the host shows what it is bound to.
Nothing at all, but curl works on the hostThe host's firewall is not letting the probe in. Allow TCP 29731 from the probe's address in firewalld, ufw, nftables or whichever this host runs.
HTTP status 401No token was sent, or it does not match. Check there is a Linux agent credential on the device or above it, holding a token listed in tokens.
HTTP status 403The probe's address is not in allow. Behind NAT or a proxy, it is the address the agent sees that counts, not the probe's own.
The service will not startjournalctl -u oversight-agent says why: usually a listen beyond the host with neither tokens nor allow, or a misspelt key in the config.
Figure slots read empty for the first minuteRates are worked out over the last minute, so a freshly started agent has none yet.
Something expected under unavailableThat check cannot see on this host. no /proc/pressure means the kernel was booted without PSI: add psi=1 to the kernel command line to enable the pressure checks.