How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message, or ask a question in your own words.
File on the probe host: has last night's backup landed?
A file sensor asks whether one file exists on the probe's own host, and if so how big it is and how old. It is for checking that something landed: last night's backup, a router's config export, a nightly report. Nothing is read from the file. The probe looks it up and returns its size and modification time, and that is all that leaves the host.
The probe host is usually not where the file lives. The idea is to mount the share that holds it on the probe host, read-only, and point the sensor at the path. Anything the operating system can mount can then be checked: an SMB share signed in with Kerberos, a folder over SSH with sshfs, WebDAV, a cloud bucket through rclone, or an older NFS server. Where the file is on an NFSv4 server or an SMB share that takes a username and password, the NFS file and SMB file sensors reach it directly with nothing to mount.
Before you start
- Mount the share on the probe host, read-only. The probe only ever looks, but a read-only mount means nothing on the probe host can change the backups either.
- The probe's own user must be able to read the path. The probe does not run as root. Every folder on the way to the file needs to be searchable by it, and a mount owned by root with no access for others will read as a permission failure.
- Mount network shares soft, with a short timeout, where the filesystem offers it. A share mounted hard waits for ever on a dead server. The sensor copes with that, as described under When it fails, but a soft mount turns it into a clear error instead.
- Mind the probe groups. A sensor in a group of several probes can be polled by any of them, and moves to another if its probe stops. Mount the share on every probe in the group, or use a group with one probe in it, or a probe that has the share will be told the file is missing whenever the sensor lands on one that does not.
- No credential is needed. Whatever access the mount needs is set up on the probe host, where the operating system holds it.
Setting it up
| Setting | What to put |
|---|---|
| File path | The absolute path on the probe host, such as /mnt/backups/edge1.cfg. A dated file can be named with tokens, which are filled in on UK time: /mnt/backups/edge1-{{yesterday}}.cfg. See Dated files below before using {{isodate}}. |
| Mount point | Optional, and worth setting. Where the share should be mounted, such as /mnt/backups. If the share drops, the folder underneath is still there, empty, on the probe's own disk, and without this the sensor can only say the file is missing. With it, the sensor sees that the path is no longer on the share and says the share is not mounted, which sends the right person to look. |
| Interval | A file that changes once a day does not need checking every minute. 15 minutes to an hour suits a nightly backup, and costs a fraction as much. A file sensor is 1 credit a reading. |
| Response timeout | How long the lookup may take. 5 to 10 seconds is ample for a local share; a lookup that takes longer means the mount is in trouble. |
| Failures before a probe calls it down | 1 is reasonable on a long interval, since a missing backup is not going to reappear on the next poll. Keep 3 on a short one. |
What it reads
| Slot | Reading |
|---|---|
| SV01 File size | Bytes. |
| SV02 File age | Seconds since the file was last modified, measured against the probe's clock. |
| SV03 Lookup time | Milliseconds for the lookup. A share getting slow shows here first. |
| SD01 Modified | The modification time, in UTC. |
| SD02 Kind | file or directory. |
| SD03 Mount | The mount point the path is on, from the kernel's own mount table. |
| SD04 Filesystem | The mount's type, such as nfs4, cifs or fuse.sshfs. |
| SD05 Mounted from | Where the mount comes from, such as //nas/backups or nas:/volume1/backups. |
The sensor is OK while the file is found, whatever its age, until rules say otherwise. A file that is there but three weeks old is still found, so for a backup the rules below are what make the sensor worth having.
A nightly backup
- In Rules, add SV02 greater than 93600 gives CRIT. 93600 seconds is 26 hours: a day plus a margin for a backup that runs a little late.
- Add SV01 less than a sensible minimum, such as 1000, gives CRIT. A backup job that ran but wrote an empty or truncated file still updates the age, and only the size catches it.
A file that is missing altogether fills no slot, so these rules do not judge it: it is a failed reading, and the failure threshold turns it into CRIT.
Dated files
Where a device writes a new file each day with the date in its name, the path can follow it with {{isodate}} (today) or {{yesterday}}. Be careful with today. From midnight until the backup runs, today's file does not exist yet, so a sensor looking for it fails every night, and an alarm at ten past midnight soon gets ignored. Look for yesterday's file instead, and a missing one means a backup really did not run. If the backup runs just before midnight, today works, with an interval that does not poll in the gap.
Where the names cannot be predicted, because they carry a time or a serial number, point the sensor at the folder instead. A folder's modification time changes whenever a file is added to it or removed from it, so the folder's File age is the time since anything last landed there, and the same 26 hour rule works on it. Something deleted from the folder counts too, so this suits a folder that only ever receives.
When it fails
| Error | What it means |
|---|---|
| No error class | The file is not there, or, with a mount point set, the share is not mounted. The message says which. |
| AUTH | The probe's user may not see the file, or may not pass through a folder on the way to it. |
| REFUSED or CONNECTTIMEOUT | For an NFS, SMB or sshfs mount, the sensor first checks that the server named in the mount answers. These mean it did not, so the mount was not touched and nothing was left waiting on it. |
| RESPONSETIMEOUT | The lookup did not come back within the response timeout, or a lookup from an earlier poll still has not. The mount has hung. Only one lookup per path is ever left waiting on a hung mount, so the probe itself carries on. |