How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
Some things can only be reached through something else. The Linux and Windows agents on remote servers are polled over the mesh, so when the mesh manager goes, every one of them loses its link. A remote site sits behind its Internet uplink, so when the uplink goes, everything at the site goes with it. Without a dependency each of those raises its own alarm. With one, the thing that actually failed alarms, and everything cut off behind it goes quiet and says why.
Setting one
Open the site, group, device or sensor that cannot be reached without something else, and pick that something in its Dependencies panel. It can be a site, group, device or sensor anywhere in the same estate, and an object can have as many as it needs.
A dependency applies to the object it is set on and everything beneath it, in the same way as a notification binding. A device inside a group that depends on the uplink, and which itself depends on the mesh, depends on both. The panel lists what reaches it from above as well as its own.
Pointing at a device is usually best: its rollup already combines its sensors, so a mesh manager watched by a port check, its API and its container is down when the device says so.
What happens when it goes down
A dependency is down while it is WARN or CRIT. Then, for every sensor beneath the object that depends on it:
- A poll that cannot reach its target is held, not failed. Those are the polls that fail with DNS, REFUSED, CONNECTTIMEOUT or RESPONSETIMEOUT. The sensor goes HELD from its first such poll rather than counting towards its failure threshold.
- A reading that does arrive is judged as usual. A mesh manager that has stopped does not break tunnels that are already up, so agents that can still be reached go on reporting, and a real alarm from one of them still comes through. A login refused or a certificate failing is the target's own answer, and is never held.
- Anything that alarmed first is held after the fact. An agent often loses its link a minute before the mesh check notices. It goes from CRIT to HELD when the dependency goes down, and if its alarm had not been sent yet, it never is.
HELD
HELD is its own state, in sand on the dashboard, so an outage shows as one red object and a spread of sand behind it rather than a wall of red.
- It never alarms and never sounds. The dependency is the alarm, and the held object's detail names it: Held: Mesh manager is CRIT, so it cannot be reached.
- It has no opinion in a rollup, like UNKNOWN. A device whose sensors are all held is HELD itself, and so on up the tree; a device with one sensor held and one still reading takes the reading.
- It is not downtime. Availability counts held time as unknown, and the downtime belongs to whatever held it.
- Held is transitive. If the mesh manager depends on the data centre's edge, and the edge fails, the mesh manager is held by the edge, and so is every agent behind the mesh. One alarm: the edge.
Coming back
When a dependency recovers, what is behind it is given time before it may alarm. First the dependency goes on counting as down for 2 minutes, while tunnels and routes re-form. After that, a sensor still failing to connect starts its failure threshold from nothing and stays HELD until it has failed that many times in a row, so a sensor polled every minute with a threshold of 3 has about five minutes in all. The first reading that arrives clears it.
Where there are several dependencies, all of them must be back. A remote site whose uplink returns while the mesh is still down stays held by the mesh, with no alarm for the agents in between.
Rules
- Something inside what it depends on ignores that dependency. A remote site's group can depend on the uplink device inside it: everything else in the group is held when the uplink fails, and the uplink's own sensors still alarm, because that failure is the one you need to hear about. The panel marks a dependency reaching an object this way as not applied there.
- No loops. A dependency that leads back, directly or further along, to the object or something containing it is refused.
- Nothing can depend on its own device, group or site. Its failure is already part of that state.
- Paused, suspended, draft and scheduled off dependencies say nothing, so pausing the mesh manager for maintenance does not hold everything behind it.
Two ways in
Where a site has two uplinks and either is enough, put both uplink devices in a group of their own, leave its ANY at OK and its ALL at CRIT, and depend on the group. One uplink down leaves the group OK and nothing is held; both down makes it CRIT and the site goes quiet.
A worked example
| Object | Depends on |
|---|---|
| Remote agents group | Mesh manager device |
| Branch office site's group | Branch uplink device, inside that group |
| Mesh manager device | Data centre edge device |
The mesh manager fails: it alarms, and each agent goes HELD as it loses its link. The branch uplink then fails too: it alarms, and everything else at the branch goes HELD. The uplink comes back: the branch recovers, except its agents, which stay held by the mesh with nothing said. The edge fails: it alarms, and the mesh manager and everything behind it are held by the edge.