How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
A ping sensor asks the simplest question there is: is something at this address answering on the network? The probe sends one ICMP echo request to the device and times the reply. It is the cheapest sensor, at 1 credit a reading, and the right first sensor on almost any device.
What it cannot tell you is whether anything on the device is working. A server whose web site, database or mail has stopped will usually still ping. Put a ping on the device to show it is reachable, and a service sensor (HTTP(S), TCP, SMTP and so on) beside it for each thing that matters.
Before you start
- The device must answer ICMP echo from the probe. Many firewalls drop it. Windows drops it by default: enable the inbound rule File and Printer Sharing (Echo Request - ICMPv4-In) for the profile the machine is on, or allow echo requests from the probe's address. On a firewall or router, allow ICMP echo request in, and echo reply out, from the probe groups' addresses only.
- IPv4 only. A hostname is looked up on the probe, so the probe's own DNS decides which address is pinged. A name with no IPv4 address cannot be pinged at all, and the sensor goes UNKNOWN until it is corrected.
- No credential is needed. Nothing is logged into, so there is nothing to set up on the device beyond letting ICMP through.
If a device is up but will never answer a ping, and you cannot change that, use a TCP sensor on a port it does listen on instead.
Setting it up
| Setting | What to put |
|---|---|
| Target | Leave the target override empty and the ping goes to the device's own address. Set it only to ping something else, such as the far end of a link. |
| Interval | 60 seconds suits most devices. Down to 10 seconds for something where a minute matters, such as a core switch or a firewall. |
| Response timeout | How long the probe waits for the reply. 1 to 2 seconds on a local network, 3 to 5 across the internet or a slow link. Waiting longer does not make a ping more likely to succeed; it only makes a dead device slower to notice. |
| Connect timeout | Not used, since a ping makes no connection. It still counts towards the check that connect plus response timeout is less than the interval, so on a short interval set it low, for example 1. |
| Failures before a probe calls it down | 3 is the default and a good one. Each poll sends a single packet with no retry, so one lost packet is one failure. A threshold of 1 turns every dropped packet into an alarm. |
| Probe groups | With more than one, the sensor's Rollup settings decide what it is when the groups disagree. By default a ping in two groups only goes CRIT when both lose the device, which is what you want for "is it up". Set ANY to WARN or CRIT on the sensor's Rollup to hear as soon as one location loses sight of it. |
What it reads
Out of the box a ping fills no slots, so there is nothing for a rule to test. The sensor is OK while the device replies and CRIT once it has failed the threshold number of times in a row. A slow reply is still a reply: a ping that comes back in 4 seconds is OK.
The round trip time is always recorded as the response time. To draw it, pick a chart under Chart, response time on the Polling panel.
Warning on slow replies
To raise an alarm when replies are slow as well as when they stop, put the round trip time into a slot and write rules on it:
- In Extraction, give SV01 the method DIRECT, the name Round trip and the unit ms. DIRECT takes the response time as the value, so no expression is needed.
- In Rules, add SV01 greater than 150 gives WARN and SV01 greater than 500 gives CRIT.
Rules are checked worst first and the worst match wins, so it does not matter which order the two are entered in. A reply of 200 ms is WARN and one of 800 ms is CRIT. Pick the numbers from what the chart shows is normal for that device: a few milliseconds on a local network, tens of milliseconds across the country. A failed ping fills no slot, so these rules only judge replies that arrived; losing the device altogether is still handled by the failure threshold.
When it fails
| Error | What it means |
|---|---|
| RESPONSETIMEOUT | No reply within the response timeout. The usual failure: the device is off or unreachable, or a firewall is dropping ICMP. |
| REFUSED | A router answered that the address is unreachable, or the packet's time to live ran out on the way. This arrives well before the timeout. It is also what a probe reports when it cannot open an ICMP socket at all, which is the probe's problem rather than the device's: if every ping sensor on one probe shows REFUSED at once, look at the probe. |
| DNS | The hostname could not be looked up on the probe. |