How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message, or ask a question in your own words.
Cisco: Catalyst, ISR, ASR and Nexus
Cisco Catalyst switches, ISR and ASR routers and Nexus switches are all best watched with SNMP. Cisco's health MIBs are old, stable and supported almost everywhere, and every failure worth an alarm (a supply, a fan, a temperature, a stack member, a BGP peer, an uplink) is one value an SNMP sensor can read. IOS-XE's RESTCONF is a good second route, and Nexus has NX-API. Meraki, being cloud managed, has its own API, covered at the end.
No Cisco MIB needs loading for any of this: every object below is given as a number to type into the OID box. Loading the MIBs (listed at the end) only adds names and labels to the MIB list.
Setting up SNMP
A read-only community, answering only the probe's address:
IOS and IOS-XE (Catalyst, ISR, ASR)
ip access-list standard OVERSIGHT-SNMP permit host <probe-address> ! snmp-server community <community> RO OVERSIGHT-SNMP snmp-server ifindex persist
On Catalyst 9300 and 9500 the last line is snmp ifmib ifindex persist instead, and it is off until set. Set it. Without it, port numbers can be reshuffled by a reload, and a sensor goes on reading a port, just not the one you meant. If a platform will not take a named access list, use a numbered one: access-list 10 permit <probe-address>, then RO 10.
NX-OS (Nexus)
ip access-list OVERSIGHT-SNMP permit ip <probe-address>/32 any ! snmp-server community <community> group network-operator snmp-server community <community> use-ipv4acl OVERSIGHT-SNMP
network-operator is the read-only role. Port numbers on NX-OS are always persistent.
In Oversight, save an SNMP credential with version 2c and the community, on the device or anywhere above it; see Credentials. Poll an in-band address such as a loopback or management VLAN interface where you can: Cisco documents the Catalyst's dedicated management port (Gi0/0) as answering only the interface and entity MIBs, so CPU, environment and stack readings may not come back through it.
Finding the row numbers
Nearly every Cisco health reading lives in a table, and the row number is not something you can guess: Cisco says outright that the environment table's numbers have no meaning of their own, and on a stack they differ from member to member. So walk each device once from any machine with net-snmp, note the numbers against what they describe, and walk again after changing hardware. Oversight itself never walks.
C=<community>; H=<device> # Ports: number and name snmpwalk -v2c -c $C $H 1.3.6.1.2.1.31.1.1.1.1 # CPU rows snmpwalk -v2c -c $C $H 1.3.6.1.4.1.9.9.109.1.1.1.1.2 # Environment: temperature, fan and power supply descriptions snmpwalk -v2c -c $C $H 1.3.6.1.4.1.9.9.13.1.3.1.2 snmpwalk -v2c -c $C $H 1.3.6.1.4.1.9.9.13.1.4.1.2 snmpwalk -v2c -c $C $H 1.3.6.1.4.1.9.9.13.1.5.1.2 # Physical parts: names, then what each is (6 power supply, 7 fan, 8 sensor) snmpwalk -v2c -c $C $H 1.3.6.1.2.1.47.1.1.1.1.7 snmpwalk -v2c -c $C $H 1.3.6.1.2.1.47.1.1.1.1.5 # Stack members snmpwalk -v2c -c $C $H 1.3.6.1.4.1.9.9.500.1.2.1.1.1 # BGP peers (the row is the peer's address) snmpwalk -v2c -c $C $H 1.3.6.1.2.1.15.3.1.2 # HSRP groups snmpwalk -v2c -c $C $H 1.3.6.1.4.1.9.9.106.1.2.1.1.15
On the device itself, show snmp mib ifmib ifindex <interface> (IOS) or show interface snmp-ifindex (NX-OS) gives a port's number.
Which MIB for which platform
Not every platform supports every MIB, and the one that matters most here is the environment MIB: Nexus and ISR 4000 do not support it. They report power supplies and fans through the FRU MIB and temperatures through the entity sensor MIB instead. From Cisco's published support lists:
| Readings | Catalyst 9300, 9500, ASR 1000 | Catalyst 9200, ISR 4000 | Nexus 9000 |
|---|---|---|---|
| Power supplies, fans | Environment MIB | FRU MIB | FRU MIB |
| Temperature | Environment MIB | Entity sensor MIB | Entity sensor MIB |
| CPU | Process MIB | Process MIB | Process MIB |
| Memory | Memory pool MIB (9300, 9500); enhanced pool MIB (ASR) | Memory pool MIB (9200); enhanced pool MIB (ISR) | Process MIB or enhanced pool MIB |
| Stack | Stackwise MIB | Stackwise MIB (9200) | n/a |
| BGP | ASR: BGP MIBs. Catalyst: not listed, so test first or use RESTCONF | ISR: BGP MIBs | BGP MIBs |
SNMP sensors worth having
All are GAUGE, as read, unless the table says otherwise. <row> is the number the walk found; <if> is a port's number; <part> is a physical part's number from the entity walk.
| What | OID | Settings | Rules |
|---|---|---|---|
| Uplink up | 1.3.6.1.2.1.2.2.1.8.<if> | Labels {"1":"up","2":"down","7":"lowerLayerDown"} | SV01 not equal to 1 gives CRIT |
| Uplink errors | 1.3.6.1.2.1.2.2.1.14.<if> | COUNTER32, change per minute | Greater than 10 gives WARN, greater than 100 gives CRIT |
| Port shut down by the switch | 1.3.6.1.4.1.9.9.276.1.1.2.1.9.<if> | Reads 2 while healthy | Equal to (5,143,157,164,165) gives CRIT: err-disabled, link flapping, channel error, BPDU guard, 802.1X violation |
| Reloaded | 1.3.6.1.2.1.1.3.0 | Multiplier 0.01, unit s | Less than 600 gives WARN |
| CPU | 1.3.6.1.4.1.9.9.109.1.1.1.1.8.<row> | Unit %. A five-minute average already. | Greater than 80 gives WARN, greater than 90 gives CRIT |
| Temperature state (environment MIB) | 1.3.6.1.4.1.9.9.13.1.3.1.6.<row> | Labels {"1":"normal","2":"warning","3":"critical","4":"shutdown","5":"notPresent","6":"notFunctioning"} | Equal to 2 gives WARN; equal to (3,4,6) gives CRIT |
| Fan (environment MIB) | 1.3.6.1.4.1.9.9.13.1.4.1.3.<row> | Same labels | Equal to 2 gives WARN; equal to (3,4,6) gives CRIT |
| Power supply (environment MIB) | 1.3.6.1.4.1.9.9.13.1.5.1.3.<row> | Same labels | Equal to (3,4,6) gives CRIT. 5 is an empty bay, normal on many units. |
| Power supply (FRU MIB) | 1.3.6.1.4.1.9.9.117.1.1.2.1.2.<part> | 2 is on and healthy | Equal to (9,12) gives WARN (on, but a fan or inline power has failed); not equal to (2,9,12) gives CRIT |
| Fan tray (FRU MIB) | 1.3.6.1.4.1.9.9.117.1.4.1.1.1.<part> | 2 is up | Equal to 4 gives WARN; equal to 3 gives CRIT |
| Sensor failed (entity sensor MIB) | 1.3.6.1.4.1.9.9.91.1.1.1.1.5.<part> | 1 ok, 3 not working | Equal to 3 gives CRIT |
| Temperature over threshold (entity sensor MIB) | 1.3.6.1.4.1.9.9.91.1.2.1.1.5.<part>.<threshold> | 1 means crossed, 2 not | Equal to 1 gives CRIT. Pick the threshold rows whose severity (...91.1.2.1.1.2) is major (20) or critical (30). |
| Stack ring complete | 1.3.6.1.4.1.9.9.500.1.1.3.0 | 1 true, 2 false | Equal to 2 gives CRIT: a stack cable or member is lost and the stack is running as a chain |
| Stack member | 1.3.6.1.4.1.9.9.500.1.2.1.1.6.<part> | 4 is ready | Not equal to 4 gives CRIT |
| BGP peer (IPv4) | 1.3.6.1.2.1.15.3.1.2.<peer address>, such as ...15.3.1.2.192.0.2.1 | 6 is established | Not equal to 6 gives CRIT |
| BGP peer (IPv6, or in a VRF) | 1.3.6.1.4.1.9.9.187.1.2.5.1.3.1.4.<IPv4>, or ...3.2.16.<16 bytes> for IPv6 | 6 is established | Not equal to 6 gives CRIT |
| HSRP group | 1.3.6.1.4.1.9.9.106.1.2.1.1.15.<if>.<group> | 6 active, 5 standby | Not equal to the state this box should be in (6 on the primary, 5 on the secondary) gives CRIT |
An IPv6 address in the BGP OID is its 16 bytes in decimal: 2001:db8::1 is 32.1.13.184.0.0.0.0.0.0.0.0.0.0.0.1.
Memory
Cisco reports memory as bytes used and bytes free, as two separate readings, and a sensor reads one. So work out the pool's total once (used plus free), then put a single sensor on free with a fixed threshold, such as CRIT below 10% of that total:
- Catalyst 9200, 9300, 9500:
1.3.6.1.4.1.9.9.48.1.1.1.6.1, the processor pool's free bytes, multiplier 0.000001 for megabytes. - ISR, ASR and Nexus:
1.3.6.1.4.1.9.9.221.1.1.1.1.20.<part>.<pool>, typed COUNTER64, as read. It is a level in a counter's clothing: never set it to a rate. - Nexus supervisor:
1.3.6.1.4.1.9.9.109.1.1.1.1.13.<row>, free memory in kilobytes.
Cisco also defines a ready-made memory percentage (1.3.6.1.4.1.9.9.48.1.2.1.2.<pool>), but support for it has not been confirmed on any platform. Try it; if the device says it has no such object, use free bytes.
When a row disappears
If a device answers that it has no such object, Oversight treats that as a settings mistake, not a failure: the sensor is suspended and shows UNKNOWN until it is saved again, and nobody is woken. That is right for a typing error, but it matters for tables whose rows can vanish:
- Err-disabled ports. Cisco's err-disable MIB only has a row while a port is err-disabled, so a healthy port reads as a missing object. That is why the table above uses the interface extension MIB's shut-down cause, which is always there.
- Stack members and BGP peers may disappear rather than read removed or idle if a member is pulled out or a neighbour is deleted. Pair member sensors with Stack ring complete, which does fire when a member goes, and treat a BGP sensor that turns UNKNOWN as a peer that has gone.
A starting set
| Device | Sensors |
|---|---|
| Catalyst switch or stack | Reloaded; CPU; memory free; each power supply; fans; temperature state; stack ring and each member (stacks); each uplink or port channel, up and errors; HSRP on a distribution pair |
| ISR or ASR router | Reloaded; CPU; memory free (enhanced pool); power supplies (FRU MIB on ISR 4000); sensor failed and temperature over threshold; each WAN, up and errors; each BGP peer; HSRP or VRRP where used |
| Nexus | Reloaded; CPU; memory free; power supplies and fan trays (FRU MIB); temperature over threshold; modules on a chassis (1.3.6.1.4.1.9.9.117.1.2.1.1.2.<part>, not equal to 2 gives CRIT); uplinks, port channels and the vPC peer link, up and shut-down cause; each BGP peer |
Plus a Ping of each, which is the cheapest way to hear that a device has gone altogether. At the default 60 seconds from one probe group, each SNMP sensor costs about £1.73 a month.
RESTCONF on IOS-XE
IOS-XE 16 and 17 (Catalyst 9000, ISR 4000, ASR 1000, Catalyst 8000) answer RESTCONF, which reads the same health as named values, in words. It suits readings that SNMP does not give on a platform, BGP on a Catalyst in particular. To switch it on:
aaa new-model aaa authentication login default local aaa authorization exec default local username oversight privilege 15 secret <password> ip http secure-server ip http authentication local restconf
RESTCONF only accepts a privilege 15 user, and there is no read-only equivalent. The credential Oversight holds is then an administrator's login on that device, so weigh that before choosing it over SNMP, and restrict who can reach the device's HTTPS service. The certificate is self-signed unless one is installed, so turn off Verify the certificate on the sensor.
Save an HTTP(S) credential with the username and password, and set each sensor to https, port 443, GET, with these headers:
Authorization: Basic {{basicauth}}
Accept: application/yang-data+json
Ask for the single value, not the list it sits in, and the reply is one named value, which a JSON row reads by that name. Paths follow /restconf/data/:
| What | Path | Extraction | Rules |
|---|---|---|---|
| CPU, five minutes | Cisco-IOS-XE-process-cpu-oper:cpu-usage/cpu-utilization/five-minutes | SV03, JSON, Cisco-IOS-XE-process-cpu-oper:five-minutes | Greater than 80 gives WARN, greater than 90 gives CRIT |
| Port up | Cisco-IOS-XE-interfaces-oper:interfaces/interface=GigabitEthernet1%2F0%2F1/oper-status | SD03, JSON, Cisco-IOS-XE-interfaces-oper:oper-status | SD03 not equal to if-oper-state-ready gives CRIT |
| BGP neighbour | Cisco-IOS-XE-bgp-oper:bgp-state-data/neighbors/neighbor=ipv4-unicast,default,192.0.2.1/session-state | SD03, JSON, Cisco-IOS-XE-bgp-oper:session-state | SD03 not equal to fsm-established gives CRIT |
| Stack member | Cisco-IOS-XE-stack-oper:stack-oper-data/stack-node=1/node-state | SD03, JSON, Cisco-IOS-XE-stack-oper:node-state | SD03 not equal to state-ready gives CRIT |
| Stack ring | Cisco-IOS-XE-stack-oper:stack-oper-data/stack-info/ring-status | SD03, JSON, Cisco-IOS-XE-stack-oper:ring-status | SD03 equal to half-ring gives CRIT |
With SV01 not equal to 200 gives CRIT on every one. Some notes:
- A slash in an interface name must be sent as
%2F:GigabitEthernet1/0/1becomesGigabitEthernet1%2F0%2F1. - The BGP path names the address family, the VRF and the neighbour, in that order. The global table's VRF is believed to be
default; to be sure, fetch.../bgp-state-data/neighborsonce with Store raw response ticked and read the names it uses. - Environment readings are under
Cisco-IOS-XE-environment-oper:environment-sensors, keyed by each sensor's name and location. Their names and state words differ by platform, so fetch the whole list once, read it, then build sensors on the ones that matter. - Large numbers arrive as text, such as
"1279097", which a numeric slot still takes as a number.
NX-API on Nexus
NX-API runs show commands and returns JSON, and is on by default on HTTPS from NX-OS 9.2(1) (feature nxapi, and nxapi use-vrf management to answer on mgmt0). A sensor POSTs to /ins with Authorization: Basic {{basicauth}}, Content-Type: application/json and a body such as:
{"ins_api":{"version":"1.0","type":"cli_show","chunk":"0","sid":"1",
"input":"show interface ethernet1/1","output_format":"json"}}
Paths then start ins_api.outputs.output.body., so a port's state is ins_api.outputs.output.body.TABLE_interface.ROW_interface.state, reading up or down. It works, but SNMP is the better route on Nexus: the JSON's shape changes between platforms and releases, a list of one comes back as a single item rather than a list, and positions in a list are not stable. Keep NX-API for something SNMP cannot tell you, and where a list must be searched, use a REGEX row on the whole reply, for instance a BGP neighbour's state:
/"neighborid":\s*"192\.0\.2\.1"[^}]*?"state":\s*"(\w+)"/
into a text slot, with not equal to Established gives CRIT.
Meraki
Meraki kit is managed from Cisco's cloud, and the Dashboard API is the way to watch it. Make an API key in the Meraki dashboard, save it in an HTTP(S) credential as the bearer token, and point an HTTPS sensor at api.meraki.com with Authorization: Bearer {{bearer}}:
| What | Path | Extraction | Rules |
|---|---|---|---|
| One device | /api/v1/organizations/<org>/devices/statuses?serials[]=<serial> | SD03, JSON, 0.status | Equal to alerting gives WARN; equal to offline gives CRIT |
| Anything offline | /api/v1/organizations/<org>/devices/statuses/overview | SV03, JSON, counts.byStatus.offline | Greater than 0 gives CRIT |
Meraki allows ten requests a second per organisation, far more than a sensor a minute per device needs. Because the API is in the cloud, use a Target override with the whole URL, since the device's own address is not where the request goes.
ASA and Firepower
ASA and FTD support the process, memory pool, enhanced pool and FRU MIBs, so CPU, memory and power supplies are read as above. The environment and entity sensor MIBs are not supported. Failover role is 1.3.6.1.4.1.9.9.147.1.2.1.1.1.3.6 (primary unit) and .7 (secondary unit): 9 is active and 10 is standby, so not equal to the state that unit should be in gives CRIT. SNMP on an ASA is granted per host and interface (snmp-server host <interface> <probe-address> poll community <community>), and on FTD from the management centre's platform settings.
Things that catch people out
- Port numbers move after a reload without index persistence (above). For an important uplink, a second sensor on
1.3.6.1.2.1.31.1.1.1.1.<if>, TEXT, with SD01 not equal to the port's name gives CRIT, catches the number moving. - Traffic counters. The 32-bit octet counters wrap in about 34 seconds at 1 Gbit/s. If you chart traffic, use
1.3.6.1.2.1.31.1.1.1.6.<if>and.10.<if>, COUNTER64, change per second, multiplier 8 for bit/s. Traffic is a figure, not an alarm. - Uptime is the SNMP agent's, not strictly the chassis's, and it wraps after about 497 days. A drop means a reload or an agent restart, both worth a look.
- NX-API's CPU figure is idle, not busy, and instantaneous. Use the SNMP five-minute figure.
MIBs
None are needed to poll. To pick objects by name, import these through Files, in this order, from github.com/cisco/cisco-mibs (the v2 folder):
- Foundations: CISCO-SMI, CISCO-TC, ENTITY-MIB, ENTITY-SENSOR-MIB, INET-ADDRESS-MIB, HCNUM-TC and RFC1213-MIB where not already loaded.
- CISCO-QOS-PIB-MIB, needed only because the memory pool MIB depends on it.
- Health: CISCO-PROCESS-MIB, CISCO-MEMORY-POOL-MIB, CISCO-ENHANCED-MEMPOOL-MIB, CISCO-ENVMON-MIB, CISCO-ENTITY-SENSOR-MIB, CISCO-ENTITY-FRU-CONTROL-MIB.
- Features: CISCO-STACKWISE-MIB, BGP4-MIB, CISCO-BGP4-MIB, CISCO-HSRP-MIB, CISCO-IF-EXTENSION-MIB, and CISCO-FIREWALL-MIB for ASA.
Leave out CISCO-ERR-DISABLE-MIB, for the reason above, and because it pulls in several others. Cisco publishes what each platform supports at cisco.github.io/cisco-mibs/supportlists/.