How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message, or ask a question in your own words.
Juniper: EX, SRX, MX and QFX
Juniper's EX and QFX switches, SRX firewalls and MX routers all run Junos, and share one SNMP agent and one set of health readings. SNMP covers most of what matters, and one pair of readings does most of the work: Junos raises its own red and yellow chassis alarms when a power supply fails, a fan stops or something overheats, and SNMP gives the number of each that are active. The Junos REST API fills the gaps SNMP cannot: an SRX cluster's failover state, and Virtual Chassis members going missing.
Setting up SNMP
set snmp community oversightRO authorization read-only set snmp community oversightRO clients <probe-address>/32 set snmp community oversightRO clients 0.0.0.0/0 restrict
The last line matters: with no client list a community answers anyone, and 0.0.0.0/0 restrict refuses everyone not listed above it. A few more points:
- A firewall filter on lo0 protecting the Routing Engine must allow UDP 161 from the probe.
- An SRX polled in-band needs
host-inbound-traffic system-services snmpon the zone or interface. The management port, fxp0, is outside the zone model and does not. - A management VRF (
mgmt_junos) should need nothing more: Juniper treats its clients as if they were in the default instance. If polls over it time out, addset snmp routing-instance-accessandset snmp community oversightRO routing-instance mgmt_junos clients <probe-address>/32. filter-interfacesis not access control. It only hides interfaces from the reply.
In Oversight, save an SNMP credential with version 2c and the community, on the device or anywhere above it; see Credentials.
Finding the row numbers
Juniper's component table is indexed by four numbers, such as 9.1.0.0 for the Routing Engine on a single-engine box. Anything else, a second engine, a power supply, a line card, a cluster's second node, needs a one-off walk to find its numbers:
snmpwalk -v2c -c oversightRO -On <device> 1.3.6.1.4.1.2636.3.1.13.1.5
Each line ends with the four numbers and names the part, for example ...13.1.5.9.1.0.0 = "Routing Engine" or ...13.1.5.2.1.0.0 = "Power Supply A". The first number is the kind of part (2 power supplies, 4 fans, 7 line cards, 9 Routing Engines), but check the walk rather than working it out: on an SRX cluster, node1's parts are renumbered, and a second Routing Engine has its own number. On the device itself, show snmp mib walk jnxOperatingDescr does the same, and show interfaces <name> | match "SNMP ifIndex" gives a port's number.
Check the number before relying on it. Juniper's component readings are 0 wherever they do not apply: a fan or power supply always reads 0% CPU. A CPU sensor pointed at the wrong row reads a healthy-looking 0 forever.
SNMP sensors worth having
All GAUGE, as read, unless the table says otherwise.
| What | OID | Settings | Rules |
|---|---|---|---|
| Red alarms active | 1.3.6.1.4.1.2636.3.4.2.3.2.0 | At least 1 gives CRIT | |
| Yellow alarms active | 1.3.6.1.4.1.2636.3.4.2.2.2.0 | At least 1 gives WARN | |
| Routing Engine CPU | 1.3.6.1.4.1.2636.3.1.13.1.8.9.1.0.0 | Unit % | Greater than 85 gives WARN, greater than 95 gives CRIT |
| Routing Engine memory | 1.3.6.1.4.1.2636.3.1.13.1.11.9.1.0.0 | Unit % | Greater than 80 gives WARN, greater than 95 gives CRIT |
| A part's state | 1.3.6.1.4.1.2636.3.1.13.1.16.<part> | 1 running, 2 standby, 3 ready, 4 fan at full speed, 5 reset, 6 down, 7 unknown | At least 4 gives WARN, at least 5 gives CRIT |
| Uplink up | 1.3.6.1.2.1.2.2.1.8.<port> | Not equal to 1 gives CRIT | |
| Rebooted | 1.3.6.1.2.1.1.3.0 | Multiplier 0.01, unit s | Less than 600 gives WARN |
The alarm counts are the ones to use, not the red and yellow lamp states beside them in the MIB: the lamps can be silenced with the front-panel alarm cut-off button, the counts cannot. Temperatures, power supplies and fans are covered by the counts, since Junos raises an alarm for each. Add a part's state sensor only for something you want named in the alarm on its own. That OID uses Juniper's ordered state, which is arranged so one pair of rules suits any part.
The CPU figure jumps about from poll to poll, and a single busy moment is not worth an alarm. Where the platform fills it in, read the five-minute average at 1.3.6.1.4.1.2636.3.1.13.1.24.9.1.0.0 instead, with the same thresholds; if that reads 0, the platform does not, and the figure above is the one to use.
SRX: two CPUs
An SRX has a control plane, the Routing Engine, and a data plane that carries the traffic. A firewall can be dropping traffic with an idle Routing Engine, so watch the data plane too:
| What | OID | Settings | Rules |
|---|---|---|---|
| Data plane CPU | 1.3.6.1.4.1.2636.3.39.1.12.1.1.1.4.<spu> | Unit % | Greater than 85 gives WARN, greater than 95 gives CRIT |
| Data plane memory | 1.3.6.1.4.1.2636.3.39.1.12.1.1.1.5.<spu> | Unit % | Greater than 80 gives WARN, greater than 95 gives CRIT |
| Sessions in use | 1.3.6.1.4.1.2636.3.39.1.12.1.2.0 | Unit sessions | 80% of the maximum gives WARN, 90% gives CRIT |
Find the <spu> numbers with snmpwalk -v2c -c oversightRO -On <srx> 1.3.6.1.4.1.2636.3.39.1.12.1.1.1.4. For the session thresholds, read the maximum once from 1.3.6.1.4.1.2636.3.39.1.12.1.3.0 and work out the two figures, since a rule compares with a number, not with another reading.
On branch SRX models the Routing Engine memory figure includes the data plane's memory, so it reads higher than the engine's own use. The 80 and 95 thresholds are then on the cautious side, which is the safe way round.
Dual Routing Engines
On an MX or EX with two Routing Engines, a sensor on 1.3.6.1.4.1.2636.3.1.14.1.7.<engine> gives each one's role: 2 master, 3 backup. Not equal to 2 gives WARN on the engine that should be master tells you a switchover has happened. The engine's four numbers come from the walk.
BGP
Juniper's BGP table is at 1.3.6.1.4.1.2636.5.1.1.2.1.1.1.2.<peer>, where 6 is established: not equal to 6 gives CRIT. Its row number is awkward to build by hand, so copy it from a walk:
snmpwalk -v2c -c oversightRO -On <router> 1.3.6.1.4.1.2636.5.1.1.2.1.1.1.2
By default Junos leaves out a length byte the standard says should be there, so an IPv4 peer 192.0.2.2 from local 192.0.2.1 in the main routing instance is ...1.1.1.2.0.1.192.0.2.1.1.192.0.2.2. Tools that decode the row by name get it wrong, hence -On. If someone later sets protocols bgp snmp-options emit-inet-address-length-in-oid, every BGP sensor's OID changes. For IPv4 peers in the default instance, the standard 1.3.6.1.2.1.15.3.1.2.<peer address> is simpler, with the same rule.
Setting up the REST API
Junos answers REST requests for any operational command, which is how an SRX cluster's failover state and Virtual Chassis membership are read. Turn it on over HTTPS with a self-signed certificate, limited to the probe, with a read-only user:
request security pki generate-key-pair certificate-id oversight-rest request security pki local-certificate generate-self-signed certificate-id oversight-rest domain-name <name> subject "CN=<name>" set system services rest https port 3443 set system services rest https server-certificate oversight-rest set system services rest control allowed-sources [ <probe-address> ] set system login user oversight class read-only authentication plain-text-password
Where management is in mgmt_junos, add set system services rest routing-instance mgmt_junos (Junos 20.3R1 and later). On an SRX, use the management port: there is no in-band rest service for a zone.
In Oversight, save an HTTP(S) credential with the username and password, and give each REST sensor https, port 3443, GET, Verify the certificate off, and these headers:
Authorization: Basic {{basicauth}}
Accept: application/xml
The path is /rpc/ and the command's name, below. Add SV01 not equal to 200 gives CRIT to each.
Why XML, not JSON
Junos can answer in JSON, but for monitoring its XML is more dependable. Junos JSON drops readings that appear more than once in a group (an SRX cluster's two nodes are exactly that), and where there is nothing to report, such as no alarms, it leaves the field out rather than giving 0, so a sensor cannot tell. With XML, XPATHCOUNT counts matching entries and gives 0 when there are none. Write the expressions with local-name(), as below: Junos puts its release number in the XML namespace, so an expression that names the namespace breaks at the next upgrade.
REST sensors worth having
| What | Path | Extraction | Rules |
|---|---|---|---|
| Major chassis alarms | /rpc/get-alarm-information | SV03, XPATHCOUNT, //*[local-name()='alarm-detail'][*[local-name()='alarm-class']='Major'] | At least 1 gives CRIT |
| Hardware not OK | /rpc/get-environment-information | SV03, XPATHCOUNT, //*[local-name()='environment-item'][*[local-name()='status']!='OK' and *[local-name()='status']!='Absent'] | At least 1 gives CRIT |
| BGP peers not established | /rpc/get-bgp-summary-information | SV03, XPATHCOUNT, //*[local-name()='bgp-peer'][*[local-name()='peer-state']!='Established'] | At least 1 gives CRIT |
| Virtual Chassis member missing | /rpc/get-virtual-chassis-information | SV03, XPATHCOUNT, //*[local-name()='member'][*[local-name()='member-status']!='Prsnt'] | At least 1 gives CRIT |
| SRX cluster node lost | /rpc/get-chassis-cluster-status | SV03, XPATHCOUNT, //*[local-name()='redundancy-group-status'][.='lost' or .='disabled' or .='ineligible' or .='unavailable'] | At least 1 gives CRIT |
| SRX cluster failed over | /rpc/get-chassis-cluster-status | SD03, XPATH, below | SD03 not equal to primary gives WARN |
The failover expression reads node0's role in redundancy group 1, the one carrying traffic:
//*[local-name()='redundancy-group'][*[local-name()='redundancy-group-id']='1']/*[local-name()='device-stats']/*[local-name()='redundancy-group-status'][preceding-sibling::*[local-name()='device-name'][1]='node0']
Change '1' to '0' for the control plane's group, and node0 to whichever node should be primary.
- A peer shut down on purpose shows as Idle and is counted. Remove it from the configuration, or watch important peers one by one with an XPATH row:
//*[local-name()='bgp-peer'][*[local-name()='peer-address']='192.0.2.2']/*[local-name()='peer-state'], into a text slot, not equal to Established gives CRIT. - Virtual Chassis membership has no SNMP state, only each member's role, and a member that leaves may simply vanish from the table. Oversight treats a row that is not there as a settings fault, suspending the sensor rather than raising an alarm. The REST sensor above keeps a missing member listed as
NotPrsnt, so it alarms properly. - SRX cluster state cannot be read by SNMP at all. Juniper's cluster MIB is traps only. Poll the node that is primary for redundancy group 0: its answer covers both nodes.
- Alarm classes: the minor count is
alarm-class='Minor', if you want a WARN for those too. System alarms (show system alarms) include housekeeping such as Rescue configuration is not set; clear those withrequest system configuration rescue saveor leave them out. - If someone has set
system export-format state-data json compact, it changes JSON only; the XML above is unaffected.
A starting set
| Device | Sensors |
|---|---|
| EX and QFX | Red and yellow alarm counts; Routing Engine CPU and memory; each uplink up; rebooted; on a Virtual Chassis, members missing (REST) |
| SRX | Red and yellow alarm counts; Routing Engine CPU; data plane CPU for each SPU; sessions in use; Routing Engine memory; the internet-facing interface up; on a cluster, node lost and failed over (REST) |
| MX | Red and yellow alarm counts; Routing Engine CPU and memory; with two engines, which is master; each important BGP peer; core links up |
At the default 60 seconds from one probe group, an SNMP sensor costs about £1.73 a month and a REST sensor about £3.46.
Things that catch people out
- Zero means not available on every Juniper component reading, as above. Confirm the row is the part you meant.
- Uptime is the SNMP agent's. A restart of the agent resets it without a reboot. The Routing Engine's own uptime is in
/rpc/get-route-engine-information. - Logical systems on an SRX may need the community given as
default@<community>for some readings. - Mist-managed EX and SRX still run Junos, so SNMP and REST work as above, but Mist owns the configuration and may overwrite changes made on the device. Add them through Mist's additional CLI commands.
MIBs
None are needed to poll. To pick objects by name, import these through Files, in this order: JUNIPER-SMI, JUNIPER-MIB, JUNIPER-ALARM-MIB, JUNIPER-EX-SMI, JUNIPER-VIRTUALCHASSIS-MIB, JUNIPER-JS-SMI, JUNIPER-SRX5000-SPU-MONITORING-MIB, JUNIPER-EXPERIMENT-MIB and BGP4-V2-MIB-JUNIPER. Juniper's MIB Explorer (apps.juniper.net/mib-explorer) has the current set for each Junos release; the single files on Juniper's documentation site are older revisions. The SPU MIB's name is easy to get wrong: it is JUNIPER-SRX5000-SPU-MONITORING-MIB, even though it covers branch SRX models too.