How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message, or ask a question in your own words.
Arista: EOS switches through eAPI and SNMP
Arista EOS switches are best watched through eAPI, the switch's own JSON interface. One HTTPS request runs one or more show commands and returns the answer as structured JSON, so power supplies, fans, temperature, MLAG, BGP and err-disabled ports all come back as plain words a rule can test. MLAG in particular has no SNMP MIB at all, so eAPI is the only way to watch it. SNMP is the fallback where eAPI is not allowed, and the natural fit for a single port's state.
Setting up eAPI
On the switch, turn eAPI on over HTTPS, in the management VRF if that is how the probe reaches it, and make a read-only user:
management api http-commands
protocol https
no shutdown
!
vrf MGMT
no shutdown
!
username oversight privilege 1 role network-operator secret <password>
Use the site's own VRF name in place of MGMT; it is often management. eAPI answers only in the VRFs it is told to, so if the probe reaches the switch through Management1 in a VRF and that block is missing, the sensor simply times out. network-operator is EOS's built-in read-only role, which allows show commands and blocks the shell. To limit who can reach eAPI at all, add ip access-group <acl> inside the VRF block with an access list permitting only the probe. show management api http-commands confirms it is running and where.
In Oversight, save an HTTP(S) credential with the username and password, then give each eAPI sensor these settings:
| Setting | Value |
|---|---|
| Scheme, port, path | https, 443, /command-api |
| Method | POST |
| Headers | Authorization: Basic {{basicauth}}Content-Type: application/json |
| Request body | The commands, below |
| Verify the certificate | Off, unless the switch has been given a proper certificate with an SSL profile |
The request body names the commands to run:
{"jsonrpc":"2.0","method":"runCmds","params":{"version":1,
"cmds":["show system environment power","show system environment cooling",
"show system environment temperature"],"format":"json"},"id":"oversight"}
The answers come back in order as result.0, result.1, result.2, so one request, and one sensor, can cover several things. Keep "version":1: it fixes the shape of each command's answer, where latest may rename fields after an EOS upgrade and quietly break your paths.
And give every eAPI sensor the guard rule SV01 not equal to 200 gives CRIT, so a refused login raises an alarm rather than reading OK with nothing in it.
eAPI sensors worth having
eAPI reports states as words, such as ok, active and Established, so read them into text slots, SD03 onwards, and test them with not equal to. The paths in the tables have been checked against real switch output.
Environment: one sensor for power, cooling and temperature
With the three-command body above:
| Reading | Slot and path | Rule |
|---|---|---|
| Power supply 1 | SD03, JSON, result.0.powerSupplies.1.state | SD03 not equal to ok gives CRIT |
| Power supply 2 | SD04, JSON, result.0.powerSupplies.2.state | SD04 not equal to ok gives CRIT |
| Cooling | SD05, JSON, result.1.systemStatus | SD05 not equal to coolingOk gives CRIT |
| Temperature | SD06, JSON, result.2.systemStatus | SD06 not equal to temperatureOk gives CRIT |
That one sensor catches a failed or unplugged supply (a supply without power reads powerLoss), a failed fan and overheating. The switch applies its own per-sensor temperature limits before it says temperatureOk, so there are no thresholds to choose. Power supplies are listed by slot number, which is why the path says .1 and .2 rather than counting from 0.
MLAG
Only on switches that run MLAG: on one that does not, show mlag says disabled and little else, and the sensor would find nothing to read. Body with "cmds":["show mlag"]:
| Reading | Slot and path | Rule |
|---|---|---|
| MLAG state | SD03, JSON, result.0.state | Not equal to active gives CRIT |
| Peer | SD04, JSON, result.0.negStatus | Not equal to connected gives CRIT |
| Peer link | SD05, JSON, result.0.peerLinkStatus | Not equal to up gives CRIT |
| Local interface | SD06, JSON, result.0.localIntfStatus | Not equal to up gives CRIT |
| Configuration consistent | SD07, JSON, result.0.configSanity | Not equal to consistent gives WARN |
| MLAG ports inactive | SV03, JSON, result.0.mlagPorts.Inactive | Greater than 0 gives WARN |
BGP
For each peer that matters, one sensor with "cmds":["show ip bgp neighbors 10.0.0.1"], reading SD03, JSON, result.0.vrfs.default.peerList.0.state, with SD03 not equal to Established gives CRIT. For a peer in another VRF, add vrf BLUE to the command and read result.0.vrfs.BLUE.peerList.0.state.
Why not show ip bgp summary, which lists every peer at once? Because it lists them by address, and an address such as 10.0.0.1 contains dots, which an Oversight JSON path uses to separate names. If one sensor for all peers is what you want, use it with a REGEX row instead, into SD03:
/^(?:(?=[\s\S]*?"peerState":\s*"(?!Established")(\w+)"))?/
It gives the state of the first peer that is not established, or nothing at all when every peer is, so SD03 matching /\w/ gives CRIT. It also fires for a peer that has been shut down on purpose, so use the per-peer form where some peers are meant to be down.
Ports
| What | Command | Slot and path | Rule |
|---|---|---|---|
| Any port err-disabled | show interfaces status errdisabled | SV03, JSONCOUNT, result.0.interfaceStatuses | Greater than 0 gives CRIT |
| One port connected | show interfaces Ethernet1 status | SD03, JSON, result.0.interfaceStatuses.Ethernet1.linkStatus | Not equal to connected gives CRIT |
Slashes in port names are fine in a path: result.0.interfaceStatuses.Ethernet3/1.linkStatus works.
Restarts
show version: SV03, JSON, result.0.uptime, unit s, with less than 600 gives WARN. The same answer carries result.0.memFree and result.0.memTotal in kilobytes, if you want a figure for memory.
Keep optional commands apart
If any one command in a request fails, eAPI answers with an error and no results at all, so every reading in that sensor comes back empty and the sensor cannot tell. Put show mlag and BGP commands in sensors of their own rather than adding them to the environment request, so that a command a switch cannot run does not blind the rest.
Setting up SNMP
ip access-list standard SNMP-RO 10 permit host <probe-address> ! snmp-server community <community> ro SNMP-RO snmp-server vrf MGMT snmp-server ipv4 access-list SNMP-RO vrf MGMT
SNMP starts as soon as a community exists. It answers only in the default VRF unless told otherwise, hence snmp-server vrf; set the access list for that VRF too, as shown. In Oversight, save an SNMP credential with version 2c and the community; see Credentials.
SNMP sensors worth having
Arista's readings are indexed by part, and the numbers are best found once with a walk:
snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.47.1.1.1.1.2 snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.31.1.1.1.1
The first names every physical part with its number, the second every port. On the switch itself, show snmp mib walk entPhysicalDesc does the same. On fixed-configuration switches the numbers follow a pattern, but walk rather than work them out:
| Part | Usual number |
|---|---|
| PowerSupply1, PowerSupply2 | 100711000, 100721000 |
| Fan Tray 1 Fan 1 | 100601110 (tray N is 10060N110) |
| Cpu temp sensor | 100006001 |
| Port EthernetN | N |
| Port-ChannelN | 1000000 + N |
| Management1 | 999001 |
| What | OID | Settings | Rules |
|---|---|---|---|
| Power supply | 1.3.6.1.2.1.131.1.1.1.3.<part>, such as ...3.100711000 | GAUGE; 3 is enabled | Not equal to 3 gives CRIT. A supply that has lost power reads 2. |
| Fan | 1.3.6.1.4.1.30065.3.12.1.2.1.2.<part> | GAUGE; 3 is enabled | Not equal to 3 gives CRIT |
| Temperature | 1.3.6.1.2.1.99.1.1.1.4.<part> | GAUGE, multiplier 0.1, unit C | At the switch's own limits: read them once from 1.3.6.1.4.1.30065.3.12.1.1.1.3.<part> (warning) and .4.<part> (critical), which are in tenths too |
| Sensor broken | 1.3.6.1.2.1.99.1.1.1.5.<part> | GAUGE | Equal to 3 gives CRIT |
| Port up | 1.3.6.1.2.1.2.2.1.8.<port> | GAUGE | Not equal to 1 gives CRIT |
| Port err-disabled | 1.3.6.1.4.1.30065.3.15.1.1.1.10.<port> | TEXT; empty when the port is fine, otherwise the reason | SD01 matching /\w/ gives CRIT |
| BGP peer | 1.3.6.1.4.1.30065.4.1.1.2.1.13.1.1.4.<IPv4>, such as ...13.1.1.4.10.252.0.1 | GAUGE; 6 is established | Not equal to 6 gives CRIT |
| CPU | 1.3.6.1.2.1.25.3.3.1.2.1 | GAUGE, unit % | Greater than 80 gives WARN, greater than 95 gives CRIT |
| Rebooted | 1.3.6.1.2.1.1.3.0 | GAUGE, multiplier 0.01, unit s | Less than 600 gives WARN |
Other sensors scale differently from temperatures: voltages and currents are hundredths (multiplier 0.01) and fan speeds whole RPM.
Do not alarm on memory from SNMP. The host resources MIB counts cache and buffers as used, so a healthy switch reads around 97% full. memFree from eAPI is the figure to use if you want one.
The Arista BGP MIB sits in Arista's experimental branch, which Arista says may change between releases. It has been stable for years, but check a BGP sensor still reads after an EOS upgrade.
A starting set
- Environment (eAPI, one sensor): both power supplies, cooling and temperature.
- MLAG (eAPI), on switches that run it.
- BGP (eAPI), a sensor per important peer, or one for all with the REGEX.
- Err-disabled ports (eAPI).
- Restarted (eAPI uptime).
- Uplinks and the MLAG peer link up (SNMP), one per port that must stay up.
Where eAPI is not allowed, use the SNMP power supply, fan, temperature and BGP sensors instead; MLAG is then only visible as the peer link's port. At the default 60 seconds from one probe group, an eAPI sensor costs about £3.46 a month and an SNMP sensor about £1.73.
Not on virtual EOS. vEOS and cEOS have no hardware to report, answer with unknown alarm levels and no power supplies, and would alarm on the environment sensor. Leave it off virtual instances.
Where CloudVision is in use, it already collects this state, but Oversight still needs eAPI or SNMP on the switches themselves.
MIBs
No MIB is needed to poll: every OID above works as a number. To pick objects by name, import these through Files, from Arista's MIB page (arista.com/en/support/product-documentation/arista-snmp-mibs, or the whole set as arista-mibs.zip) and the standard collections:
- Standard, in this order: UUID-TC-MIB, IANA-ENTITY-MIB, ENTITY-MIB, ENTITY-SENSOR-MIB, ENTITY-STATE-MIB, and BGP4-MIB if you want IPv4 BGP by the standard MIB.
- Arista: ARISTA-SMI-MIB first, then ARISTA-ENTITY-SENSOR-MIB, ARISTA-IF-MIB, ARISTA-BGP4V2-TC-MIB and ARISTA-BGP4V2-MIB.
Arista publishes its MIBs as .txt files; they upload through Files as they are.