How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
The Oversight Windows agent is a single small program that runs as a service on a Windows server, watches it continuously and decides for itself whether anything is wrong. It serves that verdict over HTTP: nothing wrong, or what is wrong and how badly. Oversight reads it with a Windows agent sensor.
It judges on the server because it knows the server far better than a threshold in a template can: how long a call to WMI or the service manager usually takes on this machine, which network ports have ever had a link, whether a service is quietly crashing and being restarted. It is read only by design, runs as its own virtual service account with no administrator rights, has no way to run commands, and nothing it serves is a secret, a process list or a user name.
It runs on Windows 10 and Server 2016 onwards, desktop or server, and finds for itself whether the machine is a domain controller or a Remote Desktop Session Host.
Installing it on the server
No script is run. In a Command Prompt opened with Run as administrator:
cd %TEMP%
curl.exe -fLo oversight-agent.exe https://oversight.im/windowssensor/oversight-agent-windows-amd64.exe
certutil -hashfile oversight-agent.exe SHA256
oversight-agent.exe --install --listen 10.1.1.100:29731 --token A-LONG-RANDOM-TOKEN --allow 10.1.1.25
The hash printed by certutil must match the one published at oversight.im/windowssensor; if it does not, do not run the file. curl.exe is part of Windows 10 1803 and Server 2019 onwards; on Server 2016, download the file in a browser instead. The same page lists every command below.
| Option | What it does |
|---|---|
--listen | The one address and port the agent answers on: this server's own address, as the probe will reach it. |
--token | A bearer token the agent accepts. Use a long random one; the same token goes in a Windows agent credential in Oversight. |
--allow | The addresses, or ranges such as 10.1.1.0/24, allowed to connect: your probe's address. Comma separate several. |
The install copies the program to C:\Program Files\Oversight\agent.exe, writes C:\ProgramData\Oversight\agent.json, creates the OversightAgent service running as NT SERVICE\OversightAgent, adds that account to the built-in Event Log Readers and Performance Monitor Users groups, opens the port in Windows Firewall to the allowed addresses only, and starts it. The config file can be read only by the service, SYSTEM and Administrators, because it holds the token.
Set --token, --allow or, better, both. Without either the agent listens on 127.0.0.1 only, and it refuses to listen beyond the server with neither.
Check it from the server before going near Oversight:
curl.exe -H "Authorization: Bearer A-LONG-RANDOM-TOKEN" http://10.1.1.100:29731/v1/status
Upgrading is the same download, check and --install: it stops the service, replaces the program and starts it again, keeping the config. Removing it is "C:\Program Files\Oversight\agent.exe" --uninstall, which leaves C:\ProgramData\Oversight to be deleted by hand.
Adding the sensor
On the device that is this server, add a sensor of type Windows agent. If the agent has a token, add a credential of type Windows agent on the Credentials panel of that device, or of a group above it to share one token across many servers, with the token in the Bearer token field. The sensor sends it for you.
| Setting | What to put |
|---|---|
| Endpoint | Status, the verdict, for nearly every sensor. Protection is Defender alone; see below. |
| Agent address | http and 29731, unless it was installed with another port, or https and the proxy's port where one is in front of it. |
| Named events | Optional, Status only: events of your own to watch for. See below. |
| Interval | 60 seconds. The agent samples every 10 seconds whether anyone asks or not, so polling faster gains nothing. |
A new Windows agent sensor is ready to use as soon as it is saved. It starts on the Status endpoint with these rows and rules already filled in. Changing the endpoint afterwards does not change the rows, because each endpoint carries different fields.
| Slot | Name | Expression | Holds |
|---|---|---|---|
| SD01 | Alarm | alarm | green, amber or red |
| SD02 | Status | status | Every current problem, worst first, for the alert text |
| SD03 | Named events | watched_events | Named events seen lately, empty otherwise |
| SD04 | OS | os | For example Windows Server 2025 Datacenter |
| SD05 | Storage | storage | Volumes full or filling, disk and file system errors, VSS |
| SD06 | Memory | memory | Commit, low memory, paging, nonpaged pool |
| SD07 | Services | services | Stopped, crashing or failing to start |
| SD08 | Platform | platform | RPC, DCOM, WMI, the Server service, LSASS, Kerberos, the domain trust |
| SD09 | Network | network | Links lost |
| SD10 | Host | host_ | Crashes, the clock, reboots |
| SV01 | Alarms | alarms | How many problems, charted |
| SV02 | Uptime | uptime_s | Seconds since boot |
| SV03 | Build | build | Windows build, for example 26100 |
| SV04 | Patch level | ubr | The monthly update, the part after the dot in 26100.6584 |
| SV05 | User sessions | sessions | Remote Desktop users, on a Session Host only. Nothing raises on it; add a rule if a host with nobody on it matters to you |
| SV06 | Agent uptime | agent_uptime_s | Seconds since the agent started |
The rules: SV20 (HTTP status) not 200 is critical, SD01 red is critical, SD03 not empty is critical, and SD01 amber is a warning. SV20 is 401 for a wrong or missing token and 403 for a probe outside allow.
The agent also reports directory (domain controllers) and rdsh (Session Hosts), which have no slot of their own: there are ten text slots. Their problems are always in SD02, Status, and they can be extracted in place of an area that does not apply to the server.
Finding servers left behind
Because Build and Patch level are numbers, a rule or report can find every server at or below a given build, an upgrade that never happened, or below a patch level, one that has stopped taking updates. Read Build together with OS: servers and desktops share build numbers (Server 2019 is 17763, 2022 is 20348, Server 2025 and Windows 11 24H2 are both 26100).
What it checks
A problem must hold for the time shown before it raises, and clear for as long before it drops. For five minutes after a reboot nothing raises except the reboot and a crash, so services and links have time to arrive.
| Check | Raises when | Holds for | Limits, with their defaults | What the alert says |
|---|---|---|---|---|
vol_full | A fixed volume 90% full (amber) or 95% (red) | 5 min | amber 90, red 95 (%) | D: 92% full, 40.1 GB free |
vol_filling | A volume on course to be full within 12 hours at the last hour's rate | 15 min | hours 12 | D: full in 9 h at the last hour's rate |
commit | Commit at 90% (amber) or 95% (red) of its limit: the real out-of-memory on Windows, which can happen with physical memory still free | 5 min | amber 90, red 95 (%) | commit 92% of its limit (29.4 GB of 32.0 GB) |
commit_rising | Commit on course to reach its limit within 3 hours: a leak, raised hours before it ends in a crash | 15 min | hours 3 | commit reaches its limit in 2.1 h at the last hour's rate, 74% now |
mem_low | Available memory under 5% | 5 min | amber 5 (% available) | available memory 3.2% (1.0 GB) |
paging | Paging in at 1 MB/s or more with under 10% memory available: thrashing | 5 min | amber 1 (MB/s), avail 10 (% available) | paging in at 4.2 MB/s with 6.1% memory available, for 7 min |
pool | Nonpaged pool over 10% (amber) or 25% (red) of memory: a driver leaking | 15 min | amber 10, red 25 (% of memory) | nonpaged pool 3.4 GB, 11% of memory: a driver leaking |
svc_auto | A service set to start automatically is stopped (delayed and trigger-started services excepted, since they stop legitimately) | 5 min | none | Print Spooler (Spooler) stopped, set to start automatically |
svc_named | A service listed in the config's services is not running, or not installed | 1 min | none | World Wide Web Publishing Service (W3SVC) stopped, MSSQLSERVER not installed |
rpc, dcom, wmi, smb, lsass, kerberos | A cheap call to each takes ten times its usual time on this server, and at least 250 ms (500 ms for WMI and Kerberos), for 3 minutes (amber), or has not answered within 30 seconds (red). This is the stage before the classic collapse: RPC slows, everything that depends on it times out, services crash-loop and the server falls over. The agent keeps answering throughout. smb is also amber when the Server service runs short of work items or rejects blocking requests | as stated | times 10 (times the usual), floor_ms 250 (500 for wmi and kerberos) | RPC (service manager) slow, 2400 ms against a usual 12.0 ms, for 4 min, WMI query not answering for 45 s, Kerberos ticket for this machine failed: ..., Server service short of work items 12 times in the last minute |
ad_services, ldap, dc_dns | Domain controllers only: NTDS, Netlogon, the KDC, W32Time, DNS and DFSR (or FRS) running; LDAP answering on the DC; the DC locator record answered by its DNS | 1 min | none | Netlogon not running, LDAP not answering on this DC: ..., DC locator record not answered by local DNS: ... |
link_down | A physical port that had a link since the agent started has lost it (Hyper-V virtual adapters, wifi and VPNs are left alone) | 30 s | none | Ethernet 2 lost link at 14:02 |
time | The clock 1 second or more out against the server's own time source (a domain controller, or its NTP server) | 10 min | amber 1 (seconds) | clock 2.3 s slow against dc01.example.local for 11 min, clock set not to synchronise (W32Time type NoSync) |
reboot_pending | A reboot pending from Windows Update, servicing or file renames | 24 h | none | reboot pending (Windows Update) for at least 26 h 10 min |
rebooted | Booted within the last 30 minutes | while it holds | none | rebooted at 14:02 |
defender | On the Protection endpoint only: see Protection: Defender | at once | sig_days 7 (days) | Defender real-time protection is off, Defender signatures 9 days old |
From the event log, each raising at once for the time shown:
| Check | Events | Alarm | What the alert says |
|---|---|---|---|
crashed | Kernel-Power 41, EventLog 6008, BugCheck 1001 | Red for 30 minutes, with the stop code | restarted without a clean shutdown, logged 03:12 (Kernel-Power 41), blue screen, stop code 0x0000009f, logged 03:12 (BugCheck 1001) |
commit_exhausted | Resource-Exhaustion-Detector 2004 | Red for an hour | Windows ran out of virtual memory at 14:02 (Resource-Exhaustion-Detector 2004) |
disk_errors | disk 7, 51, 153; a storage driver's 129 (device reset) | Amber for an hour | disk error on \Device\Harddisk1\DR1 at 14:02 (disk 153), device reset on \Device\RaidPort0 at 14:02 (stornvme 129) |
fs_errors | Ntfs 55; Ntfs 98 at warning or error (98 at information is a healthy volume) | Red for a day | file system corruption on D:, chkdsk needed (Ntfs 55), volume D: needs chkdsk (Ntfs 98) |
vss | VSS 8193, 12289 | Amber for a day: backups quietly stop when VSS does | Volume Shadow Copy error at 02:00 (VSS 8193): backups may be failing |
svc_crashing | Service Control Manager 7031, 7034 | Amber at 3 in 15 minutes | Print Spooler terminated unexpectedly, 3 times in 15 min |
svc_start_failed | Service Control Manager 7000, 7009 | Amber for an hour | Print Spooler failed to start at 14:02: ..., Print Spooler timed out starting at 14:02 |
dcom_timeout | DCOM 10010 | Amber for an hour | a COM server did not register with DCOM in time at 14:02 (DCOM 10010) |
trust | Netlogon 3210: the secure channel with the domain has failed, so every domain logon to the server fails | Red for an hour | cannot authenticate with \\DC01 for domain EXAMPLE: domain logons to this machine fail (Netlogon 3210) |
name_conflict | NetBT 4321 | Amber for an hour | NetBIOS name FS01 already in use on the network (NetBT 4321) |
sysvol | DFSR 2213, domain controllers only: SYSVOL replication paused and waiting to be resumed | Red for a week | SYSVOL replication paused after a dirty shutdown and waiting to be resumed (DFSR 2213) |
profile | User Profile Service 1511, 1515: a user given a temporary profile | Amber for an hour | a user was given a temporary profile at 14:02 (User Profile Service 1511) |
The event checks take no limits: each can be turned off, and disk_errors, svc_crashing and svc_start_failed can be told to ignore one disk or service under ignore, by the name the alert gives it.
This list is a file, events.json, signed by GEN and fetched from oversight.im when the service starts and daily after, so a new Windows release needs a new list, not a new agent. A file that does not verify is refused and the agent keeps what it had, and where outbound access is blocked it simply uses the list it was installed with. events_version and events_source in the status say which it is using.
Anything that cannot be checked on a server, such as Kerberos on a machine in a workgroup, is listed under unavailable with the reason, so a check that cannot see is never mistaken for one that found nothing.
On a domain controller
A domain controller has no local groups, and the agent's own account cannot join the domain's, so the install says so and carries on without them. Everything still works except the Server service's work-item counters, which are listed under unavailable; the share-list timing, the main part of that check, is unaffected. The session count works everywhere, since it is taken from which sessions have a signed-in user's shell running, which needs no rights.
Protection: Defender
Defender is reported on its own endpoint, /v1/protection, and never in Status, because many servers are protected by another product. Where Defender is the protection, add a second Windows agent sensor on the Protection endpoint. It is amber at worst:
- antivirus or real-time protection switched off;
- signatures 7 days old or more;
- a threat found or acted on (Defender 1116, 1117), or real-time protection turned off (5001), for an hour after.
Where Defender is in passive or EDR block mode, another product is the protection and nothing raises; unavailable says so.
{"alarm":"amber","alarms":1,"status":"Defender real-time protection is off","defender":true,
"defender_mode":"Normal","antivirus_enabled":true,"realtime_protection":false,
"signature_age_days":0,"unavailable":[]}
| Slot | Name | Method | Expression |
|---|---|---|---|
| SD01 | Alarm | JSON | alarm |
| SD02 | Status | JSON | status |
| SD03 | Mode | JSON | defender_mode |
| SV01 | Alarms | JSON | alarms |
| SV02 | Signature age | JSON | signature_age_days |
with the rules SV20 not 200 critical and SD01 amber a warning. The signature age threshold is "defender": {"sig_days": 7} in the config.
Named events
For the few things only an event shows, list them in the sensor's Named events field, up to 20, comma separated, each as channel/ID or channel/provider/ID:
System/disk/7, Application/1000, Microsoft-Windows-Windows Defender/Operational/5001
Give the provider wherever the ID alone is ambiguous: 1001 is a blue screen in System and an application error report in Application. When one is seen, it appears in SD03 and in the status text for an hour or three polls, whichever is sooner, and the rule on SD03 makes the sensor critical. The first poll after the list changes reads the last hour back, so an event just before is not missed. Only IDs travel, never the text of an event.
The config file
To quieten an alarm on one server, find its check first: the alert's detail is the agent's own words, and the tables under What it checks give the check each message comes from, so commit reaches its limit in 2.1 h is commit_rising. Turn it off, or change its limits, under checks here. To exclude one volume, adapter or service rather than a whole check, name it under ignore. It can only be done on the server: nothing in Oversight changes the agent's settings.
C:\ProgramData\Oversight\agent.json, edited as Administrator. After editing by hand, net stop OversightAgent && net start OversightAgent.
{
"listen": "10.1.1.100:29731",
"tokens": ["a-long-random-token"],
"allow": ["10.1.1.25"],
"checks": {"vol_full": {"amber": 85, "red": 95}, "reboot_pending": false},
"ignore": ["E:", "Ethernet 2", "SomeVendorUpdater"],
"services": ["W3SVC", "MSSQLSERVER"]
}
| Key | What it does |
|---|---|
checks | Per check, false to turn it off, or its limits to change them. The limits each check takes, with their defaults, are in the tables under What it checks; only those named change. A misspelt check or limit stops the agent starting, and it says which. |
ignore | Names no check ever raises on: a volume, a network adapter, a service. For the known exception on one server, rather than turning a whole check off. |
services | Services that must be running whatever their start type: the ones that matter on this server. |
When it is not answering
- SV20 401: the token in the credential does not match one in
tokens. - SV20 403: the probe's address is not in
allow. - No answer at all: the service is stopped, or the firewall rule does not cover the probe. Running
--installor--allowagain rewrites the rule from the config. - The agent logs to the Application event log, source OversightAgent.
"C:\Program Files\Oversight\agent.exe" --dump statussamples for a minute and prints what it would report.