Articles Windows agent: a verdict from the server itself setup, WINDOWS, updated 2026-09-23

The Oversight Windows agent is a single small program that runs as a service on a Windows server, watches it continuously and decides for itself whether anything is wrong. It serves that verdict over HTTP: nothing wrong, or what is wrong and how badly. Oversight reads it with a Windows agent sensor.

It judges on the server because it knows the server far better than a threshold in a template can: how long a call to WMI or the service manager usually takes on this machine, which network ports have ever had a link, whether a service is quietly crashing and being restarted. It is read only by design, runs as its own virtual service account with no administrator rights, has no way to run commands, and nothing it serves is a secret, a process list or a user name.

It runs on Windows 10 and Server 2016 onwards, desktop or server, and finds for itself whether the machine is a domain controller or a Remote Desktop Session Host.

Installing it on the server

No script is run. In a Command Prompt opened with Run as administrator:

cd %TEMP%
curl.exe -fLo oversight-agent.exe https://oversight.im/windowssensor/oversight-agent-windows-amd64.exe
certutil -hashfile oversight-agent.exe SHA256
oversight-agent.exe --install --listen 10.1.1.100:29731 --token A-LONG-RANDOM-TOKEN --allow 10.1.1.25

The hash printed by certutil must match the one published at oversight.im/windowssensor; if it does not, do not run the file. curl.exe is part of Windows 10 1803 and Server 2019 onwards; on Server 2016, download the file in a browser instead. The same page lists every command below.

OptionWhat it does
--listenThe one address and port the agent answers on: this server's own address, as the probe will reach it.
--tokenA bearer token the agent accepts. Use a long random one; the same token goes in a Windows agent credential in Oversight.
--allowThe addresses, or ranges such as 10.1.1.0/24, allowed to connect: your probe's address. Comma separate several.

The install copies the program to C:\Program Files\Oversight\agent.exe, writes C:\ProgramData\Oversight\agent.json, creates the OversightAgent service running as NT SERVICE\OversightAgent, adds that account to the built-in Event Log Readers and Performance Monitor Users groups, opens the port in Windows Firewall to the allowed addresses only, and starts it. The config file can be read only by the service, SYSTEM and Administrators, because it holds the token.

Set --token, --allow or, better, both. Without either the agent listens on 127.0.0.1 only, and it refuses to listen beyond the server with neither.

Check it from the server before going near Oversight:

curl.exe -H "Authorization: Bearer A-LONG-RANDOM-TOKEN" http://10.1.1.100:29731/v1/status

Upgrading is the same download, check and --install: it stops the service, replaces the program and starts it again, keeping the config. Removing it is "C:\Program Files\Oversight\agent.exe" --uninstall, which leaves C:\ProgramData\Oversight to be deleted by hand.

Adding the sensor

On the device that is this server, add a sensor of type Windows agent. If the agent has a token, add a credential of type Windows agent on the Credentials panel of that device, or of a group above it to share one token across many servers, with the token in the Bearer token field. The sensor sends it for you.

SettingWhat to put
EndpointStatus, the verdict, for nearly every sensor. Protection is Defender alone; see below.
Agent addresshttp and 29731, unless it was installed with another port, or https and the proxy's port where one is in front of it.
Named eventsOptional, Status only: events of your own to watch for. See below.
Interval60 seconds. The agent samples every 10 seconds whether anyone asks or not, so polling faster gains nothing.

A new Windows agent sensor is ready to use as soon as it is saved. It starts on the Status endpoint with these rows and rules already filled in. Changing the endpoint afterwards does not change the rows, because each endpoint carries different fields.

SlotNameExpressionHolds
SD01Alarmalarmgreen, amber or red
SD02StatusstatusEvery current problem, worst first, for the alert text
SD03Named eventswatched_eventsNamed events seen lately, empty otherwise
SD04OSosFor example Windows Server 2025 Datacenter
SD05StoragestorageVolumes full or filling, disk and file system errors, VSS
SD06MemorymemoryCommit, low memory, paging, nonpaged pool
SD07ServicesservicesStopped, crashing or failing to start
SD08PlatformplatformRPC, DCOM, WMI, the Server service, LSASS, Kerberos, the domain trust
SD09NetworknetworkLinks lost
SD10Hosthost_Crashes, the clock, reboots
SV01AlarmsalarmsHow many problems, charted
SV02Uptimeuptime_sSeconds since boot
SV03BuildbuildWindows build, for example 26100
SV04Patch levelubrThe monthly update, the part after the dot in 26100.6584
SV05User sessionssessionsRemote Desktop users, on a Session Host only. Nothing raises on it; add a rule if a host with nobody on it matters to you
SV06Agent uptimeagent_uptime_sSeconds since the agent started

The rules: SV20 (HTTP status) not 200 is critical, SD01 red is critical, SD03 not empty is critical, and SD01 amber is a warning. SV20 is 401 for a wrong or missing token and 403 for a probe outside allow.

The agent also reports directory (domain controllers) and rdsh (Session Hosts), which have no slot of their own: there are ten text slots. Their problems are always in SD02, Status, and they can be extracted in place of an area that does not apply to the server.

Finding servers left behind

Because Build and Patch level are numbers, a rule or report can find every server at or below a given build, an upgrade that never happened, or below a patch level, one that has stopped taking updates. Read Build together with OS: servers and desktops share build numbers (Server 2019 is 17763, 2022 is 20348, Server 2025 and Windows 11 24H2 are both 26100).

What it checks

A problem must hold for the time shown before it raises, and clear for as long before it drops. For five minutes after a reboot nothing raises except the reboot and a crash, so services and links have time to arrive.

CheckRaises whenHolds forLimits, with their defaultsWhat the alert says
vol_fullA fixed volume 90% full (amber) or 95% (red)5 minamber 90, red 95 (%)D: 92% full, 40.1 GB free
vol_fillingA volume on course to be full within 12 hours at the last hour's rate15 minhours 12D: full in 9 h at the last hour's rate
commitCommit at 90% (amber) or 95% (red) of its limit: the real out-of-memory on Windows, which can happen with physical memory still free5 minamber 90, red 95 (%)commit 92% of its limit (29.4 GB of 32.0 GB)
commit_risingCommit on course to reach its limit within 3 hours: a leak, raised hours before it ends in a crash15 minhours 3commit reaches its limit in 2.1 h at the last hour's rate, 74% now
mem_lowAvailable memory under 5%5 minamber 5 (% available)available memory 3.2% (1.0 GB)
pagingPaging in at 1 MB/s or more with under 10% memory available: thrashing5 minamber 1 (MB/s), avail 10 (% available)paging in at 4.2 MB/s with 6.1% memory available, for 7 min
poolNonpaged pool over 10% (amber) or 25% (red) of memory: a driver leaking15 minamber 10, red 25 (% of memory)nonpaged pool 3.4 GB, 11% of memory: a driver leaking
svc_autoA service set to start automatically is stopped (delayed and trigger-started services excepted, since they stop legitimately)5 minnonePrint Spooler (Spooler) stopped, set to start automatically
svc_namedA service listed in the config's services is not running, or not installed1 minnoneWorld Wide Web Publishing Service (W3SVC) stopped, MSSQLSERVER not installed
rpc, dcom, wmi, smb, lsass, kerberosA cheap call to each takes ten times its usual time on this server, and at least 250 ms (500 ms for WMI and Kerberos), for 3 minutes (amber), or has not answered within 30 seconds (red). This is the stage before the classic collapse: RPC slows, everything that depends on it times out, services crash-loop and the server falls over. The agent keeps answering throughout. smb is also amber when the Server service runs short of work items or rejects blocking requestsas statedtimes 10 (times the usual), floor_ms 250 (500 for wmi and kerberos)RPC (service manager) slow, 2400 ms against a usual 12.0 ms, for 4 min, WMI query not answering for 45 s, Kerberos ticket for this machine failed: ..., Server service short of work items 12 times in the last minute
ad_services, ldap, dc_dnsDomain controllers only: NTDS, Netlogon, the KDC, W32Time, DNS and DFSR (or FRS) running; LDAP answering on the DC; the DC locator record answered by its DNS1 minnoneNetlogon not running, LDAP not answering on this DC: ..., DC locator record not answered by local DNS: ...
link_downA physical port that had a link since the agent started has lost it (Hyper-V virtual adapters, wifi and VPNs are left alone)30 snoneEthernet 2 lost link at 14:02
timeThe clock 1 second or more out against the server's own time source (a domain controller, or its NTP server)10 minamber 1 (seconds)clock 2.3 s slow against dc01.example.local for 11 min, clock set not to synchronise (W32Time type NoSync)
reboot_pendingA reboot pending from Windows Update, servicing or file renames24 hnonereboot pending (Windows Update) for at least 26 h 10 min
rebootedBooted within the last 30 minuteswhile it holdsnonerebooted at 14:02
defenderOn the Protection endpoint only: see Protection: Defenderat oncesig_days 7 (days)Defender real-time protection is off, Defender signatures 9 days old

From the event log, each raising at once for the time shown:

CheckEventsAlarmWhat the alert says
crashedKernel-Power 41, EventLog 6008, BugCheck 1001Red for 30 minutes, with the stop coderestarted without a clean shutdown, logged 03:12 (Kernel-Power 41), blue screen, stop code 0x0000009f, logged 03:12 (BugCheck 1001)
commit_exhaustedResource-Exhaustion-Detector 2004Red for an hourWindows ran out of virtual memory at 14:02 (Resource-Exhaustion-Detector 2004)
disk_errorsdisk 7, 51, 153; a storage driver's 129 (device reset)Amber for an hourdisk error on \Device\Harddisk1\DR1 at 14:02 (disk 153), device reset on \Device\RaidPort0 at 14:02 (stornvme 129)
fs_errorsNtfs 55; Ntfs 98 at warning or error (98 at information is a healthy volume)Red for a dayfile system corruption on D:, chkdsk needed (Ntfs 55), volume D: needs chkdsk (Ntfs 98)
vssVSS 8193, 12289Amber for a day: backups quietly stop when VSS doesVolume Shadow Copy error at 02:00 (VSS 8193): backups may be failing
svc_crashingService Control Manager 7031, 7034Amber at 3 in 15 minutesPrint Spooler terminated unexpectedly, 3 times in 15 min
svc_start_failedService Control Manager 7000, 7009Amber for an hourPrint Spooler failed to start at 14:02: ..., Print Spooler timed out starting at 14:02
dcom_timeoutDCOM 10010Amber for an houra COM server did not register with DCOM in time at 14:02 (DCOM 10010)
trustNetlogon 3210: the secure channel with the domain has failed, so every domain logon to the server failsRed for an hourcannot authenticate with \\DC01 for domain EXAMPLE: domain logons to this machine fail (Netlogon 3210)
name_conflictNetBT 4321Amber for an hourNetBIOS name FS01 already in use on the network (NetBT 4321)
sysvolDFSR 2213, domain controllers only: SYSVOL replication paused and waiting to be resumedRed for a weekSYSVOL replication paused after a dirty shutdown and waiting to be resumed (DFSR 2213)
profileUser Profile Service 1511, 1515: a user given a temporary profileAmber for an houra user was given a temporary profile at 14:02 (User Profile Service 1511)

The event checks take no limits: each can be turned off, and disk_errors, svc_crashing and svc_start_failed can be told to ignore one disk or service under ignore, by the name the alert gives it.

This list is a file, events.json, signed by GEN and fetched from oversight.im when the service starts and daily after, so a new Windows release needs a new list, not a new agent. A file that does not verify is refused and the agent keeps what it had, and where outbound access is blocked it simply uses the list it was installed with. events_version and events_source in the status say which it is using.

Anything that cannot be checked on a server, such as Kerberos on a machine in a workgroup, is listed under unavailable with the reason, so a check that cannot see is never mistaken for one that found nothing.

On a domain controller

A domain controller has no local groups, and the agent's own account cannot join the domain's, so the install says so and carries on without them. Everything still works except the Server service's work-item counters, which are listed under unavailable; the share-list timing, the main part of that check, is unaffected. The session count works everywhere, since it is taken from which sessions have a signed-in user's shell running, which needs no rights.

Protection: Defender

Defender is reported on its own endpoint, /v1/protection, and never in Status, because many servers are protected by another product. Where Defender is the protection, add a second Windows agent sensor on the Protection endpoint. It is amber at worst:

  • antivirus or real-time protection switched off;
  • signatures 7 days old or more;
  • a threat found or acted on (Defender 1116, 1117), or real-time protection turned off (5001), for an hour after.

Where Defender is in passive or EDR block mode, another product is the protection and nothing raises; unavailable says so.

{"alarm":"amber","alarms":1,"status":"Defender real-time protection is off","defender":true,
 "defender_mode":"Normal","antivirus_enabled":true,"realtime_protection":false,
 "signature_age_days":0,"unavailable":[]}
SlotNameMethodExpression
SD01AlarmJSONalarm
SD02StatusJSONstatus
SD03ModeJSONdefender_mode
SV01AlarmsJSONalarms
SV02Signature ageJSONsignature_age_days

with the rules SV20 not 200 critical and SD01 amber a warning. The signature age threshold is "defender": {"sig_days": 7} in the config.

Named events

For the few things only an event shows, list them in the sensor's Named events field, up to 20, comma separated, each as channel/ID or channel/provider/ID:

System/disk/7, Application/1000, Microsoft-Windows-Windows Defender/Operational/5001

Give the provider wherever the ID alone is ambiguous: 1001 is a blue screen in System and an application error report in Application. When one is seen, it appears in SD03 and in the status text for an hour or three polls, whichever is sooner, and the rule on SD03 makes the sensor critical. The first poll after the list changes reads the last hour back, so an event just before is not missed. Only IDs travel, never the text of an event.

The config file

To quieten an alarm on one server, find its check first: the alert's detail is the agent's own words, and the tables under What it checks give the check each message comes from, so commit reaches its limit in 2.1 h is commit_rising. Turn it off, or change its limits, under checks here. To exclude one volume, adapter or service rather than a whole check, name it under ignore. It can only be done on the server: nothing in Oversight changes the agent's settings.

C:\ProgramData\Oversight\agent.json, edited as Administrator. After editing by hand, net stop OversightAgent && net start OversightAgent.

{
  "listen": "10.1.1.100:29731",
  "tokens": ["a-long-random-token"],
  "allow": ["10.1.1.25"],
  "checks": {"vol_full": {"amber": 85, "red": 95}, "reboot_pending": false},
  "ignore": ["E:", "Ethernet 2", "SomeVendorUpdater"],
  "services": ["W3SVC", "MSSQLSERVER"]
}
KeyWhat it does
checksPer check, false to turn it off, or its limits to change them. The limits each check takes, with their defaults, are in the tables under What it checks; only those named change. A misspelt check or limit stops the agent starting, and it says which.
ignoreNames no check ever raises on: a volume, a network adapter, a service. For the known exception on one server, rather than turning a whole check off.
servicesServices that must be running whatever their start type: the ones that matter on this server.

When it is not answering

  • SV20 401: the token in the credential does not match one in tokens.
  • SV20 403: the probe's address is not in allow.
  • No answer at all: the service is stopped, or the firewall rule does not cover the probe. Running --install or --allow again rewrites the rule from the config.
  • The agent logs to the Application event log, source OversightAgent. "C:\Program Files\Oversight\agent.exe" --dump status samples for a minute and prints what it would report.