Proving where network jitter lives: a three-script PowerShell toolkit

8 September 2026 · Luke Gillmore-White

“The network is slow” is one of the hardest tickets to close. It’s rarely constant, it’s often somewhere you don’t control, and a single ping test proves almost nothing. After one job where the whole question was whether a problem sat on the local network or on the long path to a distant server, I turned the commands I’d been running by hand into three PowerShell scripts. Each covers one phase of an investigation:

NetTest.ps1     Phase 1  one-shot inventory of the machine and the path
NetLoop.ps1     Phase 2  continuous sampling for as long as you need
NetCompare.ps1  Phase 3  lines several runs up against each other

They need no admin rights and no modules, and they prompt for everything at launch (target, reference host, label, duration), so they work on any job without editing. Every prompt also has a matching parameter, so they can run unattended too.

Download: NetTest.ps1 · NetLoop.ps1 · NetCompare.ps1

powershell -NoProfile -ExecutionPolicy Bypass -File .\NetTest.ps1

Phase 1: NetTest, the inventory

Every investigation starts with a snapshot, so each later run has its context attached. NetTest writes an inventory.txt covering:

  • OS, last boot, whether the session is elevated, and clock sync status. Comparing timestamps across machines is pointless if one clock has drifted, and the script tells you how to fix it if w32time isn’t running.
  • every active adapter, its IP, gateway and configured DNS servers, plus the public IP
  • Wi-Fi association (SSID, BSSID, signal, band, channel), plus a scan of every neighbouring network saved to its own file
  • how long each configured DNS server takes to answer, three times each
  • path MTU to the target and a reference host
  • traceroutes to both, with every hop labelled by country and provider

Labelling traceroute hops without an API

A bare traceroute is a list of IPs that mean nothing until you look each one up. NetTest uses Team Cymru’s IP-to-ASN service, which answers ordinary DNS TXT queries. That means no API key, no account and no HTTP calls, just DNS:

$o   = $Ip -split '\.'
$rev = "{0}.{1}.{2}.{3}" -f $o[3], $o[2], $o[1], $o[0]
$txt = (Resolve-DnsName "$rev.origin.asn.cymru.com" -Type TXT -DnsOnly).Strings -join ''
# "13335 | 1.1.1.0/24 | AU | apnic | 2011-08-11"  ->  ASN, prefix, country

A second lookup against AS<number>.asn.cymru.com turns the ASN into the provider’s name.

One gap catches everyone out: internet exchange peering LANs aren’t announced in BGP, so their addresses return nothing. NetTest carries a small table of known exchange prefixes (LINX, LONAP, DE-CIX, AMS-IX, France-IX, Equinix and others). Hops through an exchange come out as “LINX LON1 internet exchange” instead of a blank. At the end it prints a condensed path that only lists a hop where the country or provider changes, so you can see at a glance who carries your traffic and where it changes hands.

Path MTU, and what the number tells you

The MTU check sends pings with the don’t-fragment bit set, using a binary search to find the largest one that gets through. If the result is below 1500, it names the likely cause:

reduced by 24 bytes  ->  GRE tunnel
reduced by 8 bytes   ->  PPPoE
reduced by 20 bytes  ->  IP-in-IP

On one run the reference path came back at the normal 1500 bytes, while the path to the target came back at 1476. That 24-byte gap told me there was a GRE tunnel at the far end before I’d spoken to anyone about it.

Phase 2: NetLoop, the sampling run

NetLoop runs for a fixed period, an hour by default, and records everything to a CSV as it goes:

every second     ping the gateway, the target and the reference hosts
every 5 seconds  TCP connect to the ports that matter, and a Wi-Fi sample
every 30 seconds DNS timing against every configured resolver,
                 and a snapshot of live external connections

Each Wi-Fi sample records the BSSID and signal, so an access point roam shows up in the data. Each connection snapshot records the owning process, and after the run every destination is labelled with its provider. Two or more consecutive lost pings are logged as an outage, with its exact duration.

The most useful feature turned out to be the simplest. Press M while it’s running and a timestamped marker goes into the CSV. When a user says “it just froze”, you press M, and afterwards you can see exactly what every target was doing at that moment.

When the run finishes, it pulls the WLAN and RDP client event logs for the same window and prints a summary for every target: min, average, 95th percentile, max, standard deviation, loss, and the number I ended up relying on most.

Why mean delta beats the average

Average latency to a distant host is mostly distance, and there’s nothing you can do about distance. Standard deviation is better, but it mixes slow drift with genuine jitter. What breaks interactive sessions is the sample-to-sample swing, so I measure that directly:

for ($i = 1; $i -lt $vals.Count; $i++) {
    $dsum += [math]::Abs($vals[$i] - $vals[$i-1]); $dn++
}
$meanDelta = $dsum / $dn

Here’s an evening run from my own home Wi-Fi:

Target            Avg (ms)   Mean delta (ms)
Gateway              6.7          3.3
1.1.1.1              8.3          3.3
8.8.8.8              6.9          3.3
Distant host       203.1         37.0

The local targets all swing by about 3 ms between samples. The distant host swings by 37 ms, more than ten times as much. Its minimum was perfectly steady, so the floor was fine and the variance was coming from somewhere else.

NetTest’s labelled traceroute then showed where. Every hop up to the last one in the UK was clean, and the very next hop, on the far side of the long-haul link, was already jumping between 176 and 306 ms. All the jitter was entering at a single link.

Change one variable at a time

The next morning I re-ran it on a wired connection. The distant host’s mean delta dropped from 37 ms to 0.5 ms, with a standard deviation of 1.6 ms.

My first instinct was “the long-haul link is congested in the evenings.” But I’d changed two things at once, the connection type and the time of day, so that run couldn’t tell me which one mattered. Hold everything constant and change one variable per run (wired vs Wi-Fi at the same time, then the same connection at different times) before blaming anything.

Phase 3: NetCompare, like for like

Comparing runs by eye across several CSVs is where mistakes creep in, so NetCompare does it. Point it at a folder and it finds every samples.csv underneath:

.\NetCompare.ps1 -Root .\results -Csv

If the runs overlap in time, for example a wired laptop and a Wi-Fi laptop sampling side by side, it restricts every figure to the shared window, so you’re comparing identical conditions. Then it prints each target’s statistics for every run side by side, colouring a mean delta yellow above 20 ms and red above 50 ms. It also lists each run’s outages and its TCP and DNS results.

Symptom markers get the cross-run treatment too. For every press of M on any machine, it shows what every run saw in the 10 seconds either side. That’s usually the moment it becomes obvious whether a freeze was local to one device or happening everywhere at once.

The bug in my own tool

When I eventually looked at the raw timestamps, the samples weren’t anywhere near one second apart. The median gap was 0.19 seconds.

The cause was a combination of two things. Every 30 seconds the loop timed a DNS lookup against each configured resolver, one after another, inside the loop. My own machine still had a couple of resolvers configured that no longer answered: leftovers from an old ISP filtering service and a retired lab resolver. Each one had to time out before the loop could carry on, so sampling stalled for about 22 seconds. Then the scheduler tried to catch up by firing every missed tick back to back. The result was bursts of near-identical samples, which made the statistics look calmer than reality.

The rewrite follows a few rules:

  • No name resolution in the loop. Targets are resolved to IP addresses once, before sampling starts. Provider lookups for live connections happen after the run.
  • Nothing in the loop waits on the network. Pings, TCP connects and DNS lookups are all started asynchronously and collected when they finish. DNS lookups run in their own runspaces, so a dead resolver only delays its own result.
  • Missed ticks are skipped, never replayed. Ticks are scheduled against a monotonic stopwatch. If the loop ever does fall behind, the missed ticks are logged as late rows rather than fired in a burst.
$behind = [math]::Floor(($now - $nextTickMs) / $IntervalMs)
if ($behind -ge 1) { Write-Row 'late' 'loop' 'Skipped' $behind }
# ... start this tick's probes ...
$nextTickMs += $IntervalMs * ($behind + 1)

To prove it, I tested against a resolver that took six seconds to time out. The ping samples stayed between 0.94 and 1.06 seconds apart, with no ticks skipped. NetLoop now reports its own median sample spacing at the end of every run, and NetCompare flags any run, including old ones, whose spacing is off, so bunched data can’t slip into a comparison unnoticed.

The DNS timing is also what pointed me at those dead resolvers in the first place. A tool built to find other people’s DNS problems found mine first.

Result

Three scripts that turn “it’s slow sometimes” into evidence:

  • where the latency floor is,
  • where the variance enters the path, and whose network that is,
  • whether it follows the connection type or the time of day,
  • what was happening at the exact moment a user felt it,
  • whether DNS is quietly adding seconds to everything.

The bug was the most useful lesson of the lot: check the timestamps of your own data before trusting the statistics built on it.