Technical

Threat Hunting for Data Movement: What EDR Telemetry Leaves Out

June 5, 2026 Hilt 8 min

Your most valuable data leaves on access you granted on purpose, so every tool lets it through. Threat hunting techniques that catch the pattern across moves, not just the rule break.

Threat Hunting for Data Movement: What EDR Telemetry Leaves Out cover image

The hunt comes back clean. Your EDR shows a process tree with nothing out of place. Network logs show connections you expect. File events look like work. You close the ticket. Three weeks later a customer dataset shows up somewhere it should not be, and when you trace it back, every step was an authorized read by an account that was supposed to have that access.

Nothing broke. That is the problem. EDR is built to catch the rule break: the malware, the exploit, the credential that should not exist. Your most valuable data does not leave that way. It leaves on access you granted on purpose, through accounts that are supposed to have it, across cloud workloads, SaaS, user endpoints, and AI agents. Each move is permitted, so every tool you own correctly lets it through.

So the move you are hunting for never trips anything. One read does not look like a breach. Neither does the next. The breach is the shape they make together, and you can only see that shape while the data is moving, not guessed before it happens, not reconstructed after from logs. This is the gap threat hunters keep falling into, and it belongs to any company that moves data it cares about.

EDR answers the wrong question

EDR is good at the thing it was built for. It hooks the operating system through documented interfaces, forwards an event stream, and correlates it. For malware, exploitation, and execution that should not be running, reach for it.

It answers "did something break the rules?" It cannot answer "did the right person, with the right access, move the wrong data?" A service account that is supposed to read a customer database reads the whole thing. A developer with real credentials pulls a quarter of the data lake down to a laptop. An agent with a valid token fans out across systems it has never touched. The process tree stays clean through all of it, because the access was granted on purpose. That is the curated view EDR was designed to give you. It is not a view of how your data moves.

Hunt the move, not the process

Shift the question. Not "what executed?" but "what moved, on whose behalf, and does this fit?" A data movement hunt starts from the data and works back to the identity.

Picture a slow drain that trips no control. An account with legitimate read access pulls records in small, paced batches over days. EDR sees authorized reads from an authorized account. The command line is clean. Each request reads like a day's work. The operation is the rate, the breadth across tables, the destination, summed up. No single move is the breach. The shape across moves is.

Take a third-party integration with a valid token. The permission was right the day it was granted. Only the behavior changed. The integration starts reaching into systems it never touched, at volumes it never moved. Rule-based detection has nothing to fire on. A hunt that knows the integration's normal data movement sees the deviation the moment it starts.

Same story for AI agents. To a rule-break tool, an agent with broad legitimate access is invisible. To a data movement hunt, it has a profile: the systems it normally reads, the volumes it normally moves, the identities it normally acts for. The credential stays valid either way. Drift from the profile is the signal.

Deviation across every identity at once

Classic hunting techniques look for anomalies in user behavior, process relationships, network patterns. A data movement hunt aims at the moves: who moved what, where it went, whether that fits the job behind it.

Every legitimate workload carries a data movement profile. A reporting service reads a known set of tables on a known cadence. A backup job copies known volumes to known destinations. An analyst pulls samples, not the whole warehouse. Drift from the profile is either a new business process or something worth a closer look.

The leverage comes from layering. User baselines hold what one operator normally moves. Role baselines hold what their function normally moves. Workload baselines hold what a service or cluster normally moves. A move that deviates from one might be routine. A move that deviates from all three at once almost never is. The identity behind the move is resolved probabilistically and depends on the source: the more context a move carries, the more confidently it ties back to a person, a service, or an agent.

Run it on a real case. A developer account with valid credentials starts reading a production customer dataset and shipping it to an outside destination. User baseline: this account has never touched that data. Role baseline: developers in this group do not read production customer records. Workload baseline: this dataset has never left the account at this volume. Every credential checks out. The stack of deviations is what raises the hand.

Watch the move without reading the data

Fair objection: watching data movement this closely sounds like reading the data. It is not. The signal that exposes the pattern lives in the metadata of the move, not its contents.

An exfiltration has a shape you can read without ever opening a payload. A slow drain pushing a large volume out over an HTTPS POST reads as a movement pattern: which identity, which source, which destination, how much, how often, set against what is normal for that identity. You do not have to see the records to see that the wrong amount of the wrong data is heading to the wrong place. Metadata is the default and it carries the vast majority of these patterns. Content-aware inspection is there when an investigation genuinely calls for it, held in reserve rather than reached for first.

That is the unlock. You hunt across every move your organization makes, and your monitoring layer never has to read your customers' records, your source code, or your files. The pattern is in the movement. The movement is what you watch.

Where this sits next to the stack you have

Most teams already run a deep stack. EDR on endpoints. SIEM over the logs. Network monitoring at the perimeter. Posture management scanning configurations. They cover a lot. Some of it you keep, and some of it Hilt can stand in for.

What they share is the rule break. They catch the unauthorized, the malicious, the misconfigured. They leave the permitted move: the authorized account moving data it should not, the valid token reaching where it never went, the agent acting at a scale no one sized for. None of that shows up in rule-based telemetry, because no rule broke. These are anomalies in data movement, not policy violations.

Data movement governance covers the seam those tools leave open, and on the endpoint it can do more than that. Hilt can stand in for your endpoint sensor, and many clients retire their EDR once it is in place. Keep an EDR alongside only if you also want the malware and intrusion layer it provides. Keep network monitoring for traffic. When the threat is a real identity moving data the wrong way, you need something watching the movement itself and resolving it to the job behind it.

Hilt is that layer. One lightweight collector watches data movement at the kernel, metadata only by default and off the path, single-tenant in your own cloud. It resolves each move to a probabilistic, source-dependent identity, surfaces the pattern across moves instead of firing on a single event, writes the case, and responds with host-level network isolation (quarantine) from the control plane. It never sits inline. It never blocks, drops, or alters traffic. The footprint stays small enough to run everywhere: on the order of 0.1% of one core and a few megabytes of memory. Events never leave your account.

The moves that cost the most are the ones that look the most like work. A hunt that only chases rule breaks comes back clean while the data walks out on access you granted on purpose. To close that, hunt the pattern across moves, at runtime, on the movement itself, while there is still something to save. If that gap is the one your hunts keep missing, a 30-minute technical call is enough to walk through how Hilt watches it.