Guide

Watching Data Move Across AI Agents and Tools

February 24, 2026 Alexandre Genest 8 min

Your most valuable data leaves on access you granted on purpose. A paste into a chatbot or an API call to a model endpoint is a data movement before it is an AI event. Why governing data movement covers AI tools you never enumerated.

Watching Data Move Across AI Agents and Tools cover image

Your AI governance plan is a list. The approved chatbot. The sanctioned coding assistant. The model endpoint that cleared procurement. You enforce the list wherever you can reach it: a proxy, a browser extension, a CASB rule, an allowlist on the egress firewall. The list is the policy.

The list is always one tool behind. New ones ship weekly, an employee adopts one over lunch, and a model endpoint is a single API key away. Chasing them is the wrong job. The job is to notice that a paste into a chatbot is a data movement before it is an AI event, and an agent calling a model API is a data movement before it is an agent action. Watch the movement at the source and the list stops mattering, because you stopped depending on it.

Shadow AI is not a discovery problem

It gets treated like one. Find the unsanctioned tools, add them to a list, repeat. Buy a scanner that watches SaaS logs and network traffic for known AI domains, then chase the long tail forever.

That catches the tools you already knew about. It cannot catch the reason the category exists, which is that an AI tool is trivial to adopt and trivial to reach. A developer pipes a customer dataset into a local model. An analyst drops a contract into a browser tab. An agent your own team wrote reads from a production database and posts to a model endpoint that was never on any list, because it is yours.

Every one of these is permitted. The proxy waves the paste through because the destination is on the allowlist, or has not made the blocklist yet. The agent's call succeeds because the credentials are ones you issued on purpose. No rule broke. What changed was the pattern: what moved, where it went, which identity sent it, which job was behind it. That pattern reads the same whether the destination is a sanctioned tool or one you have never heard of.

The same event, seen from below

Hilt does not start from the list. It starts from the data and watches data movement at the kernel, metadata only by default, off the path. From there, the AI tools you enumerated and the ones you never did collapse into one shape: bytes leaving a host toward a destination, on behalf of an identity, through a job.

Take the case your allowlist loves most. A researcher pastes proprietary code into a sanctioned assistant. Approved destination, so the proxy waves it on. At the kernel it is still a bulk read of high-value paths flowing to an external endpoint, resolved to a person and a job. Whether that fits what the researcher normally does is a different question than whether the tool was approved, and it is the question that matters.

Now the agent your team built. You granted it read access to a customer database and an API key to a model endpoint. Both grants are correct. Then it drifts. It reads more, reaches paths it never touched, ships the result out at a volume with no precedent. No credential was misused. The job changed, and the job changing is the whole signal.

And the brand-new tool on no list at all. Nothing to enumerate, no domain you have ever seen. The move still surfaces at the source the instant it happens, resolved to who sent it and why.

In all three the AI tool is incidental. The data movement is the event, and the event is where the signal lives.

Permitted is not normal

Every move is permitted. The pattern is the breach. Your proxy, your CASB, your firewall, your AI gateway each answer one question: is this move allowed. They answer it well. None of them was built to ask whether an allowed move fits how this identity's data usually behaves.

That second question is Hilt's only question. Each move resolves to a probabilistic, source-dependent identity: which user or service, which job, which destination. Then it scores against months of how data actually moves in your environment. A quant copies a model in small off-hours chunks, through an approved channel, into a sanctioned assistant. No single action trips a rule. But the timing is wrong for that identity, the access is a bulk read of high-value paths in a tight window, and the volume to the destination has no match in the history. One of those is noise. Stacked, they are a pattern, and a pattern is a case, not an alert.

Agents are where most AI plans go quiet, and this is exactly where they get caught. An agent is a non-human identity moving data on a schedule. A permission check passes it every time, because the access was granted and the credentials are valid. A pattern check catches the drift. An agent that starts exfiltrating reads as an anomaly in how data moves, not as a policy violation that never fires.

You do not read the prompt to see the pattern

The real objection to watching AI usage is privacy. Read the paste to catch the paste and you have built surveillance, plus a second copy of the exact data you were trying to protect.

Hilt's default is metadata only. It sees that a move happened, its shape, its identity, its destination, and its fit against history. It does not read the prompt or the file. You learn that a researcher's data did something strange on its way to a model endpoint without learning a word she typed. Content-aware inspection exists for the investigation that demands it. It is never the price of seeing the pattern, and never the default.

The collector stays off the path. Single-tenant inside your own cloud, on the order of 0.1% of one core and 4 to 8 MB of memory per host. It does not sit inline between your data and a model endpoint, and it does not block, drop, or alter a prompt in flight. When a move resolves to a real anomaly, the response is host-level network isolation (quarantine) from the control plane, not filtering on the request. Events never leave your account.

What you already run still earns its keep

This replaces nothing. Not your AI gateway, not your DLP, not your CASB, and it is not an argument against enumerating tools. A gateway that routes and rate-limits sanctioned model traffic does real work. A CASB that catches the obvious upload to an unsanctioned app catches obvious uploads. Keep them.

They share one vantage: the application or network layer, where the decision is "is this destination allowed." Hilt sits a layer under that, on the data movement itself, at the source, resolved to identity and job, scored against normal. That is where the permitted-but-anomalous move lives, and it does not care whether your list of AI tools is complete, because it never read the list.

So what you gain is not one more shadow AI domain found. It is this: any move of your data toward any model endpoint, sanctioned or not, human or agent, surfaces at the source and is judged on its pattern instead of its permission.

Govern AI as a list and you are enumerating faster than people adopt, which you cannot do, and your own agents are tools you built. Govern the data movement and the list never had a vote. The paste, the API call, and the agent's read are one event seen from the kernel, and that event is where the dangerous pattern becomes legible.

Here is the test. Ask your current stack what this identity's data actually did on its way to a model endpoint, resolved to the job behind it, scored against how it normally moves. If it cannot answer, your AI coverage ends exactly where your list does.

If that sounds like your environment, the fastest way to see the gap is thirty minutes, engineer to engineer, on your own data.