Technical

Metadata by Default, Content-Aware When You Want It

February 23, 2026 Hilt 7 min

Your most valuable data leaves on access you granted on purpose. You can turn on content-aware inspection when a team wants it, single-tenant in your own cloud, but you never need it to surface the pattern. The default is metadata only.

Metadata by Default, Content-Aware When You Want It cover image

Most data security tools open the bag. To know whether data is leaving where it should not, they read the file, match the regex, classify the payload. The customs officer model: inspect every bag to find the one that is wrong.

That model answers the wrong question. "What is this data" requires looking inside. "Is this move wrong" usually does not. Your most valuable data leaves on access you granted on purpose. The credentials are real, the channel is approved, the file is one this person is allowed to touch. Every move is permitted. The breach is the pattern across moves, and the pattern lives in the shape of the movement, not the contents of the payload.

The shape gives it away, not the bytes

Take one move. Without opening the payload you already see which identity initiated it, resolved probabilistically from the source. Which job or process ran it. Which path it read from. How much moved. Where it went. When. And how all of that compares to what this identity normally does, scored against months of how data actually moves in your environment.

A quant researcher pulls proprietary code in small off-hours chunks through an approved transfer channel. No single file announces her. The shape does: a bulk read of high-value paths, in an odd window, at a volume that does not fit her baseline, through a channel that is technically allowed. Reading a byte of the code adds nothing. The deviation is already in the metadata, and the metadata is enough to write the case.

The split is structural. Content tells you what something is. Metadata tells you what is happening. Exfiltration is something happening. So the dangerous pattern sits in the metadata, and you do not have to read your data to see that a move is wrong.

Why metadata is the floor, not a downgrade

A system that reads content at rest has copied or indexed your most sensitive data somewhere. That somewhere is now a target, a compliance question, and the first thing a customer asks about in their security review. Hilt watches data movement at the kernel, metadata only by default. The events never leave your account. There is no content corpus to defend, because none was built.

Payload inspection also costs compute that scales with volume. A metadata-first collector stays off the path and stays small, around 0.1% of one core and 4 to 8 MB of memory per host. It observes the move instead of standing in it. It never sits inline. It never blocks, drops, or alters traffic. Latency-sensitive infrastructure can run it at all only because the resting state is metadata, not deep inspection of everything that moves.

And here is the part most teams miss. Content-based detection mostly catches the move that should never have been allowed in the first place. Match the regex for a customer record and you flag the channel you would have closed anyway. The move that was permitted on purpose passes every content rule, because the content is exactly what this person is cleared to handle. Reading the file confirms it is the sensitive file. It does not tell you the read was anomalous. Only the pattern does that.

So when do you reach for content

You can turn content-aware inspection on. It is real and it is useful. It is a choice you make for a reason, not the toll you pay to see exfiltration at all.

Three reasons a team flips it on. You have surfaced a case and want to confirm which classes of data were in the moves before you escalate. A regulator wants positive classification of what crossed a boundary, not just evidence that one was crossed. A workload has a genuinely ambiguous metadata signal and content would sharpen the baseline. In those spots, you scope it on, on that surface, for that reason.

When you do, it runs like everything else: single-tenant, inside your own cloud, AWS, GCP, Azure, or Ali Cloud. The inspection happens in your account. Enabling it routes nothing through a vendor SaaS, and it does not change the default for the rest of your estate. You add a deeper lens on one surface, deliberately. You do not flip the whole system into a mode where it reads everything.

Content-aware is a capability you reach for on the surfaces where it earns its cost. It is never the resting state, and never required to surface the pattern.

The question that sorts the vendors

Almost every data movement tool can inspect content. So "can it inspect content" tells you nothing. The clarifying question is: what is its default vantage, and does it have to read your data to do its core job.

If content inspection is how it detects anything at all, you are buying a system that, by design, looks inside your most sensitive data to function, carries every privacy, performance, and blast-radius cost that implies, and still misses the permitted move. If the default is metadata with content on demand, you are buying a system that surfaces the dangerous pattern without reading your data and goes deeper only where you decide the depth is worth it.

Different architectures, not different settings. You cannot turn a content-first system into a metadata-first one by flipping a feature off. A metadata-first system becomes content-aware exactly where you want, because the floor was set right.

So: you do not need to read your data to detect exfiltration. The dangerous move is the permitted one, and the permitted one beats content rules by definition. What gives it away is the pattern, the identity, the job, the volume, the destination, the timing, scored against how your data normally moves. That is metadata, and metadata is the right default. Content is the lens you choose, single-tenant in your own cloud, on the day a team wants it.

If you want to walk through where metadata is enough for your environment and where you would actually reach for content, that is a 30-minute technical call, engineer to engineer. We can map it to your real workloads on the call.