A computational biologist gives two weeks' notice. Over the next ten days she pulls the full compound-activity dataset in small batches, off-hours, through the same approved sync she uses every morning. No policy fires. No rule breaks. Each pull is something she is allowed to do. The day she starts at the competitor, the structure-activity data is already in their pipeline, and you cannot un-disclose a target.
That is how research data leaves a biotech. Not through a broken door. Through one you held open on purpose.
Shared access is the business, not a flaw
The CRO needs the raw reads to run the assay. The academic collaborator needs the structure to do the chemistry under the MTA. The partner needs the compound library for the co-development deal. The cloud workload needs to pull from the data lake to train. Lock any of these and the science stops. So you grant the access, and you should.
Then every tool you own that checks permission agrees with the access. Identity and access management confirms the user is who she says. DLP checks the file class against the destination and approves it. The cloud posture tool reports the bucket policy is correct. All three are right. The move was permitted.
A breach hides where none of them looks: the pattern across permitted moves. The departing researcher's batched off-hours pulls violate no policy on any single pull. A CRO connection that normally returns processed results, suddenly reaching back across paths it has never touched, is using access it was granted. The shape of the movement is the breach. The permission behind it is a red herring.
Predictive guesses early; forensics reads the wreckage
DLP guesses in advance. It classifies data, writes rules about where each class may go, blocks the rest. A research environment punishes this. A novel assay format is not in the classifier. A fresh collaboration looks like exfiltration. The team drowns in false alarms, loosens the rules to keep the science moving, and the permitted move sails through anyway, because by construction it was permitted.
Forensic and detection tools read the wreckage. When the partnership sours or the researcher lands at a competitor, the logs reconstruct what left. By then the dataset is gone. For biotech IP, gone is the whole loss.
Neither tool is dishonest. Each answers the question it was built for. Neither watches the movement itself, as it happens, and asks whether this pattern fits how this data normally moves. That question only has an answer at runtime, while the data is in flight, before it lands somewhere you cannot reach and after the guessing has stopped.
What governing the movement looks like
Hilt is runtime Data Movement Governance. One lightweight collector watches data movement at the kernel, where every move is visible no matter which app, channel, or cloud carried it. It runs metadata-only by default. It sees the pattern is wrong without reading the sequence data, the structures, or the assay results. Content-aware inspection is there when an investigation calls for it, never as the price of admission. The collector stays off the path: on the order of 0.1% of one core and 4 to 8 MB of memory per host, single-tenant inside your own cloud. It does not block, drop, or alter a single transfer.
Each move resolves to a probabilistic, source-dependent identity. Which user or workload, which job behind it, which destination, and whether this fits what that identity normally does with your data. When the compound-activity pulls start, the deviation surfaces across layers at once. The job is wrong for that researcher. The access reads high-value paths broadly in a short window. The volume to that destination has no precedent in months of history, even on an approved channel. Any one signal is noise. Together they hold up as a case, not an alert.
When the pattern crosses the line, Hilt quarantines the host at the network from the control plane. It never sits between your data and the instrument cluster waiting to drop packets. It watches the move, writes the case with the evidence already correlated, and isolates the host so the rest of the dataset stays put.
It sits beside your stack, not on top of it
Keep what you run. Your IAM still proves who is connecting. Your DLP still catches the careless exfiltration, the structure file emailed to a personal account, the bulk download to a USB drive. Your cloud posture tool still keeps the buckets configured. Those controls evaluate permission well. Let them.
None of them was built to evaluate the pattern of movement across permitted actions. That is the layer Hilt adds. It assumes the access is legitimate, because in research it almost always is, and watches what that access does with your most sensitive data over time. The CRO that overreaches. The collaborator integration pulling more than the MTA scoped. The AI workload that starts moving training data somewhere new. All of it lives in the gap between permitted and normal.
A biotech concentrates its value in a way most companies do not. One target, one structure-activity dataset, one negative result that saved two years can be a whole program. That value moves every day, across collaborators and CROs and clouds, on access you granted because the work demands it. Whether each move was allowed is not the useful question. What this data actually did, resolved to the job behind it and scored against how it normally moves, is.
If your stack cannot answer that, the gap is open and it is sitting on your pipeline. To see how the collector watches movement at the kernel without reading your research data or touching the path, the next step is a 30-minute call, engineer to engineer.