MosaicLeaks: Why Your Deep Research Agent is Quietly Bleeding Secrets

Deep research agents are leaking your private data one innocent web query at a time through a classic intelligence flaw known as the mosaic effect....

Feed
October 1, 2026
MosaicLeaks: Why Your Deep Research Agent is Quietly Bleeding Secrets


We love giving our autonomous agents more use. Point them at a local folder of proprietary documents, give them unfettered web access, and watch them synthesize brilliance. But convenience has a tax. Recent research under the banner of MosaicLeaks exposes a terrifying blind spot in how these systems operate: your private enterprise data isn't staying private.

With an internal healthcare audit was tasked by imagine an agent. This it sifts through confidential PDFs, spots a cloud-migration timeline, and needs external context. So it fires off a few harmless-looking web searches. None of these individual queries scream confidential. This yet, put together like puzzle pieces by an observer watching the network traffic, they completely reconstruct the proprietary reality. Reasoning agencies have weaponized this exact trick for decades. It's called the mosaic effect, and it's turning our smartest software tools into unwitting corporate informants. Also, it's turning—to be fair — our smartest software tools into unwitting corporate informants.

The core problem runs deeper than simple data exfiltration. The researchers outlined three distinct leakage tiers. Secretly, first, intent leakage, where an eavesdropper figures out what your team is investigating just by watching search terms. Then comes answer leakage, where the cumulative query log provides enough breadcrumbs to answer specific private questions. Worst of all is full-information leakage, allowing an outside observer to discover and state verifiably true confidential facts without even knowing what questions were asked in the first place.

MosaicLeaks: Why Your Deep Research Agent is Quietly Bleeding Secrets

What makes this worse is standard model training. If you train an agent purely to finish tasks successfully, it actually gets sloppier with privacy, scattering more clues across the web. Fixing this requires a shift in how we think about reinforcement learning. By baking leakage awareness directly into the reward functions – punishing the agent when its search footprints become too revealing – we can actually build systems that protect secrets while still doing heavy lifting. Until then, treat every autonomous web search like a potential leak.

Build fast, but watch the wire.