What happened
- A paper published on arXiv on September 24 presents PrivDrift, a benchmark for measuring whether sensitive data revealed to an assistant remains retrievable after the conversation has changed subject.
- The dataset has 1,000 multi-turn dialogues, with seeded secrets, content-heavy drift turns and standardized extraction probes.
- Across three models with extended context windows, dialogue-level leakage ranges from 38.7% to 54.6%, with strong variation by model, type of secret and the intensity of the persuasion used to extract it.
- The central finding, according to the text: continuing to change the subject doesn’t reliably reduce leakage within the window tested.
- The author concludes that this should be evaluated as a persistent behavioral failure, not as memorization of training data or as a one-off leak caused by filter evasion.
Why it matters
- It squarely affects any organization that has opened a shared assistant to its team. If two people use the same session, or if a long session covers several matters, data provided at the start remains available to whoever knows how to ask.
- In Chile this falls directly under Law 21.719: data revealed in a work conversation is personal data if it identifies someone, and whoever administers the tool is responsible for its processing. Closing the tab isn’t a deletion measure.
- There’s a consequence the paper doesn’t address: most corporate controls review the input message and the final answer. None audits whether the accumulated context still contains material that should already have been discarded, because there’s no record of what was said twenty turns ago.
The number
54.6% is the ceiling of dialogue-level leakage measured across the three models evaluated.
Context
It’s the third piece this month on the same blind spot. The note on the lethal trifecta in ChatGPT for work placed the risk in the combination of private data, external content and the ability to send. PrivDrift shows the first of the three is enough.
What’s next
- No timelines announced for publishing the dialogue dataset or for extending the evaluation to more models.
- The paper is in version 1 and doesn’t state peer review.
Bottom line
The public debate in Chile about postponing the data protection law was framed in terms of compliance deadlines. The problem this paper measures doesn’t wait for the deadline: it’s already in the context window of the sessions opened today.
This note describes a technical finding and does not constitute legal advice.
Sources
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.


