Conversational AI privacy requires one data-flow audit across web, mobile, backends, and third parties. A new peer-reviewed study of nine major services finds that the same account can face different tracking paths, identifiers, consent controls, and conversation exposure depending on the client. This article converts those measurements into a reproducible audit that tests what data actually moves, who can link it to a person, and whether user controls change the flow.
Reading time: 10 minutes · About 1,800 words
TL;DR
- Researchers analyzed nine web services and eight Android apps using static analysis, controlled interactions, network instrumentation, fresh profiles, and canary URLs.
- Every evaluated service contacted at least one third-party organization classified as advertising or tracking infrastructure, but presence alone does not prove an advertising purpose for every transmission.
- Six of nine web clients and three of eight mobile clients disclosed conversation-derived artifacts to third parties during normal interactions.
- Web and mobile are different privacy surfaces. Web exposed more URLs, titles, cookies, and sharing-page data; mobile exposed advertising, account, installation, and device identifiers.
- A credible audit must join permissions, traffic, identifiers, recipients, access control, retention, and user actions into one evidence ledger.
The product boundary is larger than the chat window
Users see one assistant brand. Their data moves through several systems: a browser or native app, first-party APIs, analytics SDKs, error monitoring, advertising infrastructure, identity providers, sharing pages, and server-side forwarding.
Auditing only a privacy policy misses implementation. Auditing only app permissions misses web identifiers and server-side flows. Auditing only browser cookies misses mobile SDKs and device-linked identifiers.
The cross-channel unit of analysis should be:
account × client × consent state × subscription tier × interaction × recipient
That unit turns privacy from a list of settings into a flow that can be replayed and compared.
What the new study measured
The paper, Prompt like a Butterfly, Sting like a Tracker, is scheduled for the 2027 Privacy Enhancing Technologies Symposium and is marked peer reviewed and in press by IMDEA Networks. The authors studied ChatGPT, Claude, Grok, DeepSeek, Perplexity, Gemini, Microsoft Copilot, Mistral Le Chat, and Meta AI. All nine had web clients; eight had Android clients.
For web, researchers used Chrome Developer Tools Protocol traces to record requests, storage access, cookies, JavaScript execution, and tracking pixels. Each condition started with a fresh browser profile. The team compared ignoring, rejecting, and accepting non-essential cookies, plus guest, free, and paid tiers where available.
For Android, they combined APK static analysis with an instrumented Android 12 device. The setup observed permission-protected APIs, files, loaded classes, and outbound traffic, including certificate-pinned apps. Known pseudonymous identifiers let the researchers search traffic for plaintext and common hashes.
The experiments used predefined health-related prompts, normal conversation, and shared-link actions. Canary URLs in prompts and uploaded files tested whether external systems later accessed submitted resources.
This is a strong black-box design, but it remains a point-in-time lower bound. Each deterministic configuration was run once after pilot testing. Gemini mobile traffic could not be extracted. Enterprise and government tiers, desktop clients, voice interfaces, geography variation, long-lived memory, and invisible server-side processing remain outside the measured scope.
The findings need careful separation
The paper identified 124 third-party domains attributed to 44 organizations, including 34 advertising and tracking services. Every evaluated assistant contacted at least one organization in that category.
That headline needs a qualifier from the authors: a third-party component can provide operational functions such as authentication, payments, fraud prevention, telemetry, or error monitoring. Observing a domain or SDK does not independently establish the purpose of every event.
The more decision-relevant finding concerns conversation artifacts:
| Observed behavior | Web | Android |
|---|---|---|
| services disclosing conversation-derived artifacts during normal interaction | 6 of 9 | 3 of 8 |
| conversation URL disclosed | 5 providers | 0 |
| conversation title disclosed | 3 providers | 0 |
| prompt disclosed | 1 provider | 1 provider |
| conversation screenshot disclosed | 1 provider | 0 |
The web and mobile surfaces differ. Browser trackers can observe page URLs, titles, cookies, and application state. Mobile SDKs can combine account IDs, hashed emails, advertising IDs, installation IDs, and device identifiers. The study found that 71% of distinct endpoints contacted by mobile clients originated from WebViews, so the boundary between web and native tracking is porous.
Consent and payment do not define one consistent boundary
Rejecting non-essential cookies reduced some web disclosures, but it did not remove every third-party connection. The paper reports that four of nine free-tier web services still sent data to third-party trackers after rejection. It also found little consistent difference between free and paid tiers in the set of contacted third parties.
Mobile had a different control model. Most apps required acceptance of terms and privacy policies before use, without an equivalent per-category cookie decision at launch.
This creates a practical audit rule: never copy a conclusion from one channel to another. A web setting may have no mobile equivalent. A paid plan may change product features while leaving analytics infrastructure similar. A rejected cookie banner may reduce optional advertising calls while operational telemetry remains.
The control must be tested as a before-and-after flow, not accepted as a label.
Identity linkage turns metadata into conversational exposure
A conversation ID alone can look harmless. Combined with a stable account, device, cookie, or advertising identifier, it becomes a durable join key.
The study observed combinations that included user IDs, anonymous IDs, organization IDs, email addresses or hashes, session IDs, advertising cookies, Android Advertising IDs, installation IDs, and SDK-specific identifiers. Resettable identifiers can also be rebound when transmitted with a persistent account or installation ID.
An audit should therefore model linkability, not only raw sensitivity. For each event, record:
- conversation artifact present;
- direct or pseudonymous identifier present;
- recipient and corporate owner;
- purpose claimed in documentation;
- consent state and account tier;
- persistence across sessions;
- ability to reconstruct or access the conversation.
This is closely related to a repository or agent data-flow audit: the decisive question is which information crosses a trust boundary, under which identity, and with what downstream authority.
Shared links are an access-control system
All tested services supported a mechanism for intentionally generating a publicly accessible shared conversation. The paper found nine third parties present on conversation-sharing pages, where page context can expose the conversation.
More concerning, access defaults varied by provider and tier. The researchers observed public-by-default or opt-out conversation permalinks in some configurations. In Grok sharing flows, they report Meta and TikTok receiving the latest prompt, and TikTok receiving a conversation screenshot.
Canary tests added evidence about subsequent access. The study observed 70 Grok-related activations over hours to days, from 70 IP addresses across 48 autonomous systems and 14 countries. It did not observe the same repeated pattern for other services, and explicitly warns that absence of observed access does not prove absence of access.
Shared-link auditing therefore needs four tests:
- unauthenticated access in a fresh session;
- revocation and cache behavior;
- third-party requests made by the sharing page;
- later access to embedded canary resources.
Build the audit as an evidence ledger
A reproducible cross-channel audit can follow seven stages.
- Freeze the matrix. Record versions, region, account tier, consent state, device, browser, and timestamp.
- Use fresh identities. Create isolated accounts and known pseudonymous values that can be searched safely in captured traffic.
- Replay the same tasks. Use a fixed prompt sequence, upload set, share action, and deletion action across channels.
- Capture at multiple layers. Collect browser HAR, storage, mobile static SDK inventory, runtime network traffic, permission access, and screen recording.
- Normalize recipients. Map domains, SDKs, proxy endpoints, and server-side forwarding to organizations while retaining uncertainty about purpose.
- Join artifacts to identifiers. Flag events where prompts, titles, URLs, screenshots, or conversation IDs travel with stable identity signals.
- Verify user controls. Repeat after rejection, opt-out, plan change, share revocation, account deletion, and identifier reset.
The ledger should distinguish four evidence states: declared by policy, present in code, observed in traffic, and verified through downstream access. This prevents a policy promise or an SDK inventory from being mistaken for demonstrated behavior.
What organizations should require
Teams procuring conversational AI should request a channel-specific data-flow map, not a single privacy summary. The map should identify every processor and tracking service, fields transmitted, purpose, legal basis, retention, region, account linkage, deletion propagation, and enterprise-versus-consumer difference.
They should also test controls independently. A contract saying enterprise data is excluded from training answers one question. It does not answer whether operational telemetry contains conversation identifiers, whether shared pages expose content, or whether deletion reaches third-party processors.
For sensitive use, the lowest-friction risk reduction is architectural: separate consumer and organizational accounts, disable public sharing, reject optional tracking, restrict mobile permissions, avoid placing secrets in prompts, and route high-sensitivity tasks through products with auditable organizational controls.
FAQ
Do all AI chatbots send conversation text to advertisers?
The study does not support that claim. It observed different artifacts, recipients, and purposes across services. Six web clients and three Android clients disclosed some conversation-derived artifact during normal interactions; direct prompt disclosure was much less common.
Does rejecting cookies stop AI chatbot tracking?
It reduced some optional web flows but did not eliminate every third-party connection in the study. The result depended on provider and channel, and mobile apps generally lacked an equivalent cookie-choice flow.
Is a paid AI plan more private?
Payment alone was not a reliable boundary in the measured third-party infrastructure. Enterprise tiers may have distinct contractual controls, but those tiers were outside the study and require separate verification.
Why are conversation URLs sensitive?
They can identify or retrieve a specific conversation. When transmitted with a stable user or advertising identifier, they may connect conversation context to a long-term profile.
What are the study's largest limitations?
It is a black-box snapshot of nine services from Spain in May 2026, with one run per stable configuration after pilots. It excludes several products and channels, cannot see all server-side processing, and could not extract Gemini mobile traces.
References
- Oliveira et al., Prompt like a Butterfly, Sting like a Tracker
- IMDEA Networks, accepted manuscript PDF
- Jazlan et al., Tracking Conversations: Measuring Content and Identity Exposure on AI Chatbots
- European Data Protection Board, Guidelines on Article 5(3) of the ePrivacy Directive
- OWASP, LLM02:2025 Sensitive Information Disclosure
Start with one account and one repeatable prompt sequence. Run it through web and mobile, reject optional tracking, then compare the actual recipients. The differences are the beginning of the privacy model.