Administrator
Published on 2026-09-15 / 17 Visits
0
0

LLM Relay Security: Audit the Data Supply Chain

An LLM relay can see the plaintext prompt, choose an upstream endpoint, and reshape the response before your client receives it. The model name in the UI therefore describes a requested capability, while the route determines the actual trust boundary. This guide turns that route into an auditable contract for production data and tool-using agents.

Reading time: 12 minutes | Length: about 2,400 words

TL;DR

  • Treat every LLM relay as a data supply-chain node, because each TLS-terminating hop can process prompts, credentials, tool definitions, and returned tool calls.
  • Separate confidentiality from integrity. A zero-data-retention policy governs storage; it does not prove that an intermediary preserved the upstream response.
  • Require per-request route evidence, strict provider allowlists, disabled fallbacks for sensitive workloads, scoped credentials, local action gates, and append-only logs.
  • Give production access by data sensitivity and action authority. A clean finite audit earns bounded confidence, while high-impact agents still need fail-closed controls.

A model name is not a data path

Changing an SDK's base_url can make dozens of models available behind one API. The application still displays a familiar model name, and the client still receives valid JSON. Operationally, however, the path may look like this:

application
  -> relay or reseller
  -> second-tier aggregator
  -> hosting provider
  -> model endpoint
  -> the same chain in reverse

Each intermediary that terminates one TLS connection and starts another receives application-layer plaintext. That can include system prompts, attached document text, tool schemas, tool arguments, environment fragments, and the structured action returned by the model.

This produces two separate security questions:

  1. Confidentiality: who can read, retain, classify, resell, or train on the request?
  2. Integrity: who can modify the request or returned tool call before execution?

HTTPS authenticates the endpoint your client selected. It does not authenticate the semantics of a response across a multi-hop application path. A valid certificate for a relay proves that you reached that relay.

The distinction matters most for agents. A rewritten paragraph can mislead a person; a rewritten Bash or run_command argument can change the machine directly. Model-level jailbreak defenses operate inside inference, while relay-side rewriting can happen before the model sees a request or after it emits a response.

Read the evidence at the right strength

Three evidence layers establish the risk without turning every relay into a presumed attacker.

First, the April 2026 paper Your Agent Is Mine examined 28 paid and 400 free commodity routers. The authors reported malicious code injection in 1 paid and 8 free routers, adaptive triggers in 2, access to researcher-owned AWS canary credentials from 17, and one transfer from a researcher-owned Ethereum key. The on-chain loss was below $50 by experimental design.

The same study used a research proxy against four agent frameworks. Its rewritten tool calls arrived in framework-native formats, but these were compatibility measurements, not end-to-end compromise rates. Local permission prompts and sandboxes could still stop execution. Its defense results also came from generated, controlled corpora rather than production traffic.

Second, Real Money, Fake Models audited three representative shadow APIs selected from a larger set of 17. It found large utility divergence and identity-verification failures in 45.83% of fingerprint tests. The authors lacked backend ground truth, so fingerprints, metadata, performance, and safety behavior establish inconsistency signals rather than a courtroom-grade identity verdict.

Third, component trust can change after procurement. Datadog Security Labs reported that real PyPI releases of LiteLLM, versions 1.82.7 and 1.82.8, contained malicious code on March 24, 2026. This incident demonstrates a supply-chain compromise of a legitimate routing component. It does not establish exposure for every LiteLLM deployment.

Anthropic's September 2026 threat report adds a provider-side view. Anthropic says it observed customer requests silently forwarded from other services to Claude, including requests containing corporate data, source code, and live credentials. Those named-actor findings are Anthropic's high-confidence threat-intelligence attribution. They should remain labeled as vendor findings until independently confirmed.

The combined conclusion is precise: a relay is an intentional plaintext intermediary, malicious behavior has been measured in a commodity sample, identity inconsistency has been measured in selected shadow APIs, and legitimate gateway software can be compromised. These facts justify controls. They do not justify a universal claim that every aggregator is unsafe.

Draw the relay as a supply chain

A vendor questionnaire that asks only which model is used captures the last box and misses the route. Build an inventory of every processing hop instead.

Supply-chain object Evidence to collect Risk when missing
Client and SDK Version, configured endpoint, plugins, local logs Hidden history replay or client-side storage
Relay operator Legal entity, service domain, subprocessors, jurisdictions No accountable data processor
Routing layer Provider allowlist, fallback rule, policy version Silent endpoint or model changes
Transformation layer Compression, guardrails, caching, server tools Prompt or response semantics change
Credential plane Key owner, scope, budget, rotation, upstream account Cross-tenant exposure and unattributed spend
Model endpoint Actual provider, model/version, region, retention policy Requested model differs from serving endpoint
Response path Request ID, attempts, response hash, action decision No incident reconstruction or tamper signal

The key unit is the endpoint, not merely the model family. A provider's default retention policy may differ from a particular hosted endpoint or enterprise agreement. Plugins, web search, file storage, observability exports, and client memory add their own processors and retention rules.

OpenRouter's current documentation illustrates why configuration belongs in the evidence set. Its provider-routing guide supports ordered providers and disabling fallbacks. Its ZDR guide says policies are endpoint-specific, unknown policies receive a conservative classification, plugins and tools sit outside inference ZDR, and in-memory prompt caching still fits OpenRouter's definition of zero retention. These are operator-documented controls and definitions, not independent assurance.

Require a minimum route evidence contract

A production route should generate evidence for six promises. Each promise needs a failure behavior as well as a configuration value.

Contract field Minimum evidence Fail-closed behavior
Route identity Requested model, selected provider/endpoint, region, attempts Reject an endpoint outside the approved set
Route policy Allowlist and policy version or hash Reject an unknown policy or silent fallback
Data use Retention, training, caching, abuse review, subprocessors Block data above the approved classification
Transformations Compression, guardrails, server tools, schema conversion Escalate an undeclared transformation
Credential scope Workload identity, budget, expiry, upstream account class Revoke or stop a key that crosses scope
Audit linkage Local request ID, relay ID, provider ID where available, redacted hashes Quarantine a response that cannot be linked

Some commercial routers expose parts of this contract. For example, OpenRouter's opt-in router metadata can report the requested model, selected endpoint, strategy, fallback attempts, BYOK status, and material pipeline stages. Its guardrails can restrict providers and models, enforce ZDR, and block or redact sensitive patterns.

Capture these fields in your own telemetry. A dashboard inside the intermediary is useful for operations, while an independently retained, redacted record is useful after the intermediary itself becomes suspect.

Two fields deserve special caution:

  • ZDR: retention and training policies govern what happens after processing. Every in-path service still processes plaintext at request time. ZDR also leaves integrity unanswered.
  • Response hash: a client-side hash proves which bytes the client received. Without a provider-signed canonical response, it does not prove which bytes the upstream model produced.

The April study says major tool-use APIs and the current MCP specification do not expose deployed provider signatures over tool-call arguments. Current client controls can reduce risk and preserve evidence; they cannot establish end-to-end provenance.

Grant access by data sensitivity and action authority

Relay approval works better as a matrix than a binary safe label.

Tier Typical workload Route requirement Agent authority
0 Public text, disposable experiments Known operator; basic route log No external side effects
1 Internal, low-impact content Approved endpoints; retention policy; scoped key Read-only tools or human approval
2 Confidential code or business data Contracted processor; ZDR where required; fixed providers; fallback disabled; DLP Sandboxed tools; narrow allowlists; explicit commit gate
3 Live credentials, regulated records, production administration, signing Direct contracted endpoint or fully governed private gateway; strong regional and audit controls Dedicated identity; least privilege; independent approval; rapid revocation

The matrix combines two dimensions that teams often mix together. Data sensitivity determines who may see the prompt. Action authority determines what a modified response can do.

A Tier 0 chatbot and a Tier 3 coding agent can call the same model while requiring entirely different routes. The model is a capability choice; the route is a governance choice.

Live secrets should stay out of prompt bodies. Secret managers can inject credentials at the action boundary after the model has proposed an operation. This reduces what the model, relay, provider, logs, and error handlers can expose. The same principle underpins an installation trust boundary for coding agents: the model can suggest an action, while a deterministic local gate decides whether it may execute.

Run a verification ladder

A single clean probe measures one moment and one trigger condition. The paper observed conditional behavior after warm-up counts, for selected languages, and in autonomous modes. Build a recurring ladder instead.

  1. Verify the accountable entity. Record the legal operator, domain ownership, subprocessors, support path, incident-notification term, privacy policy, DPA, and applicable jurisdiction. Treat an unexplained reseller chain as a different risk class from a contracted managed service.

  2. Freeze the route. Use provider and model allowlists, disable fallback for sensitive jobs, and make unknown endpoints fail closed. Test failure behavior by making an approved endpoint unavailable in staging. Success means a visible error, not an invisible route expansion.

  3. Use synthetic canaries. Send unique, non-production markers through a scoped test account. Monitor for unexpected reuse or appearance in returned content. Never place live cloud keys, wallet keys, customer records, or proprietary code in an audit corpus.

  4. Collect consistency signals. Compare returned model metadata, request IDs, latency distributions, behavior on a frozen benchmark, stream structure, and billing records. Identity self-reports and fingerprints can reveal contradictions; corroborating sources are required for a provider-level accusation.

  5. Constrain actions locally. Put shell, package installation, network writes, database changes, and signing behind deterministic allowlists and human or policy gates. Sandboxing limits blast radius; package pinning and hashes help with dependency identity. Each control covers a different failure mode.

  6. Preserve redacted evidence. Keep timestamps, endpoint, route metadata, policy version, tool name, approval result, token/cost data, and salted or keyed hashes. Store prompt bodies only when a defined debugging need, access policy, and retention schedule justify them.

  7. Exercise revocation. Rotate a route key, disable a provider, remove a team member, and trace affected sessions. Measure time to stop traffic and time to identify every dependent workload.

An open-source tool such as API Relay Audit v2.4.0 can collect local anomaly signals. Review and pin the code first, use a low-limit test credential and synthetic traffic, and retain its stated boundary: the report is evidence from a finite run, not a safety certificate.

This ladder also clarifies what each control can prove. Contract review establishes promises. Route metadata establishes what the intermediary reports for a request. Canaries reveal some unauthorized use. Local gates stop defined action classes. Logs support incident scope. Provider-backed signatures would add origin integrity when the ecosystem deploys them.

Frequently asked questions

Does an LLM relay keep prompts?

Policies vary by relay, endpoint, account setting, and enabled feature. OpenRouter, for example, documents prompt and response storage as opt-in at its own layer, while its privacy policy says inputs are transmitted to selected providers whose practices also apply. Verify the exact endpoint, logging settings, plugins, observability exports, and contract used by your workload.

Is zero data retention enough for production secrets?

ZDR addresses retention. The service still processes data in transit, metadata may remain, tools may follow separate policies, and a compromised intermediary can modify a response. Production secrets also need minimization, endpoint controls, integrity controls, scoped identities, and revocation.

Does HTTPS prevent a relay from changing tool calls?

HTTPS protects each transport segment. A configured relay legitimately terminates the client's TLS connection and creates another connection upstream, so it can process application-layer content. End-to-end tool-call integrity requires evidence beyond ordinary TLS to the relay.

Can model fingerprints prove substitution?

Fingerprints, self-identification, latency, output quality, and metadata are consistency signals. They become stronger when several independent signals agree and remain stable across a frozen benchmark. Backend records or provider-signed provenance would support a stronger identity claim.

Is OpenRouter safe for sensitive data?

The answer depends on the configured endpoint, provider policies, organizational contract, region, fallbacks, logging, plugins, and your application's own controls. OpenRouter documents provider allowlists, ZDR, guardrails, and route metadata that can make a route more governable. A production decision should verify those settings against the data tier and retain evidence outside the platform.

Does a self-hosted gateway remove the supply-chain risk?

Self-hosting gives your team control over one routing hop. Its packages, container images, plugins, cloud, upstream providers, and operational credentials remain dependencies. The LiteLLM incident shows why build provenance, pinning, secret isolation, and update response still matter.

Next action

Pick one production LLM request and draw its real path from client to serving endpoint. If the team cannot name every processor, freeze the provider list, or correlate a response with route evidence, keep that workload below Tier 2 until the contract is complete.

For adjacent controls, see Encrypted Reasoning Blocks Are Bearer Secrets and The AI Financial Control Plane Behind Model Routing.

References


Comment