← Back to Blog

Residency Isn't Sovereignty: What 'EU Data Residency' Actually Buys Your LLM Traffic

Flipping on an 'EU data residency' toggle checks a compliance box, but residency (where the bytes sit) and sovereignty (whose law can compel them) are different guarantees. As of mid-2026, the gap between them is where the real risk lives.

August 2, 2026


There is a moment that happens in a lot of compliance reviews right now. Someone asks, “Where does our LLM traffic go?” An engineer flips on the “EU data residency” setting in their provider dashboard, points at the region label, and everyone moves on. Box checked. Data stays in Europe. We’re covered.

Except that setting quietly answers a narrower question than the one that was asked. It tells you where your prompts and completions are stored. It does not tell you whose legal system can reach in and compel them. Those are two different guarantees, and as of August 2026 almost every LLM provider sells you the first while letting you assume you also bought the second.

This piece is about the gap between them. It matters more this year than last, because enforcement is arriving, a court order just demonstrated the gap in public, and the honest answer for teams that genuinely cannot let certain data leave their jurisdiction is more complicated than any dashboard toggle admits.

Two words that get used as if they mean the same thing

Start with the vocabulary, because the whole confusion lives here.

Data residency is about location. It is the promise that your data physically sits (and ideally gets processed) in a particular geography. “Your prompts stay in the EU” is a residency claim.

Data sovereignty is about jurisdiction. It is the question of whose laws govern that data and whose courts can compel access to it, regardless of where the bytes physically live. “No government outside the EU can force disclosure of our prompts” is a sovereignty claim.

Here is why the distinction is not academic. Most of the big LLM providers are US-headquartered companies. Under the US CLOUD Act, a US company can be compelled by a US court to hand over data it controls no matter where in the world that data is stored. On top of that, FISA Section 702 authorizes bulk collection of non-US persons’ data held by US companies, with no adversarial process and no notice to the customer whose data was collected. The legal test is only that the target be reasonably believed to be a non-US person located outside the United States.

Put those together and the uncomfortable conclusion is this: a US-headquartered provider storing your data in Frankfurt has given you residency, but the data is still reachable through the US company that controls it. The address changed. The jurisdiction did not. Analysts have written about this at length, and the phrase that captures it best comes from a piece titled “Two Sovereign Clouds, One Legal Wall”: the legal reach follows the corporate parent, not the server rack. Firms like SoftwareSeni and Kiteworks have laid out the same exposure in detail.

I want to be careful here, because this is exactly the kind of claim that gets overstated. Nobody has litigated a CLOUD Act warrant against, say, the new AWS European Sovereign Cloud and its German operating entity to see what actually happens. It is a live legal question, not a settled one. But “we don’t know if the wall holds” is a very different risk posture from “there is no wall,” and the residency toggle invites you to believe the second.

The fine print inside “residency” itself

Even if you set sovereignty aside and only care about residency, the toggle promises less than it looks like it does. Three traps show up repeatedly once you read the actual documentation.

Storage at rest is not the same as in-region processing. When OpenAI introduced EU data residency, the default guarantee (as reported in a June 2026 Wavect analysis, since OpenAI’s own page blocks direct fetching) is about where data is stored at rest. Keeping the actual inference (the model doing the computing) inside the region is a separate, narrower commitment that you often have to confirm you’re getting. Azure OpenAI has arguably the strongest EU posture through Microsoft’s EU Data Boundary, which covers data at rest and in transit, but even there, per the same analysis, batch jobs default to global processing unless you specifically pick the EU Data Zone variant. Google’s Vertex AI offers regional endpoints too, though its EU regions tend to lag the US by a quarter or two on the newest model variants, according to a 2026 architecture writeup from digitalapplied. The point is that “residency” is not one switch. It is several, and some of them are off by default.

“Zero data retention” is usually gated, not given. Zero data retention, or ZDR, means the provider doesn’t store your requests and responses after serving them. It is the thing you actually want for sensitive traffic. But ZDR is frequently approval-gated or sales-gated rather than a self-serve checkbox, and it is generally not available on consumer or lower-tier plans at all. So the team that turned on “residency” in a self-serve dashboard has very likely not turned on ZDR, because you usually cannot.

“We don’t train on your data” answers a question you didn’t ask. This is the big one, because a real event in the last year proved it. In May 2025, in the copyright case New York Times v. OpenAI, a court issued a preservation order that forced OpenAI to retain output logs it would otherwise have deleted, overriding users’ own deletion settings and privacy preferences. Only customers on ZDR API access and certain Enterprise and Edu tiers were exempt. The order was later narrowed and then lifted in October 2025 (covered by Engadget and analyzed by Terms.Law in November 2025), but the precedent is the lesson. A provider’s retention policy is a promise the provider makes. A court can override that promise. “We don’t train on it” and “we don’t keep it” and “no one can compel it” are three separate statements, and buying the first tells you nothing about the third.

Why this is heating up right now

Part of what makes August 2026 a real inflection point is regulatory. The EU AI Act’s rules for general-purpose AI models (the big foundation models) started applying in August 2025, and the Commission’s supervision and enforcement powers switch on on August 2, 2026, with models released before the 2025 cutoff getting until August 2027 to comply. So the machinery that turns “you should” into “you must, or else” is coming online essentially now.

And the “or else” is not hypothetical. European regulators have been active: Italy’s data protection authority fined OpenAI €15 million, and separately the Dutch DPA fined Uber €290 million specifically over improper transfers of data to the US, per a 2026 roundup from crescendo.ai (which also notes GDPR fines have now totaled billions cumulatively). The legal ground the US-transfer story stands on is also shakier than it looks. The EU-US Data Privacy Framework, which is what currently makes many transfers lawful, was upheld by the EU General Court in September 2025, but privacy advisers are still keeping Standard Contractual Clauses in place as a backup because a “Schrems III” challenge could unwind it again, as secureprivacy noted in 2026.

This is not only an EU story either. China’s Cybersecurity Law amendments took effect on January 1, 2026, with higher penalties and expanded extraterritorial reach. India finalized its DPDP rules in late 2025 with a “negative list” approach to cross-border transfers, where data can flow anywhere the government hasn’t specifically restricted, with full compliance expected by May 2027 (and, as of mid-2026, no restricted-country list published yet). If you route LLM traffic for users in multiple countries, the map of what’s allowed is getting redrawn in several places at once.

The actual options, ranked by how much control they give you

So what can a team that genuinely cannot send certain payloads to a third-party API actually do? The honest answer is a spectrum, not a product. Roughly, from most convenient to most controlled:

Regional endpoints from the big providers (OpenAI EU, Azure’s EU Data Boundary, Vertex AI EU, AWS Bedrock EU regions). These buy you real residency and a strong contractual posture. For a lot of GDPR and data-residency requirements, that is genuinely enough. What they do not buy you, as covered above, is full escape from US legal reach. Two provider-specific notes worth knowing: Anthropic’s own first-party API does not offer a guaranteed EU-only processing region, so the practical way to run Claude inside the EU is through AWS Bedrock or Vertex AI in EU regions where the model runs on the cloud provider’s infrastructure (Requesty documents this). And the newer “Claude Platform on AWS,” launched in May 2026, processes customer data outside the AWS security boundary, which is actually a step backward for strict-residency teams versus plain Bedrock (InfoQ and isimplifyme covered the distinction). The lesson repeats: read which specific product you’re actually on.

Sovereign cloud offerings. AWS European Sovereign Cloud went generally available on January 14, 2026, operated by a dedicated German legal entity with EU-resident staff and physical separation from the main AWS regions (MassiveGRID has a good summary). Microsoft has rolled out expanded sovereign cloud tiers too. These get you meaningfully closer, because the legal-structure separation is real. Two caveats keep them from being a clean answer: the US-parent question is exactly the untested one from earlier, and the frontier models you might want may not be available in these regions yet.

EU-headquartered providers and self-hosted open-weight models. This is the only path that keeps both residency and sovereignty fully in your hands. An EU-headquartered provider like Mistral sits under EU jurisdiction rather than US. And running open-weight models yourself (families like Llama, Qwen, DeepSeek, Mistral, GLM, and Kimi) means the data never leaves infrastructure you control. The good news is that the quality gap has narrowed a lot: several 2026 comparisons (wavect, datavlab) show open-weight models trading blows with frontier proprietary ones on benchmarks. I’d treat the specific scores as directional rather than gospel, because they come from vendor and blog roundups rather than primary leaderboards, and benchmark parity is not the same as production parity. Running your own fleet means owning the ops burden, the tool-calling reliability testing, and the long-context stability work yourself. The control is real. So is the cost.

One aside for the healthcare crowd, since it comes up constantly: there is no such thing as a “HIPAA-compliant model.” Compliance is a property of the whole deployment, not the model. It requires a signed Business Associate Agreement, zero data retention, encryption, access controls and audit, data minimization, and breach notification, all together. Anthropic, OpenAI Enterprise, Azure OpenAI, Vertex, and Bedrock will all sign BAAs (swfte walks through the requirements), with ZDR being default-on for Anthropic’s paid contracts and available on the others by arrangement. Consumer tiers do not offer BAAs, which means putting real patient data through a consumer ChatGPT or Claude subscription is a violation, full stop.

The part nobody’s dashboard handles: it’s per-payload

Look at that spectrum again and the real problem comes into focus. It is not “pick the sovereign option.” Most teams cannot, because the fully sovereign options are the most expensive and the least capable, and most of your traffic does not need them.

The thing is, sensitivity is not uniform across your requests. A prompt asking a model to refactor a public code snippet and a prompt containing a folder of patient records do not need the same guarantee. Forcing every request down the strictest, most sovereign, most expensive path because a few of them require it is how you end up with an AI capability that is too slow and too costly to use. Letting everything take the cheap, convenient path and hoping the sensitive stuff doesn’t slip through is how you end up in front of a regulator.

So the actual engineering problem is routing. Send the payloads that can leave to the cheapest capable model. Keep the payloads that cannot on infrastructure whose jurisdiction you trust. Doing that by hand, per call, across providers that each speak a slightly different API protocol, is exactly the kind of friction that pushes teams toward one of the two bad extremes.

Where a gateway fits

This is the seam that a gateway sits in, and it is worth being precise about what it can and cannot do.

A gateway like AJNT is a single wire-compatible endpoint that sits between your agents or applications and every model provider they call. Because it is one control point that all your LLM traffic flows through, the residency decision stops being a code change scattered across services and becomes a routing decision made in one place. You can route a request down a cost-and-preference-ordered ladder that ends wherever you need it to end: a mainstream provider for the non-sensitive majority, or a sovereign region, an EU-headquartered provider, a bring-your-own endpoint, or self-hosted hardware you operate for the traffic that cannot leave. If a preferred rung degrades, it falls through to the next automatically. For platform leaders, the same choke point is where team rules get enforced (rules set once by an admin ride every session), and where per-project attribution of what ran where actually becomes possible.

The honest limit: a gateway does not repeal the CLOUD Act. Routing to a sovereign region or a self-hosted box is still subject to whatever the underlying jurisdiction actually offers, and the untested legal questions stay untested. What the gateway changes is the cost of acting on the distinction. It makes the residency-versus-sovereignty tradeoff incremental and reversible. You can try the sovereign or self-hosted rung for a slice of traffic, keep the endpoint under your control, and fall back safely if it doesn’t hold up, all without rewriting your agent loop. When the decision is that cheap to change, you stop having to make one all-or-nothing bet for your whole system.

The one question worth asking

If there is a single thing to take away, it is a swap in the question you ask during that compliance review. Stop asking “is our LLM traffic in the EU?” That is a residency question, and the dashboard toggle will happily answer yes while leaving the harder thing unaddressed.

Ask instead: “Whose court can compel this specific payload, and is that acceptable given what’s in it?” That question forces you to notice that different payloads have different answers, that the label on the region is not the whole story, and that the honest posture is a routing policy rather than a checkbox.

As of August 2026, several of the load-bearing facts here are genuinely unsettled. We don’t know if a sovereign cloud’s legal-entity structure survives a CLOUD Act warrant, because no one has tested it. We don’t know if the EU-US data transfer framework survives another legal challenge. We don’t know whether open-weight parity holds up under real production load or just on benchmarks. Anyone who tells you those are solved is selling you a toggle. The useful move is not to wait for certainty. It is to build so that when the answers change, and they will, changing your routing is a config decision and not a rebuild.


← Back to Blog