Back to blog
Compliance·May 25, 2026·7 min read

GDPR compliance for LLMs in France — pseudonymization, DPA, and sovereignty

A complete checklist for deploying Claude / Gemini / GPT in GDPR compliance: PII pseudonymization code, whom to sign an Article 28 DPA with, Zero Data Retention directly through Anthropic and Google, the Mistral and AI sovereignty question, log hosting, and informing data subjects.

Deploy an LLM (Claude, Gemini, GPT) in a French company in compliance with the GDPR — the complete checklist: data pseudonymization, processing agreement (DPA), Zero Data Retention, log hosting, and the question of Mistral / sovereignty.

The regulatory minimum

  1. Pseudonymize PII before every external LLM call
  2. DPA Article 28 GDPR signed with each processor that processes the data. Kunavo does not offer one: if your processing requires it, contract directly with the model provider
  3. Up-to-date processing register, including the purpose of each LLM call
  4. Inform data subjects: your privacy policy must mention the use of AI

1) Pseudonymization — ready-to-use code

Before the prompt leaves your backend, we replace direct identifiers with tokens. The mapping table stays local:

pseudonymiser.py
# Pseudonymiser les données personnelles avant l'appel LLM
import re
from openai import OpenAI

PATTERNS = [
    (re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+"), "<EMAIL>"),
    (re.compile(r"\b(?:\+33|0)\s*[1-9](?:\s*\d{2}){4}\b"), "<PHONE_FR>"),
    # Numéro de sécurité sociale français (15 chiffres)
    (re.compile(r"\b[12]\d{2}(?:0[1-9]|1[012])(?:\d{2})(?:\d{3})(?:\d{3})(?:\d{2})\b"), "<NIR>"),
    # IBAN FR
    (re.compile(r"\bFR\d{2}\s?(?:\d{4}\s?){5}\d{3}\b"), "<IBAN_FR>"),
]

def pseudonymiser(texte: str) -> tuple[str, dict]:
    masque = texte
    mapping = {}
    for i, (pat, tag) in enumerate(PATTERNS):
        for m in pat.finditer(texte):
            cle = f"{tag}_{i}_{m.start()}"
            mapping[cle] = m.group(0)
            masque = masque.replace(m.group(0), f"[{cle}]")
    return masque, mapping

The LLM sees “[EMAIL_0_42] ordered [IBAN_FR_3_158]” instead of the real identifiers. If the response needs the values inserted again (for example, in a confirmation email), we perform the inverse substitution afterward, never in the API call.

2) DPA — who needs to sign what

  • You ↔ Kunavo: Kunavo does not offer a GDPR Article 28 DPA. Your requests are processed in the United States and transmitted to upstream providers, whose own privacy and data-retention policies apply to the content
  • You ↔ model provider: if your processing requires a DPA, sign it directly with Anthropic, OpenAI, or Google under their enterprise terms
  • Standard Contractual Clauses (SCCs): for transfers to the USA (Anthropic, OpenAI), verify that the SCCs approved by the European Commission are included in the DPA you sign with the provider

3) Zero Data Retention (ZDR): directly with the provider

By default, Anthropic and Google retain prompt logs for 30 days (Anthropic) / 60 days (Google) for abuse monitoring. For GDPR-sensitive workloads:

  • Kunavo: no ZDR. Requests pass through our gateway in the United States and then through upstream providers, whose own retention policies apply
  • Anthropic ZDR: offered by Anthropic to enterprise customers under a direct contract with Anthropic
  • Google Vertex AI EU: under a direct contract with Google Cloud, routing to europe-west1 (Belgium) or europe-west4 (Netherlands) is possible—data does not leave the EU. Kunavo offers no EU region
  • OpenAI: ZDR on the Enterprise tier, under a direct contract with OpenAI

The question of Mistral and sovereignty

Many French teams want to reduce their dependence on US providers. Three angles to consider:

  • Mistral through Kunavo: No. No Mistral model or other European provider is in the catalog, and there is no routing restricted to European providers
  • Mistral directly: For genuinely sensitive cases (public sector, healthcare), Mistral offers on-premises deployment. Kunavo does not replace this need
  • Combine the two: Mistral directly for sensitive cases, Claude through Kunavo for strong reasoning. These are two providers and two keys, and your code chooses where each call goes

Log hosting and data residency

  • Kunavo routing: In the gateway, in a single region (US-East, Ashburn). Only the anycast edge in front terminates TLS near you
  • Account and billing persistence: A single primary database in US-East, with daily encrypted volume snapshots. Neither EU-only routing nor EU-region storage is available, even on request
  • Your application logs: you must host them in the EU (OVH, Scaleway, Hetzner). Do not be the weak link

Informing data subjects — what to say

If you have personal data processed by an LLM, your privacy policy must mention this. Minimal template:

We use artificial intelligence models (Anthropic Claude, Google Gemini, OpenAI GPT) as subprocessors through Kunavo (https://kunavo.com) for the following purposes: [classification, summarization, response generation, etc.]. Personal data is pseudonymized before transmission. Kunavo does not use this data to train models; each model provider processes it according to its own privacy and retention policies. You have the right of access, rectification, objection, and erasure under Articles 15 to 21 of the GDPR.

What to check before production

  • DPA signed with each subprocessor (directly with the model provider: Kunavo does not offer one)
  • Processing register up to date
  • Pseudonymization tested on a real sample (leakage rate < 0.1%)
  • Logs in an EU region
  • Procedure for responding to Article 17 requests (right to erasure)
  • Privacy policy updated and accessible
  • DPO consulted if processing is large-scale or sensitive

Ready to get started? Sign up for free, then top up from $10.