Back to blog
Compliance·May 25, 2026·8 min read

GDPR-compliant LLM deployment — PII masking, DPA, and Zero Data Retention

A practical guide for German companies that want to use Claude / Gemini / GPT in production in compliance with the GDPR. PII masking code, whom to sign a data processing agreement with, Zero Data Retention directly through Anthropic / Google (Enterprise), audit logs with request-ID correlation, and the Schrems II question.

GDPR-compliant LLM use in German companies — what you need before Claude or GPT can go into production. A practical setup covering PII masking, audit logs, Zero Data Retention with Anthropic / Google, and the DPA question.

The three required components

  1. Mask PII before the upstream call — You must not send unfiltered personal data in plain text to US providers
  2. Data Processing Agreement (DPA / Article 28 GDPR)—directly with the model provider (Anthropic, Google, OpenAI); Kunavo does not offer a DPA
  3. Audit trail — Every call is traceable and documented with its purpose

1) PII masking — the practical part

Before a prompt leaves your own backend, it passes through a masking step. Emails, phone numbers, personal names, IBANs, and social security numbers are replaced with tokens:

pii_mask.py
# Vor jedem LLM-Aufruf PII maskieren — niemals Klartext-Personendaten
# an die Upstream-API senden.
import re
from openai import OpenAI

PII_PATTERNS = [
    (re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+"), "<EMAIL>"),
    (re.compile(r"\b\d{4,}\b"), "<NUMBER>"),
    (re.compile(r"\b(?:\+49|0)\s*\d{2,4}\s*\d{4,8}\b"), "<PHONE_DE>"),
]

def redact(text: str) -> tuple[str, dict[str, str]]:
    redacted = text
    mapping: dict[str, str] = {}
    for i, (pat, tag) in enumerate(PII_PATTERNS):
        for match in pat.finditer(text):
            key = f"{tag}_{i}_{match.start()}"
            mapping[key] = match.group(0)
            redacted = redacted.replace(match.group(0), f"[{key}]")
    return redacted, mapping

def restore(text: str, mapping: dict[str, str]) -> str:
    for key, original in mapping.items():
        text = text.replace(f"[{key}]", original)
    return text

The mapping table stays local. If the LLM response needs to return the personal data that was inserted again (e.g., “Hello Ms. Schmidt”), we perform the substitution only after the response. The upstream always sees anonymized text.

2) DPA — who needs which agreement

  • You ↔ Kunavo: Kunavo does not offer an Article 28 GDPR DPA. Requests are processed in the United States and forwarded to upstream providers; what they do with the content is governed by their own privacy and retention terms (see /legal/privacy)
  • You ↔ model provider: if your workload requires a DPA, conclude it directly with Anthropic, OpenAI, or Google under their business or enterprise terms and call the model there directly. Through Kunavo, you do not inherit a subprocessor chain
  • Schrems II risk: Anthropic and OpenAI are based in the United States. Without a DPA, Kunavo also has no Standard Contractual Clauses (SCCs) for transfers to the United States. If your risk assessment requires a purely EU-based provider, Kunavo is not suitable: no provider in our catalog is based in the EU, and every request passes through our gateway in US-East

3) Audit log — log every call

Per request: request ID, user ID, model, hashed prompt, token usage, latency, request ID as a correlation key in your own logs:

audit.py
# Jeder LLM-Aufruf wird in einem unveränderlichen Audit-Log festgehalten
import json, time, uuid
from datetime import datetime

def call_with_audit(user_id: str, prompt: str, model: str) -> dict:
    request_id = str(uuid.uuid4())
    redacted_prompt, mapping = redact(prompt)

    # Audit-Log VOR dem Call schreiben
    audit_log.write({
        "request_id": request_id,
        "user_id": user_id,
        "timestamp": datetime.utcnow().isoformat(),
        "model": model,
        "prompt_redacted": redacted_prompt,
        "prompt_hash": hashlib.sha256(prompt.encode()).hexdigest(),
        "pii_tags_count": len(mapping),
    })

    resp = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": redacted_prompt}],
        extra_headers={
            "X-Request-ID": request_id,  # nur für deine eigenen Logs; Kunavo wertet ihn nicht aus
        },
    )

    response_text = resp.choices[0].message.content
    # Bei Bedarf PII zurück mergen (nur in der lokalen DB)
    final_text = restore(response_text, mapping)

    audit_log.write({
        "request_id": request_id,
        "response_redacted": response_text,
        "tokens": resp.usage.model_dump(),
        "duration_ms": int((time.time() - start) * 1000),
    })
    return {"text": final_text, "request_id": request_id}

The X-Request-ID is your own key: Kunavo does not evaluate or store the header, and our Usage Dashboard does not show which upstream provider handled a call. It shows Kunavo's own call ID for each call, with the timestamp, model, tokens, cost, and latency; reconcile it with your audit log using the timestamp and model.

Zero Data Retention (ZDR) — only directly with the model provider

By default, Anthropic and Google store prompt logs for 30 days (Anthropic) and 60 days (Google) for abuse monitoring. For GDPR-sensitive workloads, you can request ZDR directly from the provider under your own contract. Kunavo offers neither ZDR routing nor a ZDR tier: requests through Kunavo are processed in the United States and forwarded to upstream providers, whose own retention rules apply.

  • Anthropic ZDR: available to Enterprise customers under a direct contract with Anthropic. Not available through Kunavo
  • Google Vertex AI: similar terms under your own Google Cloud contract, plus the option to choose an EU region (europe-west4 / Netherlands). Kunavo does not route through an EU region
  • OpenAI: Zero Data Retention is available only for the Enterprise tier, under a direct contract with OpenAI. Not available through Kunavo

Data residency

  • Kunavo routing: In the gateway, in a single region (US-East, Ashburn). Only the upstream anycast edge terminates TLS near you
  • Account data persistence: A single primary database in US-East with daily encrypted volume snapshots. We do not offer EU-only routing
  • VAT: Checkout does not display VAT or issue an invoice; for a receipt for your accounting records, email contact@kunavo.com

What most compliance teams overlook

  • Prompt content is personal data if it contains names — even if the user didn’t enter “Müller,” but only said “my name is Petra.” Masking should heuristically capture first names and proper nouns
  • Training on your data is disabled by default at Anthropic and Google (unlike consumer ChatGPT). Kunavo itself does not train on your content; what the upstream provider does with it is governed by that provider's own terms. Kunavo provides no contractual assurance on this
  • Hash output content — store only the hash on your side, not the full text, if your DPA requires it
  • Right to erasure — in response to a GDPR Art. 17 request, you must be able to delete the prompt from your logs. This works only if your audit logs are deletable (i.e., not stored in append-only blockchain structures)

Imprint & cookie banner

For a complete compliance trail, see our /legal/impressum under § 5 DDG and our privacy policy at /legal/privacy. You manage the cookie banner on your own website (e.g. with Usercentrics, Cookiebot, or Klaro) — Kunavo itself sets only functional cookies.

Ready? Free registration, top up from $10, pay as you go, and credits never expire. If your workload requires a DPA, ZDR, or EU data residency, enter into the contract directly with the model provider.