I ZDR you not

I wrote before about companies without frontier models pushing propaganda about the importance of owning your AI. The fears are mainly around the risk of AI labs using the enteprises' proprietary data and wiping out their business alpha.

So I took a look at the Terms of Service, Privacy Notices, and Data Processing Addendums (DPAs) to understand the actual legally binding promises the AI labs make to the enterprise customers of their APIs.

Let's first align on the commonly used terminology in the context of this post.

  • Model training on customer data is the lab's use of customer content, including prompts, files, and model responses to train or improve its models.
  • Data retention is the period the lab stores customer content after the model returns the response, often for the purpose of abuse monitoring.
  • Zero data retention (ZDR) is an arrangement under which the lab does not store customer content at rest after serving a request.
  • Data residency is the geographic location where customer content is processed during inference and where it is stored at rest.

Key findings from the labs' published commitments:

LabModel trainingData retentionZero data retentionData residency
OpenAINo30 daysApprovalUnspecified
Choice: Approval
AnthropicNo30 daysApprovalUS
Choice: Self-service
GoogleNo
Free Gemini API: Yes
None
Except safety violations
ApprovalUnspecified
Choice: Self-service
xAINo
Free: Yes
30 daysSelf-serviceUnspecified
Choice: Self-service
MetaNo
Discounted: Yes
UnspecifiedNoUnspecified
Choice: No
DeepSeekYesIndefiniteNoChina
Z.aiNoNoneDefaultSingapore
Moonshot AINo
Personal accounts: Yes
IndefiniteApproval
Personal accounts: No
Singapore

Highlights:

  • All five US labs rule out training their models on enterprise customer content sent through paid API plans. Google, xAI, and Meta do use data from free or discounted API tiers for model training.
  • By default OpenAI, Anthropic, and xAI store customer content for 30 days, Google stores none unless its safety classifiers flag it, and Meta keeps it vague with "as needed". DeepSeek and Moonshot AI keep the content for as long as the account exists.
  • OpenAI, Anthropic, and Google gate ZDR behind an approval process, while xAI lets a team admin just turn it on.
  • The providers' ZDR implementations have limitations. For example, OpenAI and Anthropic exclude file uploads and batch jobs, and Google excludes Grounding with Google Search from ZDR.
  • Pinning inference and storage to a specific location is available at all US labs except Meta, while the Chinese labs offer no such choice.
  • The Chinese labs, except the straightforward DeepSeek, promise not to use enterprise data for model training.

My take:

  1. The US labs make binding promises not to train models on enterprise content. Obviously, it is also a question of trusting their statements, but I can hardly imagine Google, for example, risking a Cloud business that made $24.8B last quarter.
  2. The labs offer zero data retention (ZDR) to organizations that want to reduce data exposure. In exchange, a lab takes on more risk, because it gives up its ability to investigate AI misuse and abuse violations. Therefore, OpenAI, Anthropic, and Google grant ZDR only upon approval, while xAI makes it self-service.
  3. Anthropic, Google, and OpenAI are the most ready to serve customers in regulated industries. Google and Anthropic have a standard Business Associate Agreement (BAA) ready for organizations under HIPAA to accept. For finance, Anthropic is the most ready, with a 48-hour breach notice and customer audit rights, while Google and OpenAI leave the breach deadline to negotiation.
  4. Z.ai attracts enterprises with ZDR by default and is the only Chinese lab that offers a DPA. But its parent company, Beijing-based Zhipu, has been on the US Entity List since January 2025, and I recently wrote about its involvement in massive AI distillation campaigns.
  5. The DOJ bulk data rule (28 CFR Part 202) restricts giving China-owned vendors access to bulk US sensitive personal data. The low bulk thresholds, for example personal identifiers of 100,000 people or biometric identifiers of 1,000 people in any 12 months, mean that using Chinese models on their own platforms carries a high risk of breaking the rule.

Sources:

  1. OpenAI: Services Agreement
  2. OpenAI: Data controls in the OpenAI platform
  3. OpenAI: Data Processing Addendum
  4. Anthropic: Commercial Terms of Service
  5. Anthropic: Data Processing Addendum
  6. Anthropic: API and data retention
  7. Anthropic: How long do you store my organization's data?
  8. Anthropic: Data residency
  9. Google Cloud: Service Specific Terms
  10. Google Cloud: Cloud Data Processing Addendum
  11. Google Cloud: HIPAA Business Associate Addendum
  12. Google Cloud: Gemini Enterprise Agent Platform and zero data retention
  13. Google Cloud: Abuse monitoring
  14. Google Cloud: Data residency
  15. Google: Gemini API Additional Terms of Service
  16. xAI: Terms of Service - Enterprise
  17. xAI: Enterprise FAQ
  18. xAI: FAQ - Security
  19. xAI: Regional Endpoints
  20. Meta: Model API Terms of Service
  21. DeepSeek: Terms of Use
  22. DeepSeek: Open Platform Terms of Service
  23. DeepSeek: Privacy Policy
  24. Z.ai: Terms of Use
  25. Z.ai: Privacy Policy and Data Processing Addendum
  26. Moonshot AI: Terms of Service for Kimi OpenPlatform
  27. Moonshot AI: Privacy Policy
  28. Moonshot AI: What is Zero Data Retention (ZDR)?
  29. 28 CFR Part 202: Access to U.S. Sensitive Personal Data and Government-Related Data by Countries of Concern or Covered Persons
  30. Federal Register: Addition of Entities to and Revision of Entry on the Entity List
  31. CISA: China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies
  32. Alphabet: Second Quarter 2026 Results