I ZDR you not
I wrote before about companies without frontier models pushing propaganda about the importance of owning your AI. The fears are mainly around the risk of AI labs using the enteprises' proprietary data and wiping out their business alpha.
So I took a look at the Terms of Service, Privacy Notices, and Data Processing Addendums (DPAs) to understand the actual legally binding promises the AI labs make to the enterprise customers of their APIs.
Let's first align on the commonly used terminology in the context of this post.
- Model training on customer data is the lab's use of customer content, including prompts, files, and model responses to train or improve its models.
- Data retention is the period the lab stores customer content after the model returns the response, often for the purpose of abuse monitoring.
- Zero data retention (ZDR) is an arrangement under which the lab does not store customer content at rest after serving a request.
- Data residency is the geographic location where customer content is processed during inference and where it is stored at rest.
Key findings from the labs' published commitments:
| Lab | Model training | Data retention | Zero data retention | Data residency |
|---|---|---|---|---|
| OpenAI | No | 30 days | Approval | Unspecified Choice: Approval |
| Anthropic | No | 30 days | Approval | US Choice: Self-service |
| No Free Gemini API: Yes | None Except safety violations | Approval | Unspecified Choice: Self-service | |
| xAI | No Free: Yes | 30 days | Self-service | Unspecified Choice: Self-service |
| Meta | No Discounted: Yes | Unspecified | No | Unspecified Choice: No |
| DeepSeek | Yes | Indefinite | No | China |
| Z.ai | No | None | Default | Singapore |
| Moonshot AI | No Personal accounts: Yes | Indefinite | Approval Personal accounts: No | Singapore |
Highlights:
- All five US labs rule out training their models on enterprise customer content sent through paid API plans. Google, xAI, and Meta do use data from free or discounted API tiers for model training.
- By default OpenAI, Anthropic, and xAI store customer content for 30 days, Google stores none unless its safety classifiers flag it, and Meta keeps it vague with "as needed". DeepSeek and Moonshot AI keep the content for as long as the account exists.
- OpenAI, Anthropic, and Google gate ZDR behind an approval process, while xAI lets a team admin just turn it on.
- The providers' ZDR implementations have limitations. For example, OpenAI and Anthropic exclude file uploads and batch jobs, and Google excludes Grounding with Google Search from ZDR.
- Pinning inference and storage to a specific location is available at all US labs except Meta, while the Chinese labs offer no such choice.
- The Chinese labs, except the straightforward DeepSeek, promise not to use enterprise data for model training.
My take:
- The US labs make binding promises not to train models on enterprise content. Obviously, it is also a question of trusting their statements, but I can hardly imagine Google, for example, risking a Cloud business that made $24.8B last quarter.
- The labs offer zero data retention (ZDR) to organizations that want to reduce data exposure. In exchange, a lab takes on more risk, because it gives up its ability to investigate AI misuse and abuse violations. Therefore, OpenAI, Anthropic, and Google grant ZDR only upon approval, while xAI makes it self-service.
- Anthropic, Google, and OpenAI are the most ready to serve customers in regulated industries. Google and Anthropic have a standard Business Associate Agreement (BAA) ready for organizations under HIPAA to accept. For finance, Anthropic is the most ready, with a 48-hour breach notice and customer audit rights, while Google and OpenAI leave the breach deadline to negotiation.
- Z.ai attracts enterprises with ZDR by default and is the only Chinese lab that offers a DPA. But its parent company, Beijing-based Zhipu, has been on the US Entity List since January 2025, and I recently wrote about its involvement in massive AI distillation campaigns.
- The DOJ bulk data rule (28 CFR Part 202) restricts giving China-owned vendors access to bulk US sensitive personal data. The low bulk thresholds, for example personal identifiers of 100,000 people or biometric identifiers of 1,000 people in any 12 months, mean that using Chinese models on their own platforms carries a high risk of breaking the rule.
Sources:
- OpenAI: Services Agreement
- OpenAI: Data controls in the OpenAI platform
- OpenAI: Data Processing Addendum
- Anthropic: Commercial Terms of Service
- Anthropic: Data Processing Addendum
- Anthropic: API and data retention
- Anthropic: How long do you store my organization's data?
- Anthropic: Data residency
- Google Cloud: Service Specific Terms
- Google Cloud: Cloud Data Processing Addendum
- Google Cloud: HIPAA Business Associate Addendum
- Google Cloud: Gemini Enterprise Agent Platform and zero data retention
- Google Cloud: Abuse monitoring
- Google Cloud: Data residency
- Google: Gemini API Additional Terms of Service
- xAI: Terms of Service - Enterprise
- xAI: Enterprise FAQ
- xAI: FAQ - Security
- xAI: Regional Endpoints
- Meta: Model API Terms of Service
- DeepSeek: Terms of Use
- DeepSeek: Open Platform Terms of Service
- DeepSeek: Privacy Policy
- Z.ai: Terms of Use
- Z.ai: Privacy Policy and Data Processing Addendum
- Moonshot AI: Terms of Service for Kimi OpenPlatform
- Moonshot AI: Privacy Policy
- Moonshot AI: What is Zero Data Retention (ZDR)?
- 28 CFR Part 202: Access to U.S. Sensitive Personal Data and Government-Related Data by Countries of Concern or Covered Persons
- Federal Register: Addition of Entities to and Revision of Entry on the Entity List
- CISA: China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies
- Alphabet: Second Quarter 2026 Results