Skip to content
Putting technology to work.
Insights to guide decisions and action.

Search articles

Does Input Data Persist? — Understanding Zero Data Retention in Generative AI

Table of contents · 6 items

Fields like this have increased on client security assessment questionnaires: "Whether generative AI is utilized in operations" and "Handling of input data." When discussing internally what to write here, have you ever found that answers varied from person to person?

"The free tier trains on data, but paid tiers are fine," "APIs don't retain data," "It should disappear after 30 days." All of these are partially correct and partially outdated. OpenAI's announcement expanding Zero Data Retention (ZDR) on August 19, 2026, provides a good opportunity to untangle this ambiguity. What matters more than the new feature itself is parsing and distinguishing the different kinds of commitments.

The three commitments are distinct from one another

Data handling in generative AI services breaks down into at least the following three categories. They are frequently conflated, but the scope of protection differs entirely.

CommitmentWhat it meansWhat happens if not met
Do not use for trainingSubmitted data is not routed to model training.The possibility of traces appearing in outputs to other companies in the future cannot be ruled out.
Humans do not viewPersonnel at the provider cannot review it.Content could be read under justifications such as abuse investigations.
Do not retainThe data itself does not remain after processing.Your company's data falls within the blast radius of outages or breaches on the provider's side.

The newly announced Zero Data Retention means that for eligible API customers, prompts and model completions are not retained after processing requests. It also indicates that content will not be accessible for review by OpenAI personnel, and that enterprise customer data will not be used for training unless explicitly opted in.

When all three are aligned, explaining it becomes simple; however, they are only aligned under specific usage arrangements that meet certain criteria. This brings us to the next point.

It does not happen by default

This is where misunderstandings occur most frequently. Zero Data Retention is an agreement applied upon application by eligible API customers; it is not something that takes effect automatically just by creating an account. Nor does it apply to general ChatGPT users.

In other words, in order to state internally that your company is protected, you must be able to state the following three things:

  1. Which route are you using? Are individuals opening ChatGPT on the web, calling the API from internal systems, or using AI features embedded inside SaaS products? The third is the most frequently overlooked. What powers the backend of a tool you use is usually stated only in its terms of service.
  2. Which agreement applies to that route? For arrangements requiring an application, they do not apply unless you have submitted one.
  3. Who verified it? Unless records show who verified it and on what date, you will find yourselves revisiting the exact same debate six months later.

Decision diagram showing how data handling diverges by usage route

We previously covered planning procurement without locking into a single AI provider in Recounting Dependencies in LLM Procurement. Data handling directly impacts the friction of switching vendors.

The direction of monitoring without viewing content

Another development likely to impact operations was announced simultaneously: the preview of Private Safety Processing.

When "do not retain" and "humans do not view" are enforced rigorously, detecting abuse becomes difficult for the provider. To resolve this dilemma, they outlined an approach that establishes a mechanism inspecting only cross-session interaction patterns without accessing individual contents. Coverage is scheduled to expand through September, accompanied by technical documentation detailing architecture and safety measures.

The key takeaway here is that it does not mean "nothing is observed whatsoever." Stating that "absolutely nothing remains" in internal briefings or client responses will force you to make corrections later. It is safer to frame it at the granularity that content is not retained, but usage facts and safety signals are processed.

Things the ordering party must decide

When embedding AI features into your company's systems, this issue cannot simply be left to development contractors. There are three things the client must decide before implementation.

The scope of data permitted for input. Customer personal information, contractual terms, unreleased business plans. How much can be passed to AI is a business decision, not a technical one. If this is left undefined, developers will lean toward either the most conservative or the most convenient extreme.

Whether to retain logs on your own side. The fact that the provider does not retain data also means records of what was sent will exist only on your company's side. If you need to trace later which input generated a particular response, you must design internal logging. Principles for maintaining credentials and audit trails were organized in Credentials Passed to AI Agents and Incident Response.

Where to state it in the contract. Unless explicitly stipulated as a requirement that Zero Data Retention must be applied, it depends entirely on the goodwill of the implementers. Including it as an acceptance testing item ensures verification requires only a single check.

What to do next

First, map out the routes through which AI is used within your company. Simply listing what tools are used across three categories—personal use, internal custom development, and SaaS integrations—is sufficient. Discussing data handling policies without this list will lead nowhere.

Next, prepare a single standardized response for business partners. This should be a concise statement detailing which of the three commitments applies to your company across each route. Having this prepared in advance versus researching it after being asked makes a tangible difference in the speed of closing business deals.

GleamHub offers consultations regarding custom development, AI, and automation to assist with integrating AI features into business systems, designing data handling boundaries, and implementing logging and auditing. Because system architecture varies depending on the types of data handled and existing infrastructure, please consult with us individually. Reach out via Contact Us.

Sources

Share this articleXFacebook
Kakeru Suzuki

Fascinated by the possibilities of technology, has had a deep interest in programming and digital art since student days

Turn this article's theme into your company's next step

Concrete steps forward for your organization.

We organize your desired architecture, legacy systems, and operational requirements to formulate your next steps toward execution.

  • Desired architecture
  • Integration with existing environments
  • Operational requirements
Consult on development & operations initiatives

You can consult with us from the initial conceptual stage. Details from this article will be carried over to the inquiry form.

Receive the latest articles by email