"Generative AI proved so convenient when we started using it for work that employees now use it freely. But recently, a client asked us: 'Is your company entering our information into AI, and which country is that processed in?' We couldn't answer." This was a consultation from the owner of an SMB. When adopting generative AI, companies eagerly compare accuracy and pricing, yet often begin using tools without ever examining where inputted information goes or how it is treated.
This is not simply a matter of foreign AI being hazardous. The problem is that without internal boundaries, companies have not defined who may share what information with which AI. Regulatory discussions requiring AI to run within domestic borders are advancing worldwide, and mechanisms allowing businesses to choose where their AI operates are expanding. In this article, we outline a practical framework from a project owner's perspective for SMBs to determine what data can be shared with which AI.
Why processing location matters
When you enter text or files into generative AI, that information is typically transmitted to external servers—often located overseas—and processed there. There are two critical points to grasp here.
First, inputted information may be retained and, in some cases, used for model training. Depending on the service and subscription tier, some providers explicitly commit to not retaining input data or using it for training, whereas others do not. Policies frequently differ between free personal tiers and enterprise plans. Second, which country's laws govern your data. The physical country where servers reside determines applicable jurisdictions as well as protocols for legal disclosure and data retention during emergencies. If your client contracts stipulate that information must not leave the country, this directly translates into an operational violation. Ultimately, providing data to an AI means entrusting data to an external party, and sharing confidential details without knowing where they are hosted is much like storing documents in a safe whose keyholder is unknown. Options for keeping AI entirely within your company are covered in our article on internal private LLMs and our article on local AI environments.
Classifying data into three tiers instead of asking "is it safe or dangerous?"
Where many companies falter is framing the question as "is this AI safe or dangerous?" That question cannot be answered. Even with the same AI tool, problems may or may not arise depending entirely on the content of the data provided.
A practical approach is not to screen individual AI tools, but to categorize your company's data into three tiers and determine which level of AI is permissible for each.
| Data classification | Examples | Permissible AI |
|---|---|---|
| Publicly disclosable information | Public documents, published text, drafts for external release | No restrictions; general generative AI tools are fine |
| Internal-only information | Internal documents, meeting minutes, unreleased project plans | Restricted to enterprise AI with contractual terms prohibiting model training |
| Information that must never be entered | Customer personal data, client confidential information, data contractually barred from cross-border transfer | As a rule, do not enter into AI; use internal-only solutions if necessary |
The beauty of this classification is that employees are never left guessing. They can make immediate determinations: "This is internal-only, so it must stay within enterprise-tier AI," or "This is customer data, so it cannot be entered." By attaching criteria to the information itself rather than vetting every single AI tool, rules can be applied effectively in everyday operations. Governance encompassing advanced use cases where AI sends emails automatically is also addressed in our article on Gemini Spark governance.
Checking contract plans and "do not train" settings
Once tiers are established, the next step is to verify how the AI you use actually handles data. Assumptions are dangerous here; assuming that using an enterprise plan automatically guarantees safety can easily trip you up.
Focus on three primary points. First, whether settings explicitly exclude inputted data from model training. Second, whether you can select the country or region where data is processed (some enterprise offerings allow specifying data storage regions). Third, how long conversation histories and uploaded files are retained, and whether they can be deleted. While these can be verified in admin consoles and service terms, they are packed with technical jargon that can be difficult for project owners to interpret alone. Getting an external perspective on ambiguous areas is often the quickest path. Concrete methods for governing AI chat histories are covered in our article on conversation history governance.
Case study: A company shifting from a total ban to categorized adoption
Here is a specific example. Fearing data breaches, one business (company name withheld) instituted a total ban on generative AI for workplace tasks. In reality, employees were secretly using it on personal smartphones, creating an unmanaged environment where confidential information was actually being exposed outside company oversight. The ban was worsening the situation.
What they did instead was adopt the three-tier classification above, explicitly establishing that "public and internal information may be used with enterprise-contracted AI that opts out of training, while customer information must never be entered." Concurrently, they verified the enterprise AI's data storage region and training opt-out settings, distributing a single-page decision matrix to employees. As a result, the motivation for covert usage vanished, bringing AI use back under governance. What proved effective was not advanced technology, but condensing data boundaries into a single-page guideline that avoided both outright bans and unchecked free-for-alls.
Start by defining what data must never be entered
When adopting generative AI for business, attention often gravitates toward comparing capabilities and pricing, but where inputted data is processed and how it is treated directly affects client trust and contracts. The solution is not choosing a single safe AI tool, but classifying your data and determining which tier of AI is acceptable for each category. Even if you cannot establish everything at once, simply defining the information that must never be entered will prevent worst-case incidents.
If you are experiencing concerns like "employees are using AI on their own and management cannot track it," "we couldn't answer our client when questioned about data handling," or "we banned AI but suspect staff are using it discreetly," please reach out via GleamHub's development, AI, and automation consultation. From setting up data tiers and verifying AI configuration settings to crafting clear decision sheets and evaluating on-premise setups, we support you with solutions tailored to your operational realities.









