If you run regulatory affairs, quality, compliance, or engineering inside a company that handles controlled, confidential, or export-sensitive documents, you already know this. The harder problem is that the people selling you AI often do not. So the conversation starts in the wrong place. This is an attempt to start it in the right one: what the constraint actually is, why the usual workarounds do not dissolve it, and what you are left with when you take it seriously.

"We can't use the cloud" is usually a legal statement

There is a meaningful difference between would rather not and may not. A preference can be negotiated against convenience. A legal or contractual prohibition cannot. It sits upstream of the technology decision and constrains it.

For a large share of regulated work, the prohibition is the real situation. The data your team works with every day -- engineering specifications, batch records, complaint files, policy manuals, contract archives -- carries an obligation that travels with it. The obligation does not care how good a given AI product is. It cares about where the data goes and who can reach it.

The question is not "cloud or not." It is narrower and more useful: whose computers does your confidential data get processed on, and who can access it while it is there.

What "using cloud AI" actually does with your data

A cloud AI tool -- whether a chat assistant or an enterprise platform -- works by sending your text to a model running on infrastructure you do not own or control. To answer a question about your documents, the relevant content has to leave your environment and be processed somewhere else. Many products add retrieval: they index your documents so the model can pull in the right passages, which means a copy of that content lives in the provider's system too.

None of this is a flaw. It is simply how the architecture works, and for most companies it is fine. The problem is specific to organizations where "leaves your environment" is the exact thing a rule forbids. For them, the convenience of the cloud and the requirement of the regulation point in opposite directions, and the regulation wins.

Four constraints that decide it for you

These are not the only rules that apply, but they are clear, named, and common across the industries where this comes up. None of this is legal advice; confirm your own obligations with counsel. The point is to show that the constraint has a name -- and that it is the name, not a mood, doing the work.

Export-controlled technical data (ITAR)

If your organization handles defense-related technical data, ITAR governs where it can go. A 2020 State Department rule created an encryption carve-out: storing or transmitting unclassified technical data is not treated as an export if it is secured end-to-end with FIPS 140 cryptographic modules and the means of decryption are never given to a third party, as Wiley and Baker McKenzie both documented when the rule landed. The catch is in that last clause. Standard cloud providers hold the keys to your data, and their staff -- who may be foreign persons -- can reach it. The carve-out protects the encrypted pipe, not an arrangement where the provider can decrypt. That is why so many defense and aerospace teams conclude the ordinary cloud is off the table for this material.

Controlled Unclassified Information (CUI) and CMMC

Defense and government contractors that handle CUI are subject to NIST SP 800-171, and CMMC Level 2 requires full implementation of all 110 of those controls, per the Department of Defense's own CMMC model overview. Meeting that bar in a shared cloud environment -- and being able to prove you met it on the specific systems that touched the data -- is a heavy lift. For a lot of contractors, keeping CUI inside a controlled environment they can fully attest to is simpler and safer than stretching a cloud deployment to fit.

Personal data after Schrems II (GDPR)

If you hold personal data on people in the EU, GDPR Chapter V restricts moving it outside the European Economic Area without specific safeguards. The 2020 Schrems II decision invalidated the EU--US Privacy Shield, as Pinsent Masons and others have detailed, and pushed organizations toward data localization -- keeping data inside a jurisdiction rather than risking a transfer that cannot be defended. Routing that data through a US-operated cloud AI service is precisely the kind of transfer that became hard to justify.

Critical infrastructure (NERC CIP)

For electric utilities, NERC CIP effectively keeps medium- and high-impact systems off the cloud today -- not because the standard bans it outright, but because cloud providers cannot produce the per-device evidence auditors require, as industry analysis from Industrial Defender lays out. A standards effort to allow cloud-based systems is underway, but realistic timelines put usable rules years out, around the end of the decade. Until then, the practical answer for sensitive operational data is unchanged.

Other regimes hand you the same conclusion from a different direction: financial-sector confidentiality and supervisory expectations, GxP and data-integrity rules in pharma and chemical manufacturing, and ordinary contractual confidentiality clauses that prohibit disclosing a customer's information to third parties -- a cloud AI provider being a third party.

Why "private cloud" and bolt-on controls don't resolve it

Two answers usually come back at this point, and both deserve a straight response.

The first is "use the private or government version of the cloud." That can narrow the gap, and for some obligations it is enough. But it does not change the underlying fact that your data is processed on infrastructure operated by someone else, under their access model and their staff. For ITAR key control or for evidence you can fully attest to, "operated by a vendor" is frequently the sticking point, not the data center's certifications.

The second is "we'll add controls" -- data-loss prevention, redaction, tokenization. These reduce exposure, but they do not turn a prohibited transfer into a permitted one. The State Department, for instance, treated tokenization as distinct from encryption and did not extend the ITAR carve-out to it. Bolt-ons manage risk around a model that still sends your content out. They do not move the processing inside your walls, which is what the rule is actually about.

The options you're actually left with

Strip it down and there are three honest paths.

Do not use AI on the sensitive corpus. Safe, and increasingly expensive as the work piles up and competitors move. This is the default many teams are stuck in, not because they chose it but because no one offered them a third option.

Push to use cloud AI anyway and manage the exposure. Workable only where the obligation genuinely allows it. Where it does not, this is risk you are absorbing on behalf of the organization, and it tends to surface at the worst possible moment -- an audit, an incident, a contract review.

Run the AI inside your own environment. Put the model on infrastructure you control, keep the documents where they already live, and let the data stay inside your network boundary -- no egress to a third party. This is the path that actually matches the constraint instead of negotiating against it. It is more involved to stand up than signing up for a service, but it is the only one of the three that lets a regulated team use modern AI on its real documents without taking on a problem it cannot defend. It is also the approach we build at Metellus Partners: a sovereign deployment that runs entirely inside the client's own infrastructure.

When the answer is the third path, the system can still do what you want: answer questions from your own documents, show the source passage behind each answer so a person can verify it, log who asked what, and restrict who can see which material. The capability does not have to be surrendered to keep the data home. What changes is where it runs.

What to check before you let any AI near your documents

If you are evaluating any AI on a sensitive corpus, these are the questions that separate a serious option from a liability -- useful to hand to whoever on your side has to approve it:

  • Where is our data processed, physically and legally? Inside our boundary, or on a third party's infrastructure?
  • Who can access the content, and can we name them? Including the vendor's staff and their subprocessors.
  • Can we prove it? Logs, access records, and per-system evidence an auditor will accept.
  • Can a human verify each answer? Does the system cite the source document, or does it ask you to trust an unsourced response?
  • Does it match the specific obligation by name? ITAR, CUI/CMMC, GDPR, NERC, GxP, or your own contracts -- not a generic "we're secure."

If a tool cannot answer the first three cleanly, the rest does not matter for regulated work. The constraint already made the decision; the job is to find the approach that respects it rather than one that hopes to be forgiven.

"We can't put this in the cloud" is not the end of the AI conversation for a regulated organization. It is the start of a better one -- about building the capability where the data already has to stay.

Metellus Partners builds that capability as private, on-premise AI for regulated industries.

Sources

  1. U.S. State Department, ITAR encryption carve-out (2020) -- analysis by Wiley and Baker McKenzie.
  2. U.S. Department of Defense, CMMC Model Overview v2.0 (Dec 2021) -- CMMC Level 2 and the 110 NIST SP 800-171 controls.
  3. Pinsent Masons (Out-Law), International data transfers and Schrems II -- GDPR Chapter V and the 2020 CJEU decision.
  4. Industrial Defender, Does NERC CIP Allow Use of The Cloud? -- current limits on cloud for medium/high-impact BES Cyber Systems.