Three nested outlines around a lit core, one dotted line running out to a small node, on a petrol-to-mint gradient
Oct 2, 2026

OpenAI Private Intelligence: what it promises, and what you can verify today

OpenAI now says AI inference belongs in confidential computing. The announcement is a preview with open questions. Here is what it contains, what it leaves out, and how to check any provider's claim, including ours.

Johannes Hötter

Johannes Hötter

VP Growth

On September 29, 2026, at DevDay, OpenAI introduced Private Intelligence, described as a way for businesses to use its models "with stronger data controls". It has two parts. The first, zero data retention with private safety processing, is rolling out now. The second, Private Inference, is a preview "coming this fall" that "combines confidential computing with strict, verifiable controls".

This matters to us at Edgeless Systems for the same reason Apple's move of Private Cloud Compute to third-party hardware did in June. The company behind ChatGPT is now essentially saying that processing sensitive data in plaintext isn't always good enough, and that the answer is technical operator exclusion through hardware-based confidential computing. That is the idea Privatemode has been built on since 2024.

What OpenAI announced

The announcement, as reported by VentureBeat and Decrypt, covers two things.

Zero data retention with private safety processing, available now for approved API projects. OpenAI already offered zero data retention, which excludes customer content from its logs. What is new is how abuse monitoring works under that setting: automated safety reviews run inside what OpenAI calls a hardware-attested runtime, so that OpenAI personnel do not see the content. Encrypted safety records are written to storage the customer controls, on AWS, Azure, or Google Cloud, with a 30-day time to live. Zero retention on OpenAI's side therefore does not mean zero storage; it means the copies live in your bucket, not theirs.

Private Inference, a preview announced for this fall. Inference is the moment the model reads your prompt and generates an answer, and today that happens in plaintext inside OpenAI's infrastructure. Private Inference is meant to move it into confidential computing: hardware that keeps memory encrypted and isolated while the model runs, so that even the cloud operator cannot read it, combined with what OpenAI calls verifiable controls. You can register interest; there is no product to evaluate yet.

What the announcement does not say

A preview is allowed to be incomplete, and OpenAI may well answer these points when Private Inference ships. But if you are deciding how to handle sensitive data in the next six months, these are the questions that are open today:

  • Which hardware, and which trusted execution environment. Confidential computing on CPUs (AMD SEV-SNP, Intel TDX) and on GPUs (Nvidia's confidential-computing mode) are different mechanisms with different guarantees. The announcement names neither.
  • Which models and endpoints. Private Inference will almost certainly serve OpenAI's own proprietary models; that is the point of the offering. Which ones, on which endpoints, and whether the full line-up is covered is not stated.
  • Who verifies. "Verifiable controls" can mean two very different things: the provider attests its own deployment and publishes a statement, or the customer's own client receives the attestation evidence and checks it against published reference values before sending data. The first is a promise with a hardware signature; the second removes the need to trust the provider. The announcement does not say which one it is.
  • What is published. Whether the code running inside the enclave is public, whether its builds are reproducible, and whether anyone outside OpenAI can confirm that the attested software is the published software.
  • Where it runs, and what it costs. No regions, no data-residency statement for the inference itself, no pricing.

None of this is a criticism of the direction. It is a description of where the preview stands, and a reminder that "confidential computing" on a slide is not yet a property of a service you can use.

Attestation is not verification

The word that will decide how much Private Intelligence is worth is "verifiable", so it pays to be precise about it. Here is the everyday version first.

You have sensitive documents and want to hand them to a processing room you are not allowed to enter. The room's manufacturer put a seal on it when it was built. The seal records exactly how the room was built and which machine is inside, and it cannot be forged. Before you hand your documents through the slot, you want to know that the seal is right. There are three ways to find out:

  1. Weak. The room's operator has its own inspector look at the seal, inside the building, and the courier tells you: "We checked, it's fine." You never see the seal. You are trusting the operator about the operator's own room.
  2. Strong. The courier shows you the seal at your door. You send a photo of it to an independent inspector you chose, and they tell you whether it matches. You are trusting that inspector, not the operator.
  3. Very strong. The courier shows you the seal at your door, and you hold it against the manufacturer's published picture of the correct seal yourself. You compare, you decide, and only then do you hand the documents over. You are trusting the seal, a picture anyone can look up, and your own eyes.

That third version is what the BSI's C5:2026 calls Type 1, and it is how Privatemode works: the room is the confidential VM, the seal is the signed attestation report from the CPU and GPU, the published picture is the manifest of reference values, and the person at the door is the client-side proxy running on your machine. It compares before your first byte leaves.

The weak and the very strong version side by side. Left: the operator's own inspector looks at the seal and tells you "it's fine". Right: the seal comes to your door and you compare it with the manufacturer's published picture.

In technical terms: attestation is something the hardware does: a CPU or GPU produces a signed report of exactly which software booted inside the confidential environment. Verification is something a party does with that report: checking the signature against the chip vendor's certificate chain and comparing the measurements against the values the software is supposed to have. A provider can attest its own deployment all day; the question is who performs the verification, and whether the reference values they compare against were derived from code anyone can inspect.

If the provider verifies on your behalf, you have moved your trust from the provider's privacy policy to the provider's attestation service. That is progress, but the provider is still the party you have to believe. If your own client verifies, using reference values you or an auditor can regenerate from public source, the provider has removed itself from the list of parties that could lie to you. Apple's Private Cloud Compute sits in between: it publishes binaries and a transparency log so that researchers can verify on everyone's behalf.

This is not a distinction we invented. Germany's Federal Office for Information Security wrote it into C5:2026, the current edition of its cloud security catalogue. Criterion OPS-33 makes remote attestation mandatory wherever a cloud offers confidential computing, and it grades attestation by who collects the evidence and who verifies it. Type 1, "very strong": the customer collects the evidence from the trusted execution environment and verifies it on their own systems. Type 4, "weak": the provider collects the evidence, verifies it with a service under its own control, and hands the customer a result. Most confidential-computing offerings from hyperscalers operate as Type 4 today. Privatemode is Type 1: the client-side proxy collects the evidence and verifies it on your systems, before any data is sent, which is why it already meets a criterion that becomes binding for German cloud providers in June 2027. "Verifiable controls" can mean anything on that scale. Which type Private Inference will be is the single most important open question in the announcement.

What you can run today

Privatemode is an inference API built by Edgeless Systems, a German company based in Bochum, and designed around the second model: the customer verifies. It serves open-weight models on purpose, not as a stopgap until something better comes along. Here is how it maps onto the questions above, so you can hold us to the same standard.

  • Hardware. Workloads run in confidential VMs on AMD SEV-SNP and Intel TDX CPUs, with Nvidia H100 and B200 GPUs in confidential-computing mode. Memory stays encrypted while the model runs. Both the CPU and the GPU produce attestation reports.
  • Who verifies. A client-side proxy, running on your machine or in your infrastructure, requests the attestation evidence and verifies it before a single byte of your prompt leaves your environment. The reference values it checks against come from a published manifest. Neither Edgeless Systems nor Scaleway, the infrastructure provider, can forge the report: it is signed with a key fused into the hardware and chains to AMD's, Intel's, or Nvidia's certificate authorities. In the terms of C5:2026, this is Type 1, the "very strong" form of remote attestation.
  • What is published. The source code of every trusted component is public and builds reproducibly with Nix. Anyone can rebuild it, regenerate the manifest, and confirm that the software admitted into the confidential environment is exactly the software that was published. The verification does not depend on us.
  • Where it runs, and who runs it. On Scaleway, in Europe, operated by a German company under European law. No US parent, no data leaving the EU for inference. Prompts and responses are encrypted end to end; by design, neither we nor the infrastructure provider can access them.
  • Models. Open-weight models only, including OpenAI's own gpt-oss-120b, Gemma 3, Qwen3, GLM-5.3-Flash, and Whisper, served through an OpenAI-compatible API. Open weights mean no lock-in to a single model vendor: the same model can be run elsewhere if you ever need to. Switching an existing integration is a base-URL change.

And the honest limit: Privatemode does not serve OpenAI's proprietary frontier models. If your use case needs GPT-5 specifically, Private Inference is the thing to wait for, and we hope it arrives with client-side verification and public code for the serving stack. If an open-weight model does the job, the property OpenAI is previewing is one you can have, and check, this afternoon.

Three questions for any Confidential AI provider

Confidential computing is becoming the baseline for AI on sensitive data. Apple runs it on Google Cloud, OpenAI is building it, and more will follow. As it becomes standard, the mechanism itself stops being the differentiator. What remains is how much of the claim you can check. Whoever you talk to, these three questions separate a property from a promise:

  1. Who verifies the attestation, and where? You, in your environment, before data is sent, or the provider, on its own infrastructure? In C5:2026 terms: Type 1 or Type 4?
  2. Can I reproduce what runs? Is the code inside the confidential environment public, and do the builds reproduce so that the measurements can be independently derived?
  3. Who stays in the trusted base? After the hardware isolation, which parties could still read your data: the operator, the cloud, the model vendor, a safety-review pipeline?

Ask them of OpenAI when Private Inference ships. Ask them of us now: the security page and the documentation are written to answer exactly these, and the API has a free tier so you can run the verification yourself.

Articles

Further reading

Explore other articles

Pale strands entering from the left at scattered heights and settling into one narrow band, one of them lime, on a dark petrol-to-green gradient

Oct 6, 2026

When 90% means 90%: calibrating one-token LLM decisions

Raw one-token probabilities are overconfident. One temperature fixes most of it without labels. About 100 labels also make more answers correct and put an error bound on the answers you automate.

A plume of pale filaments rising from a single point, one of them mint and reaching furthest, on a teal-to-steel gradient

Sep 24, 2026

Turn GLM-5.3-Flash into a Jev-like System One model

Typed decisions in a single forward pass, about as accurate as Jev and faster from Europe. With a playground and a benchmark on 29 datasets.