AI decisions Architecture and technology
Sovereign or on-prem AI vs cloud AI APIs: where should your models run?
Decide by data, not by ideology. Cloud AI APIs are the default for most workloads because they give strong models with little to operate; sovereign or on-prem deployment is justified when the data, the regulator or the contract requires that processing stays under your control and EU jurisdiction. Most European organisations end up hybrid: sensitive workloads on infrastructure they control, the rest on cloud APIs under a governed gateway.
The options
Sovereign / on-prem AI
Models and retrieval run on your own data centre, a private cloud or a provider under EU jurisdiction that you control contractually and technically.
Cloud AI APIs
Models consumed as a managed service from a hyperscaler or model provider, possibly in EU regions, under that provider's terms.
Side by side
| Criterion | Sovereign / on-prem AI | Cloud AI APIs |
|---|---|---|
| Where data is processed | Inside a perimeter you define and can audit. | In the provider's infrastructure; EU regions are often available, depending on the service. |
| Jurisdiction | Can be kept fully under EU law if the operator is also EU-based. | A non-EU provider may be subject to third-country laws with extraterritorial reach, even with data stored in the EU. |
| GDPR international transfers | Transfers can be avoided by design. | Requires a valid transfer basis (adequacy, standard contractual clauses) where personal data may leave the EEA, and a transfer impact review. |
| Model choice | Mostly open-weight models you can host; frontier proprietary models are rarely available this way. | Access to the strongest proprietary models as soon as they are released. |
| Operational burden | Hardware or capacity, MLOps, security patching and monitoring are yours. | Largely managed by the provider. |
| Cost profile | Upfront and fixed; pays off with steady volume and long horizons. | Pay per use; flexible, but can be hard to predict at scale. |
| Auditability | Full logs and change control under your rules. | Depends on what the provider logs, exposes and contracts. |
| EU AI Act | Deployer obligations are the same; location does not reduce them. | Same deployer obligations; you also depend on the provider's documentation as GPAI or system provider. |
Choose Sovereign / on-prem AI when…
- You process special categories of personal data, classified information or core intellectual property that policy says cannot leave your control.
- You are a public body or critical operator with security frameworks that constrain where and by whom data may be processed (in Spain, for example, the Esquema Nacional de Seguridad).
- A sector regulator, auditor or contract requires full traceability of where inference runs and which model version answered.
- The workload is steady and high-volume, and a well-tuned open-weight model meets the quality bar.
Choose Cloud AI APIs when…
- The data is public, internal-low-sensitivity or can be anonymised before it leaves your perimeter.
- You need frontier capability that is only offered as a service.
- Demand is uncertain and you want to validate value before investing in infrastructure.
- Your team cannot yet operate models in production securely.
When to combine them
A hybrid pattern works for most organisations. Classify each workload by data sensitivity and regulatory impact; keep retrieval over sensitive documents and the most sensitive inference on infrastructure you control; anonymise or minimise data locally before any call to an external model; and send the rest to cloud APIs through a single gateway that enforces identity, logging and routing rules. The rule should live in the gateway, not in each team's code.
Common mistakes
- Confusing data residency with sovereignty: an EU region does not by itself remove third-country jurisdiction over the provider.
- Going on-prem for everything and ending up with weaker models, high fixed cost and slow delivery.
- Sending sensitive data to a public API during a pilot and discovering the legal blocker only before production.
- Assuming that running a model locally exempts you from the EU AI Act; deployer duties depend on the use case.
- Leaving security, legal and the data protection officer out of the architecture until the end.
How Thinkia approaches it
We start with a classification, not a platform. Each use case is placed by data sensitivity, regulatory exposure and required capability, and that decides where it runs. Security, legal and the data protection officer join the architecture from the first session, because they are the ones who will approve or block production.
Synapse is built for this hybrid reality. Its RAG repository runs on the client's own servers, so documents and embeddings do not leave the corporate perimeter; its gateway enforces corporate SSO and logs every call; and its routing prioritises confidential local models before commercial endpoints. Because it is LLM-agnostic, a workload can move between a hosted open-weight model and a cloud API without rewriting the application.
On regulation, we treat sovereignty and compliance as separate questions. GDPR governs where personal data goes and on what legal basis; the EU AI Act (Regulation (EU) 2024/1689) governs what the system does and who is accountable. A sovereign deployment can still be high-risk, and a cloud deployment can be compliant. For public bodies we also consider national frameworks and the supervisory authority, in Spain AESIA. This is practical orientation, not legal advice.
Thinkia products involved
- SynapseGoverned agentic platform: agents, models, costs and data in one place.
- Enterprise Knowledge AIGoverned knowledge layer: source-grounded answers with citations and audit trail.
- EU AI Act governance guideRisk tiers, timeline, roles and a 20-point checklist. Not legal advice.
Related AI solutions
- AI governance, risk & controlAI your board, legal team and regulators can sign off on.
- MLOps & AI infrastructureTurn experiments into production
- AI-ready data platformA platform AI use cases ship on
- AI patient & citizen assistantThe first response that's always right — in any language, at any hour.
- Enterprise Knowledge AIYour organisation knows more than it can find.
Frequently asked questions
Is hosting in an EU cloud region enough to be “sovereign”?
Not always. Residency means the data is stored and processed in the EU; sovereignty also covers who can be compelled to access it and under which law. If the provider is subject to non-EU legislation with extraterritorial reach, assess that risk explicitly, together with contractual and technical safeguards such as encryption with keys you control.
Can we use cloud AI APIs with personal data under GDPR?
Yes, if you have a legal basis, a data processing agreement, appropriate safeguards for any transfer outside the EEA and, where required, a data protection impact assessment. Transfer frameworks have been struck down before, so keep a fallback option for the most sensitive workloads.
Does on-prem AI mean worse models?
Often it means different models. The strongest proprietary models are rarely available to host yourself, but open-weight and smaller specialised models perform well on many bounded tasks. Test on your own cases before deciding the quality gap matters.
What does the EU AI Act change for this decision?
It does not dictate where models run. It sets obligations by risk level and role: as deployer you need human oversight, logging and relevant input data for high-risk uses, wherever the model is hosted. Annex III high-risk obligations apply from December 2027, after the Digital Omnibus on AI entered into force on 27 July 2026; check the consolidated text on EUR-Lex or the AI Act Service Desk. This is not legal advice.
Is sovereign AI mandatory for the public sector?
There is no single rule. It depends on the data, the security category of the system, national frameworks and procurement conditions. Many public bodies combine both: citizen-facing information services on cloud APIs and case files or sensitive records on controlled infrastructure.
Where should we start?
Inventory current and planned AI use cases, including unsanctioned ones, and classify them by data sensitivity. That map usually shows that only a minority of workloads need sovereign hosting, and those are the ones to design first.
Keep exploring
Related decisions
- Open-source vs proprietary LLMs: how should an enterprise choose?
- Single AI vendor vs multi-model strategy: should you bet on one provider?
- EU AI Act provider vs deployer: which role are you, and what does each one owe?
- Centralised vs federated AI governance: who should decide what in your organisation?
Sectors where this decision comes up
- Public Sector & Smart Cities
- Finance
- Insurance
- Healthcare & Life Sciences
- Energy & Utilities
- Manufacturing
Key terms
Thinkia articles
- Secure On-Premises AI: Why the Mainframe Is Making a Strategic Comeback
- On-Device AI: The New Standard for Enterprise Data Privacy
- Hybrid AI Strategy: Why Open-Source Models Are Now Essential
- Frontier Model Compliance: The EU AI Act's Hidden Liability
Whitepapers
- The AI Act already applies.The Digital Omnibus postponed Annex III to December 2027. It did not touch Article 4 or Article 50. Five questions for your next committee meeting.
- Shadow AI. Govern it, don't ban it.Three in four employees already use AI where IT cannot see it. Banning removes the witnesses, not the risk: make visible, govern, and enable.