Secure AI in an offline environment: how to build an air-gapped enterprise LLM architecture
The same sentence is heard in regulated sectors, at manufacturers bound by confidentiality agreements and at retailers holding customer data: “We want AI, but the data will not leave.” The technical answer to that sentence is not an API key but an architecture: the model, the context, the connectors and the logs all run inside the company's own network boundary — and, when required, that network has no connection to the internet at all.
This article walks through the components of an air-gapped enterprise LLM installation, how a question is answered inside, how updates are applied without internet and which evidence a KVKK review asks for. Where a product is named we use the Zzeti Zeka Platform as the reference; the architectural principles apply to any local LLM deployment.
Air-gapped, on-premise and private cloud: three different network boundaries
The three terms are often used interchangeably, but the network boundary differs in each. In an on-premise installation the platform runs on your hardware; the network may still allow the outbound connections you permit — to download a model, say, or to route a non-sensitive task to a cloud provider. In a private cloud inside Türkiye the hardware sits in a data center: data does not leave the country, but it is not in your building either. In an air-gapped installation there are no external connections at all: platform, models, context and logs live in a fully isolated network.
Which one is right depends on the data class. Most organizations hold three classes at once: general content that may leave, personal data that must stay in the country and trade secrets that must not leave the building. A good architecture treats these classes as layers working side by side rather than as rivals — and keeps the air-gapped mode permanently on for the most sensitive class.
- On-premise: your hardware, the outbound connections you allow.
- Private cloud (Türkiye): your tenancy, inside the country; data never crosses the border.
- Air-gapped: zero external connections; updates and model files arrive through a maintenance process.
- Decision criterion: which data class requires which boundary — a data inventory, not a technology preference.
Reference architecture: six components
A local model alone is not an enterprise system; an inference engine and a model file come up in a few days. What takes time and determines security is the set of layers around the model. In an air-gapped installation every one of those layers has to run inside the network — there is no piece that can be left on the cloud side and “added later”.
The six components below map one-to-one onto the layers of the Zzeti Zeka Platform; if you are building a different stack, the same list works as a checklist.

- Model serving: open-weight models (Qwen, Llama, Mistral, Gemma and variants tuned for Turkish) served with vLLM, Ollama or Llama.cpp; model files kept in a local model registry.
- LLM gateway: the layer through which agents and applications call models via one endpoint — load balancing, semantic routing, caching, rate and budget limits; in air-gapped mode the cloud-provider route is closed.
- RAG: embedding model, vector store and document collections — all inside. When a question arrives, the relevant chunks are retrieved from those collections and handed to the model as context.
- MCP connectors: governed tool calls into ERP, CRM and the data warehouse; no direct database connection. The integration proxy decides which outbound connections exist — in an air-gapped network, none.
- Guardrails and access control: PII masking, secret detection, input/output policies and a role hierarchy (Store → Region → HQ, for example) run before the model.
- Observability: traces, token and latency counts, cost, quality scores and the audit log — again inside. No telemetry leaves.
Data flow: how a question is answered inside
Nothing explains an architecture better than the path of a single request. A user asks a question from the console, over WhatsApp or from the company's own application; the request reaches the gateway with a user and agent identity. The role hierarchy decides which collections, which connectors and which model that identity may reach.
Then the guardrails run: personal data is masked, secrets are caught, policy rules are applied. If needed, the agent pulls data from the ERP through MCP and retrieves the relevant chunks from the RAG collections; inference runs on the local GPU. The answer passes the output policies, returns to the user, and every step — which source was used, which SQL ran, how many tokens were spent — is written to the audit log. No link in this chain crosses the network boundary.
- Identity → scope: user and agent role determine reachable data and models.
- Guardrails: PII masking and policy checks before the model.
- Context: MCP tool calls + RAG chunks.
- Inference: local model, local GPU.
- Record: trace, tokens, sources and actions in the audit log — inside the installation.
Updating without internet: model files, connectors, patches
The most frequent question about air-gapped installations is: “What happens when a new model comes out?” The answer is a defined maintenance process, not a live connection. Model files, platform updates and new connectors are packaged outside, verified with a checksum and brought in through a process the organization approves — physical media or a one-way transfer point. Inside, they are written to the local model registry and switched on with a routing rule in the gateway.
Designing this process well removes the fear that an air-gapped installation will “fall behind”. Comparing the new model against the same test cases before go-live (Playground and AI Judge) and keeping a rollback plan ready are just as possible inside as they are in the cloud.
- Packaging: model files + platform version + connectors, with a version number and checksum.
- Approval: the security team reviews the package; import follows the organization's change-management process.
- Verification: the new model is compared with the old one in the Playground on the same test cases.
- Go-live: a single routing rule; agents do not change; rollback is a rule change.
- No internet is needed for licensing or telemetry.
Hardware sizing: how many users, which model, how much memory?
Three questions drive sizing: which models will run, how many users and agents will generate requests at the same time, and how large the RAG collections are. The model's parameter count and quantization level determine GPU memory; concurrent requests determine the number of inference replicas and the load balancing. The embedding model and the vector store need memory too — and are frequently forgotten.
A single NVIDIA DGX Spark on the desk is enough to start and for many mid-size deployments; larger loads run on GPU servers in the data center or a Kubernetes GPU pool; Apple Silicon can serve as an edge node. The right approach is to measure during discovery, validate in the pilot and grow in steps — not to buy “the biggest server” on day one.
- Model choice: the smallest model that is good enough for the task; a bigger model does not always answer better, but it always costs more.
- Concurrency: expected requests at peak; load balancing and caching reduce it.
- RAG: collection size drives embedding and vector-store memory.
- Growth path: DGX Spark → GPU server → Kubernetes GPU pool; the platform stays the same.
Evidence for a KVKK review
Keeping data in Türkiye and under the organization's control removes the cross-border transfer question, but a review does not stop there. It wants to see that personal data is masked before it reaches the model, who accessed which data, which queries ran and that all of it is logged. The advantage of an air-gapped architecture is that all of this evidence is already being produced inside.
In practice, two documents complete the picture: a data inventory showing which data class sits in which collection and is reachable by which role, and an AI usage policy defining what agents may and may not do.
- PII masking records: which fields, under which rules.
- Access matrix: role → collection → connector → model.
- Audit log: session, trace, function call, executed SQL.
- Data inventory and AI usage policy: delivered with the installation.
An air-gapped enterprise LLM is not a model cut off from the internet; it is all six layers — model serving, gateway, RAG, MCP connectors, guardrails and observability — running inside the organization's network boundary. A well-designed maintenance process removes the fear of falling behind, and an architecture that produces its own evidence turns the KVKK review from a preparation project into a report.
Frequently asked questions
Which open-source models run in an air-gapped installation?
Any open-weight model on Hugging Face — Qwen, Llama, Mistral, Gemma and variants tuned for Turkish — can be served with vLLM, Ollama or Llama.cpp. Model files are brought in through the maintenance process and kept in the local model registry; several models can run behind one gateway with load balancing.
What is the difference between an air-gapped and an on-premise installation?
In an on-premise installation the platform runs on your hardware but may have the outbound connections you permit; in an air-gapped one there are none. The difference is in the network boundary and the maintenance process, not in the platform: same platform, same controls, different outbound policy.
Where are the documents for RAG stored?
In collections inside the installation, together with the embedding model and the vector store. File shares, document systems and ERP or data-warehouse queries through MCP feed the collections; access to a collection follows the role hierarchy.
Is an air-gapped installation more expensive?
Hardware and installation cost depend on the models and the number of users, not on the network boundary; air-gapping needs no extra license. What differs is the operating model: updates are imported in a maintenance window. Because local models carry no token cost, the total under heavy use usually ends up below an API bill.
The same architecture, installed as a fixed-scope project
Zzeti Zeka brings the gateway, RAG, MCP connectors, agents and observability ready-made; Skyloop installs it on your infrastructure — on-premise, air-gapped or private cloud — with the scope and the metric defined up front, not as hourly consulting.