Technical overview • LLM Gateway • RAG • MCP • Observability

Zzeti Zeka Platform architecture: a model-agnostic LLM gateway, RAG, agents and observability — on your own infrastructure

This page summarizes, for technical teams, the layers of the Zzeti Zeka Platform developed by Zzeti, Skyloop's AI spin-off: which controls a request passes through after it enters from a channel, how a model is chosen, where data comes from, what is measured, and how the platform is installed on-premise, in a closed air-gapped network, in a private cloud or in a hybrid topology. Skyloop Cloud, the authorized reseller and deployment partner in Türkiye, installs and operates it as a fixed-scope project.

6
Layers, one platform
30+
LLM providers behind one gateway
50+
MCP connectors
0
External calls when air-gapped
Architecture

Six layers, one platform: the path from channel to model.

Each layer can be reviewed on its own; together they are what turns a local model into an enterprise system.

1

Channels & interfaces

The Zeka Chat console, WhatsApp, voice, Boards and the REST API with webhooks. Every channel passes through the same identity, access and guardrail layer; your own application calls the same agents through the API.

2

Agent layer

Expert and worker agents; skills, goals and personas; a visual workflow designer with conditional logic; the Actions Inbox for human approval (human-in-the-loop); a debug trace for every step. Agents bind to the gateway, not to a model — swapping the model does not change the agent.

3

Knowledge layer: Context Lake & RAG

The Context Lake ingests documents, tables and query results into collections; RAG retrieves the relevant chunks from those collections and hands them to the model as context. Query Bench runs data-warehouse and ERP queries through MCP, and the results become reusable context. Embedding models and the vector store run inside the installation too.

4

Model layer: the LLM gateway

The model-agnostic LLM gateway (the AI Gateway in Zzeti): local open-weight models served with vLLM, Ollama or Llama.cpp and — only if you allow them — 30+ cloud providers behind one endpoint. Model Hub and GPU Servers pull and run the models; Fine-tune Studio adapts them to your data with LoRA; load balancing, semantic routing and semantic caching are applied at the gateway.

5

Integration layer: MCP connectors

Agents reach ERP, CRM, data-warehouse and office tools through governed MCP (Model Context Protocol) connectors — Nebim V3, Logo, SAP, Microsoft 365, Google Workspace and custom systems; there is no direct database connection. The MCP Inspector and the Functions IDE let you test and publish new connectors and functions; the integration proxy decides which outbound connections exist at all.

6

Governance & observability

Security Hub and the Access Wizard (role-based access, IAM policies), guardrails and PII masking, quality evaluation with AI Judge, sessions, traces, gateway traffic and latency in Monitor, per-LLM budgets and rate limits, an audit record for every access — all on your own instance.

Request lifecycle

How the LLM gateway handles a request.

1

Identity & scope

The request arrives with a user and agent identity; the role hierarchy (for example Store → Region → HQ) decides which data, which connectors and which models may be reached.

2

Guardrails

Input policies, PII masking and secret detection run before the model; output policies run after the answer. The rules are independent of the model behind the gateway.

3

Routing

The gateway routes by policy: the local model by default; semantic routing by task; a semantic-cache hit never reaches a model at all; a cloud provider only when explicitly allowed and within cost and rate limits.

4

Inference & context

Chunks from RAG collections and the results of MCP tool calls are added to the context; inference runs on your GPU with vLLM or Ollama. Load balancing puts several model replicas behind one endpoint.

5

Trace & cost

Every step is traced — token count, latency, sources used, SQL executed, agent decisions; cost is attributed per team, agent and model; AI Judge scores quality on a sample. The records stay on your instance.

What model-agnostic means in practice: agents, workflows and guardrails do not depend on the model behind the gateway. When a new open-weight model is released, it is pulled from Model Hub, compared in the Playground against the same test cases and put into service with a single routing rule — no agent is rewritten.

Screens from the platform

Insight board, MCP connector and gateway topology.

Representative screens with sample data — click to open at full size.

Zzeti Insight Board: findings produced by the Cube Analyst agent from OLAP cubes, with a critical issue turned into an assigned action (representative screen, sample data)
Agents → insights → assigned action. The Cube Analyst agent scans the OLAP cubes, surfaces findings and turns the critical one into a task for the regional manager.
Zzeti Nebim V3 MCP connector: agents reach Nebim V3 data through standard MCP tools; every call is logged (representative screen, sample data)
MCP connector. Agents reach Nebim V3 (MS SQL Server) through named tools — query, list_tables, describe_table — over a connection pool in read-only mode; which agent called which tool is traced and logged.
Zzeti AI Gateway topology: agents, MCP servers and open-weight LLMs on local GPUs, with latency measured per call (representative screen, sample data)
Gateway topology in your own data center. Open-weight LLMs run on local GPUs; the AI Gateway is the hub between agents, MCP servers and models — every agent, model and MCP call is traced and its latency measured from one place.
Model observability & LLMOps

What is measured: every call, every agent step, every unit of cost.

Monitor, AI Judge and the audit log run on your own instance. Observability is what turns a pilot into an operated system — and what a KVKK review asks to see.

Traces
Session → agent step → model or MCP call, with a debug view for every step.
Tokens & latency
Per model, agent and connector; zero token cost on local models.
Cost
Budgets per LLM, monthly limits, cost per request, rate limiting.
Quality
AI Judge test cases with pass thresholds, user feedback, integrity checks.
Health
Agent fleet, success rate, live topology map, alerts.
Audit
Who asked what, which data was read, which action was taken — every record on your instance.
Models, serving & hardware

Which models, which engines, which hardware?

ComponentOptionsNote
Open-weight modelsQwen, Llama, Mistral, Gemma and other open-weight models on Hugging Face, including variants tuned for TurkishPulled by Model Hub; adapted to your data with LoRA in Fine-tune Studio
Inference enginesvLLM, Ollama, Llama.cppGPU Servers: local, remote or Kubernetes GPUs behind one endpoint
Cloud models (optional)OpenAI, Anthropic, Google, AWS Bedrock, Azure and 30+ providersOnly when allowed in the gateway, with PII masking in front; never in an air-gapped installation
Embeddings & vector storeEmbedding models and the vector store run inside the installationRAG collections live in the Context Lake
HardwareNVIDIA DGX Spark, GPU servers in your data center, a Kubernetes GPU pool, Apple Silicon as an edge nodeSkyloop sizes it in the discovery phase and can supply the DGX Spark
Platform runtimeKubernetes or Docker, installed with Terraform IaCSecurity teams review the whole installation as code
Deployment topologies

On-premise, air-gapped, private cloud or hybrid — the same platform.

The difference between the topologies is the network boundary and the maintenance process, not the platform. KVKK controls are on by default in all of them.

On-premise

The platform runs on Kubernetes or Docker in your data center; models run on your local GPUs; outbound connections are limited to what the integration proxy allows.

Air-gapped (closed network)

Zero external calls. Model files, updates and connectors are brought in through the maintenance process you approve; no internet is needed for licensing or telemetry.

Private cloud in Türkiye

When data must stay in the country but not necessarily in your building: the AWS Istanbul Local Zone or a private cloud tenancy in Türkiye — the same platform, the same controls.

Hybrid, local-first

Sensitive data stays on local models; non-sensitive tasks go through the gateway to the cloud providers you allow — with the same masking, access rules and cost limits.

Data flow step by step, the deployment-model comparison and the KVKK controls: on-premise local LLM guide →

FAQ

Questions technical teams ask about the architecture

What does 'model-agnostic' (LLM-agnostic) mean in Zzeti?

Agents, workflows, guardrails and evaluations are bound to the gateway, not to a specific model. Local open-weight models served with vLLM or Ollama and — if you allow them — cloud providers are interchangeable behind the same endpoint; a model is replaced with a routing rule, not by rewriting agents. The same test cases in the Playground and AI Judge show whether the new model is better before it goes live.

Is the LLM gateway the same thing as the AI Gateway?

Yes. The AI Gateway in Zzeti is an LLM gateway: a single API in front of local and cloud models that handles authentication, routing, load balancing, semantic caching, rate limits, budgets and logging. Applications and agents call the gateway; which model answers is a policy decision made inside your installation.

What does model observability cover?

Traces for every session, agent step and model or MCP call; token counts and latency per model, agent and connector; cost per team, agent and model with budgets and limits; quality scores from AI Judge test cases and user feedback; fleet health, success rates and alerts in Monitor; and an audit record of who asked what and which data was read. All of it is stored on your own instance — nothing is sent out.

Which components run in an air-gapped installation, and how are updates applied?

Everything: the gateway, local models, embeddings and the vector store, agents, MCP connectors inside your network, guardrails and the monitoring stack. Model files, platform updates and new connectors are packaged, checked and brought in through the maintenance process you approve — not over a live connection. Licensing and telemetry need no internet.

How do agents reach ERP data — what is MCP?

MCP (Model Context Protocol) is the open standard through which an agent calls tools and data sources in a governed way. In Zzeti, each ERP, CRM or data-warehouse system is exposed as an MCP connector with named functions and a schema — Nebim V3, Logo, SAP, Microsoft 365, Google Workspace or a custom system. The agent never opens a database connection; every call is scoped by role, logged with its trace, and write operations can require approval in the Actions Inbox.

Where do RAG documents live, and which sources can feed them?

In the Context Lake, as collections stored inside your installation together with the embeddings and the vector store. Collections can be fed from file shares and document systems, Microsoft 365 or Google Workspace through MCP, and from Query Bench results over ERP and data-warehouse queries. Access to a collection follows the same role hierarchy as everything else.

What is the relationship between Zzeti and Skyloop?

The Zzeti Zeka Platform is developed by Zzeti (Zzeti FZCO, Dubai), Skyloop's AI spin-off, and is an independent platform; it goes by 'Zzeti' globally and 'Zzeti Zeka' in Türkiye. Skyloop Cloud is the authorized reseller and deployment partner in Türkiye: it sizes the hardware, installs the platform on the customer's infrastructure, builds the integrations and operates it — as fixed-scope projects, not hourly consulting.

Is Kubernetes mandatory? Can our own team operate it?

No — the platform installs on Kubernetes or on Docker, with Terraform, so your security team can review the installation as code. After the fixed-scope deployment project, your own team can operate it; Skyloop provides updates, model upgrades and support inside your perimeter with the access you grant, for as long as you want.

See the architecture in your own environment.

Tell us which systems will connect, which data must never leave and how many users you expect. Skyloop adapts the reference architecture to your topology and comes back with a sizing and a fixed-scope deployment project — not an hourly consulting quote.