Technical series · Article 3 of 3

LLM gateway and model observability: no single-model dependency, every call measured

·8 min read·Skyloop Cloud

The most expensive decision in an enterprise AI project is the one taken without noticing: binding agents, prompts and integrations tightly to a single model and a single provider. When the model's price changes, a better open-weight model appears or a data-residency rule tightens, that binding means rewriting everything.

Two concepts dissolve the dependency. An LLM gateway is a single layer between agents and models: which model answers becomes a policy decision. Model observability measures the trace, cost and quality of every call; data, not guesswork, says what works. This article explains both, with their counterparts in the Zzeti Zeka Platform.

The cost of depending on a single model

Dependency shows up in three places. Technical: prompts and tool calls are tuned to one model's behaviour; change the model and the agent breaks. Commercial: price, quota and terms of service are set by one provider. Compliance: where the data goes depends on the provider's infrastructure, which for KVKK means a fresh assessment at every change.

A model-agnostic architecture loosens all three at once: the agent binds to the gateway, the gateway chooses the model, and data stays on the local model — a cloud model is used only for permitted tasks and behind masking.

  • Technical: prompts and tools bind to the gateway contract, not to a model.
  • Commercial: changing providers is a routing rule.
  • Compliance: sensitive data on the local model; cloud only when policy allows.

What is an LLM gateway, and what does it do?

The LLM gateway (the AI Gateway in Zzeti) is the single API through which applications and agents call models. Behind it sit local open-weight models served with vLLM or Ollama and — if allowed — more than 30 cloud providers. The gateway handles authentication, routing, load balancing, semantic caching, rate and budget limits and logging; PII masking runs in front of it.

Routing follows policy: the local model by default; semantic routing by task (short classification to a small model, long analysis to a large one); a semantic-cache hit never reaches a model at all; if a model fails to respond, the defined fallback takes over. In an air-gapped installation the cloud route is closed entirely.

  • One API: agents and applications call the gateway and need not know the model.
  • Routing: local-first, semantic by task, with fallback.
  • Load balancing: several model replicas behind one endpoint.
  • Semantic cache: repeated questions answered without reaching a model.
  • Limits: per-LLM budgets, monthly limits, rate limiting.
  • Security: PII masking and policy checks in front of the gateway.

Model-agnostic in practice: what changes when the model changes?

In a well-built architecture the answer is “almost nothing”. Agents, workflows, guardrails and evaluation scenarios are bound to the gateway's contract. When a new open-weight model appears, it is pulled from Model Hub, served on a GPU server and compared with the current model in the Playground on the same test cases. If it clears the AI Judge threshold, a single routing rule puts it into service; if it does not, nothing changes.

The same mechanism applies to fine-tuning: a model adapted to the organization's data with LoRA in Fine-tune Studio is just another option in the gateway — the agents never notice.

  • What changes: the model file in Model Hub and one routing rule.
  • What does not: agents, prompts, tool calls, guardrails, test cases.
  • Decision mechanism: Playground comparison + AI Judge threshold.
  • Rollback: restore the previous rule.

Model observability: what is measured, and where?

Observability is classic application monitoring adapted to LLMs, but what is measured differs. A session breaks down into agent steps; each step into model and MCP calls. For every call, token count, latency, sources used (which collection, which function), executed SQL and the agent's decision are recorded. Cost is attributed per team, agent and model; local models carry zero token cost, but GPU time is still visible.

All of this is collected under Monitor in Zzeti: gateway traffic, the agent fleet, latency by LLM, MCP and agent, success rate, alerts and a live topology map. The records are kept on the organization's own instance; no telemetry leaves. It is also the answer to a KVKK review's question of who accessed what, and when.

Zzeti AI Gateway topology: agents, MCP servers and open-weight LLMs on local GPUs, with latency measured per call (representative screen, sample data)
Gateway topology in your own data center. Open-weight LLMs run on local GPUs; the AI Gateway is the hub between agents, MCP servers and models — every agent, model and MCP call is traced and its latency measured from one place.
  • Trace: session → agent step → model/MCP call; debug view.
  • Tokens and latency: per model, agent and connector.
  • Cost: team/agent/model, with budgets and limits.
  • Health: success rate, agent fleet, alerts, topology.
  • Audit: who asked what, which data was read, which action was taken.

Measuring quality: AI Judge, test cases, feedback

Measuring latency and cost is easy; measuring quality takes discipline. The method is to keep a test-case set for every agent: real questions, expected behaviour and pass thresholds. AI Judge scores those cases with an evaluator model; user feedback (approvals and corrections) and integrity checks complete the score. A new model, a new prompt or a new connector passes the same set before it goes live.

Without this discipline “model-agnostic” stays a slogan: you can switch, but you cannot tell whether it got better.

  • Test set: real questions + expected behaviour + threshold.
  • AI Judge: scoring with an evaluator model; pass threshold.
  • Feedback loop: user corrections enter the test set.
  • Rule: every change that goes live passes the same set.

Cost governance: budgets, limits, cache

The gateway does not only measure cost, it governs it. Per-LLM budgets and monthly limits stop an agent or a team from exhausting the budget with an unexpected load; rate limiting protects the system at peak; the semantic cache answers repeated questions without a model call. Because local models produce no token cost, the gateway's local-first routing is the most effective savings rule of all.

Good practice is to report cost by business unit: which team, which scenario, which model. Without that report, AI cost remains an “IT expense”; with it, it becomes a line that can be compared with the leak each scenario closes.

  • Budgets and limits: per LLM, per month; alert or stop on overrun.
  • Rate limiting: protection at peak.
  • Semantic cache: zero model cost on repeated questions.
  • Local-first routing: removes the token cost itself.
  • Reporting: by team, scenario and model.
Takeaway

The LLM gateway separates agents from the model; observability makes every call measurable. Together they are the foundation of LLMOps: agents do not change when the model does, cost is visible by business unit, quality is proven with a test set rather than a hunch — and every record stays on the organization's own instance.

Frequently asked questions

Is an LLM gateway the same as the AI Gateway?

Yes. The AI Gateway in Zzeti is an LLM gateway: one API in front of local and cloud models that handles authentication, routing, load balancing, semantic caching, limits and logging. Applications and agents call the gateway; which model answers is a policy decision made inside the installation.

Does observability data leave the organization?

No. Traces, token and cost counts, quality scores and the audit log are kept on the organization's own instance. There is no telemetry and no phone-home; an air-gapped installation has no external connection at all.

Can we put our existing cloud LLM subscription behind the gateway?

Yes, if you allow it. The gateway routes non-sensitive tasks to the cloud provider while sensitive data stays on the local model; the same masking, access rules and cost limits apply. Many organizations start fully local and add cloud models later — or never.

Who evaluates model quality — another model?

AI Judge scores the test cases with an evaluator model, but the threshold, the test set and the acceptance decision belong to the organization. User feedback and integrity checks complete the score. The evaluator model can run locally too.

How Zzeti implements this

The same architecture, installed as a fixed-scope project

Zzeti Zeka brings the gateway, RAG, MCP connectors, agents and observability ready-made; Skyloop installs it on your infrastructure — on-premise, air-gapped or private cloud — with the scope and the metric defined up front, not as hourly consulting.