AI Gateway

KONEKT AI Gateway

On-prem or in a private cloud. Combination of local and public AI models: advanced smart routing, billing & budgets, fine-tuning, governance, central login, and immutable audit. Connects to Workspace via a VPN + TLS 1.3 tunnel.

What Gateway is

Model control in one place

KONEKT AI Gateway sits between your applications and AI models — local, custom, and public — according to company rules. It deploys on-prem or in a private cloud.

Applications and KONEKT Workspace connect to the Gateway over a secure VPN + TLS 1.3 encrypted tunnel. Routing, budgets, fine-tuning and security policies stay under your full control.

On-premise

Gateway in your data center or rack.

Private cloud

Gateway in your private cloud environment.

3. Full provider abstraction & Custom Models

Swap the backend model without rewriting apps, or run purpose-built (fine-tuned) models with LoRA adapters trained on your company terminology.

KONEKT AI Gateway — One console. Full control.
Gateway Evaluation Local model fine-tuning

One interface for applications

Applications connect to the Gateway. The backend model changes by policy, without rewriting business software.


4. Virtual Models & 5. Intelligent Routing

Automatic routing by policy

The Gateway dynamically scores cost, latency, and the request's security level. Besides the base aliases, you create extra virtual models by department, confidentiality, budget, time, or fine-tuned model.

Virtual models and routing

Extra virtual models under your conditions

The four base aliases (konekt-auto, konekt-fast, konekt-smart, konekt-private) are the starting point. In the Gateway console you create more — the app still calls the same name, and rules change without code changes.

Department and role

legal-smart, hr-private, it-coding

Confidentiality level

local only, masked, or a public model with permission

Budget and quota

cheaper model after the limit, or a hard stop

Speed and free GPU

local GPU if free, otherwise another allowed path

Time and priority

business hours, night batch, urgent requests

Fine-tuned / evaluation

an alias to an internal model that passed evaluation


Governance features

Full control of cost, access, and models

6. Privacy Policies

Automatic check whether content may leave to an external API or must stay on the internal network.

7. Department budgets

Monthly token and spend quotas for HR, marketing, sales, and legal — with alerts and hard stops.

8. Rate Limiting

Protection against congestion (RPM / TPM limits) and priority for critical services.

9. Detailed Usage Tracking

Real-time usage metrics: prompt tokens, completion tokens, latency, and cost per call.

10. KONEKT Console

A web admin panel for API keys, analytics, audit logs, and service configuration.

11. Private Models (GPU)

Integration with local vLLM models (DeepSeek, Llama, Mistral, Qwen, Kimi, MiniMax), treated as first-class endpoints in the catalog.

12. Strict Fallback Rules and Protection

Fallback rules keep operations running when a cloud provider is down (switch to an alternative model of the same class). For konekt-private the rule is absolute: no cloud failover.

Supported AI ecosystem and providers