An AI gateway and an inference engine occupy different parts of the request path. The gateway can authorise callers and choose a route; the engine loads a model, validates inputs and schedules GPU work. Clients should evaluate both components and assign an owner to the boundary between them.

This guide covers vLLM, SGLang, Ollama and LM Studio. Their roles and deployment choices differ. Use the official documentation and scoped cases to evaluate the configuration you actually run, alongside the gateway that routes requests to it.

Choose the right evaluation boundary

Component Deployment boundary to establish Capacity evidence to request
vLLM Public inference endpoints, distributed channels, enabled input features and model-loading configuration. Queue time, preemption, cache reuse, TTFT and token latency for the model and topology.
SGLang Inference API, administrative operations, distributed workers and approved model/adapter sources. Cache-hit and miss workloads, session lifecycle, memory pressure and worker recovery.
Ollama Local API versus cloud execution, network access and model-management permissions. Model residency, context allocation, parallel requests, queue rejection and CPU/GPU placement.
LM Studio API-token permissions, server binding and the device selected by LM Link. JIT versus explicit model loading, idle TTL, Auto-Evict and cold/warm request latency.

The table is an evaluation framework. It does not establish integration support, relative speed or the security of a product from its name. Record the release, model, enabled features and effective configuration with each answer.

SGLang: separate API, administration and worker trust

Current server arguments distinguish --api-key from --admin-api-key for designated control endpoints. Both are optional configuration fields. Establish which endpoints the installed release protects, and isolate worker communication independently. Model revisions and any use of --trust-remote-code also need approval.

The SGLang case separates the March serialization/replay findings from a later July coordinated notice. Their prerequisites and remediation records differ; the case preserves conflicting version information rather than presenting one universal fixed release.

Session-aware radix caching uses session references to influence eviction. Closing a session removes its references without immediately freeing all reusable KV; references can still be evicted under pressure. Evaluate cache behaviour and tenant permissions separately. A session_id used for cache lifecycle is not proof of an authorised tenant boundary.

Ollama: distinguish local service and cloud execution

Ollama documents no authentication for the local API, while access to its cloud API uses credentials. Its deployment FAQ describes loopback binding by default, network configuration and a local-only setting such as OLLAMA_NO_CLOUD=1. Identify the selected model and execution mode; a local API address alone does not establish the entire processing route.

Separate prompt submission from model creation, imports and exports. The GGUF model-import case explains a reported tensor-validation issue and why a registry patch field needs artifact evidence.

Use ollama ps and the configured context, parallelism, queue and keep_alive settings in the capacity review. The context guidance explains that larger context requires more memory; current documentation can evolve, so record the effective value rather than assume one default applies to every release and machine. Test model switching and overload as well as repeated requests to a warm model.

LM Studio: verify access and the execution device

Native API-token authentication is documented for 0.4.0 or newer and can be required with selected permissions. The vendor recommends authentication when serving beyond localhost. Confirm the effective setting on the API and SDK paths your application uses.

With LM Link, a localhost request can be served by a linked remote machine. Record that device and its approved data boundary. The vendor's offline-operation guidance describes local use of downloaded models; model discovery, runtime downloads and enabled integrations deserve their own network review.

The LM Studio configuration case is labelled documented behaviour. It covers access, execution location and model residency. Idle TTL and Auto-Evict affect JIT-loaded models differently from explicit loads, so measure first-load latency and model-switching behaviour under the settings you will use.

vLLM: inventory public and internal interfaces

Current vLLM security documentation describes API-key protection for /v1, /v2, /inference and /cohere prefixes. Other inference, operational, development and profiling endpoints have different exposure conditions. Record which endpoints exist under your enabled features and how ingress policy controls them.

The same guidance warns about insecure defaults in several inter-node communication paths. Map the actual backend and topology rather than describe every deployment as a single ZeroMQ channel. The V0 PyNcclPipe case is specifically about KV-cache transfer and has a historical version range and fix.

Approve model assets and input features separately

The Python optimisation case requires loading malicious model configuration while Python uses -O or PYTHONOPTIMIZE=1. These settings differ from vLLM's performance flags. Model approval, runtime configuration and request authorisation are separate controls.

The input-validation case collects prompt-embedding validation and video-frame limits with their separate evidence. Optional features deserve separate limits and validation. The concurrency follow-up establishes an invariant-check bypass; its reproduction did not establish a live-server crash or code execution.

Distinguish tenant-isolation mechanisms

vLLM documents optional per-request cache salting to separate prefix-cache reuse between trust groups and reduce timing inference concerns. That is distinct from process isolation, GPU allocator behaviour and kernel correctness.

The GGUF case concerns integer truncation in specific dequantisation kernels, leaving portions of an output tensor uninitialised. The current vendor advisory lists 0.24.0 as patched. It is not a general finding that every user's KV cache is left unwiped, and cache salting is not its remediation.

Measure cache pressure and scheduling tradeoffs

The current V1 tuning guidance describes decode-prioritised scheduling and chunked prefill enabled where possible. Under KV-cache pressure, requests can be preempted and later recomputed. The effect on latency and throughput depends on the workload.

Measure time to first token, inter-token latency, end-to-end latency, queue time, preemption and memory under your input/output mix. Vary admitted concurrency and token batching deliberately. Avoid universal claims such as “concurrency drops to zero” or a fixed CPU request-per-second ceiling.

Cache-aware routing introduces another tradeoff. Production Stack documents prefix-aware routing and load-aware routing. Reusing a warm prefix can reduce computation, but concentrating traffic on a warm worker can increase queue pressure.

Technical detail: distributed recovery and startup

vLLM's parallelism documentation supports multi-node Ray and multiprocessing options. Ray is not universally required. Ray Serve documents fault tolerance when the deployment and KubeRay recovery mechanisms are configured appropriately.

Record tensor, pipeline and data-parallel layout; node and process failure domains; and who restores workers or the head service. Evaluate tokenisation and API CPU work separately from GPU inference. For startup, record weight storage, download caching, compilation caches, memory profiling and warm readiness. Publish numeric startup or throughput claims only with the hardware, software and reproducible workload.

Questions for your evaluation

Use the inference page of the worksheet. Attach an endpoint inventory, topology, model and feature configuration, isolation evidence and a workload-specific capacity report. Record unsupported features and untested outcomes explicitly.

Applying this to OneVir

OneVir describes separate local and provider execution paths. Our implementation evidence record identifies local GGUF execution through llama.cpp bindings, a separate inference implementation from vLLM.

When evaluating a separately configured upstream such as vLLM, SGLang, Ollama or LM Studio, request backend evidence alongside the OneVir route configuration. Gateway access and budget policy do not establish the runtime's patch status, execution location, internal-network isolation or GPU correctness. This guide does not claim that OneVir includes these runtimes or supports every endpoint they offer. A local engine choice also needs its own model-format, memory-isolation and capacity evaluation.