# AI gateway provider directory

Published by OneQuill, developer of OneVir. This is a documentary review of selected public sources, not an independent vendor security audit.

Product coverage is a researched snapshot, not an exhaustive inventory. Category and deployment mode describe the cited offering; editions, contracts and configuration can change the scope. Selected cases illustrate evaluation questions. Advisory counts are not security rankings.

Sources reviewed: 2026-10-07

## agentgateway

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

Rust gateway for AI and MCP. Local rate limits are held in memory; a remote rate-limit service supplies a different sharing boundary.

Evaluation question: Who can author policies and reference credentials across namespaces?

- [Product / project](https://agentgateway.dev/)
- [Documentation / policy](https://agentgateway.dev/docs/standalone/latest/documentation/configuration/resiliency/rate-limits/)

- [Related case](https://onequill.dev/resources/agentgateway-namespace-isolation)

## AI/ML API

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Hosted model API. Service key management does not replace review of provider processing and retention terms.

Evaluation question: Can keys be scoped and revoked without interrupting all applications?

- [Product / project](https://aimlapi.com/)
- [Documentation / policy](https://docs.aimlapi.com/api-references/service-endpoints/api-key-management)



## AIHubMix

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Hosted model aggregation. Review the service policy alongside the actual selected upstream.

Evaluation question: What provider, geography and content-retention terms govern this model?

- [Product / project](https://aihubmix.com/developers)
- [Documentation / policy](https://docs.aihubmix.com/en/terms-and-privacy/Privacy)



## Apache APISIX

Category: API platforms. Deployment: Self-hosted. Reviewed: 2026-10-07.

AI plugins extend the API gateway. Documented streaming timeouts can close a stream without a DONE event.

Evaluation question: How does the client detect an incomplete stream and avoid unsafe replay?

- [Product / project](https://apisix.apache.org/ai-gateway/)
- [Documentation / policy](https://apisix.apache.org/docs/apisix/plugins/ai-proxy/)



## Apigee / Model Armor

Category: Cloud gateways. Deployment: Managed. Reviewed: 2026-10-07.

Model Armor is a configured integration. Inspection limits, supported regions and service latency affect coverage.

Evaluation question: What is the explicit fail-open or fail-closed policy for inspection errors?

- [Product / project](https://docs.cloud.google.com/model-armor/model-armor-apigee-integration)
- [Documentation / policy](https://docs.cloud.google.com/apigee/docs/api-platform/tutorials/using-model-armor-policies)



## AWS Bedrock AgentCore Gateway

Category: Cloud gateways. Deployment: Managed. Reviewed: 2026-10-07.

AgentCore Gateway now documents inference targets as well as tool integration. Policy-language and target support still have limits.

Evaluation question: Which inference and tool targets are in scope for each policy?

- [Product / project](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway.html)
- [Documentation / policy](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway-targets-inference.html)



## Azure API Management

Category: Cloud gateways. Deployment: Managed, Self-hosted. Reviewed: 2026-10-07.

Token-limit counters are independent across specified gateway, region and workspace boundaries. A newer AI gateway preview has separate limitations.

Evaluation question: Is your budget global, or is it enforced independently per region?

- [Product / project](https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities)
- [Documentation / policy](https://learn.microsoft.com/en-us/azure/api-management/llm-token-limit-policy)



## Bifrost

Category: Specialist gateways. Deployment: Self-hosted, Managed. Reviewed: 2026-10-07.

Gateway with MCP and custom-plugin administration. The September advisories distinguish dynamically linked builds from published static Docker images.

Evaluation question: Is the management API authenticated and isolated from application callers?

- [Product / project](https://www.getbifrost.ai/)
- [Documentation / policy](https://docs.getbifrost.ai/)

- [Related case](https://onequill.dev/resources/bifrost-management-api-boundaries)

## Braintrust AI Proxy

Category: Specialist gateways. Deployment: Managed, Customer-hosted. Reviewed: 2026-10-07.

Proxy and observability sit within Braintrust's architecture. Data-plane hosting and logging choices need to be assessed together.

Evaluation question: Where are request logs and evaluation datasets stored?

- [Product / project](https://www.braintrust.dev/)
- [Documentation / policy](https://www.braintrust.dev/docs/platform/architecture)



## Cloudflare AI Gateway

Category: Cloud gateways. Deployment: Managed. Reviewed: 2026-10-07.

Account-level AI Gateway permissions can cover all gateways, including stored BYOK credentials. Logging behaviour changed for new customers from 24 September 2026.

Evaluation question: How are tenants separated and which logging generation applies to your account?

- [Product / project](https://developers.cloudflare.com/ai-gateway/)
- [Documentation / policy](https://developers.cloudflare.com/ai-gateway/configuration/authentication/)



## CometAPI

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Hosted model aggregator. Published privacy claims need a deployment-specific contractual scope.

Evaluation question: Which upstreams receive content and which contractual exceptions apply?

- [Product / project](https://www.cometapi.com/about/)
- [Documentation / policy](https://www.cometapi.com/privacy-policy/)



## Databricks AI Gateway

Category: Cloud gateways. Deployment: Managed. Reviewed: 2026-10-07.

Unity-oriented gateway services and legacy Mosaic endpoints have different controls. Inference-table payload capture is a separate configuration.

Evaluation question: Which service generation and payload-capture settings are enabled?

- [Product / project](https://docs.databricks.com/aws/en/ai-gateway/)
- [Documentation / policy](https://docs.databricks.com/gcp/en/ai-gateway/model-services)



## Eden AI

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Hosted aggregation across AI services. The data-processing agreement and provider exceptions matter alongside security descriptions.

Evaluation question: Does the DPA cover each selected service and subprocessor?

- [Product / project](https://www.edenai.co/)
- [Documentation / policy](https://www.edenai.co/dpa)



## Envoy AI Gateway / Agent Router

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

Project naming and repository have moved towards Agent Router. Advisories retain the Envoy AI Gateway name and component scope.

Evaluation question: Which release and MCP request-processing path are actually deployed?

- [Product / project](https://aigateway.envoyproxy.io/)
- [Documentation / policy](https://github.com/theagentrouter/agent-router)

- [Related case](https://onequill.dev/resources/envoy-ai-gateway-mcp-message-smuggling)
- [Related case](https://onequill.dev/resources/envoy-ai-gateway-mcp-request-limits)

## F5 NGINX Gateway Fabric

Category: API platforms. Deployment: Self-hosted. Reviewed: 2026-10-07.

AI Guardrails integration is a configured stack feature, not a property of every NGINX proxy. Verify policy acceptance and request-size handling.

Evaluation question: Which guardrail integration and enforcement status are active?

- [Product / project](https://docs.nginx.com/nginx-gateway-fabric/)
- [Documentation / policy](https://docs.nginx.com/nginx-gateway-fabric/how-to/f5-ai-guardrails/)



## FastRouter

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Hosted router with routing and privacy descriptions. Confirm the effective configuration rather than infer it from feature names.

Evaluation question: What happens to data and cost when a route falls back?

- [Product / project](https://fastrouter.ai/features)
- [Documentation / policy](https://fastrouter.ai/privacy)



## GitLab AI Gateway

Category: Related products. Deployment: Managed, Self-hosted. Reviewed: 2026-10-07.

Application-specific gateway for GitLab Duo, not a general interchangeable model router. Hosted and self-hosted patch responsibilities differ.

Evaluation question: Which Duo access and flow-template capabilities are exposed?

- [Product / project](https://docs.gitlab.com/administration/gitlab_duo/)
- [Documentation / policy](https://docs.gitlab.com/releases/patches/other-patches/patch-release-gitlab-ai-gateway-19-4-1-released/)

- [Related case](https://onequill.dev/resources/gitlab-ai-gateway-template-sandbox)

## Gravitee

Category: API platforms. Deployment: Self-hosted, Managed. Reviewed: 2026-10-07.

AI policies and masking are edition and flow-order dependent. Payload logging introduces its own memory and retention needs.

Evaluation question: Is masking applied before every logger and export in this edition?

- [Product / project](https://www.gravitee.io/platform/ai-gateway)
- [Documentation / policy](https://documentation.gravitee.io/apim/create-and-configure-apis/apply-policies/policy-reference/data-logging-masking)



## Helicone

Category: Specialist gateways. Deployment: Managed, Self-hosted. Reviewed: 2026-10-07.

Observability can capture requests automatically. Retention and self-hosting depend on the selected service and configuration.

Evaluation question: Which content fields are logged, for how long, and under whose account?

- [Product / project](https://www.helicone.ai/)
- [Documentation / policy](https://docs.helicone.ai/getting-started/quick-start)



## Higress

Category: API platforms. Deployment: Self-hosted. Reviewed: 2026-10-07.

Gateway with AI provider and security-guard plugins. Plugin selection and policy order define the effective behaviour.

Evaluation question: Which request and response paths are covered by the selected plugins?

- [Product / project](https://higress.io/en/ai-gateway)
- [Documentation / policy](https://github.com/higress-group/higress/security)



## Hugging Face Inference Providers

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Hugging Face's own request handling and metadata retention differ from upstream provider policies. Dedicated Inference Endpoints are a separate service.

Evaluation question: Which provider receives the request, and are its retention terms approved?

- [Product / project](https://huggingface.co/docs/inference-providers/en/index)
- [Documentation / policy](https://huggingface.co/docs/inference-providers/en/security)



## IBM API Connect

Category: API platforms. Deployment: Managed, Self-hosted. Reviewed: 2026-10-07.

AI support and limits vary between SaaS and software versions, including watsonx integrations.

Evaluation question: Which AI features and limitations apply to your deployment and release?

- [Product / project](https://www.ibm.com/docs/en/api-connect/cloud/saas?topic=applications-using-ai-gateway-support-watsonxai-apis)
- [Documentation / policy](https://www.ibm.com/docs/en/api-connect/software/12.1.1?topic=overview-known-limitations)



## kgateway

Category: API platforms. Deployment: Self-hosted. Reviewed: 2026-10-07.

Kubernetes gateway with AI extensions. Envoy-based documentation is versioned; prompt guards and rate limits need configured policies.

Evaluation question: What happens if an external guard or rate-limit service is unavailable?

- [Product / project](https://kgateway.dev/)
- [Documentation / policy](https://kgateway.dev/docs/envoy/2.1.x/ai/about/)



## Kong AI Gateway

Category: API platforms. Deployment: Managed, Self-hosted. Reviewed: 2026-10-07.

Traditional AI Proxy plugins and newer AI Gateway offerings have distinct version and deployment scopes.

Evaluation question: How is streaming usage reconciled against provider invoices?

- [Product / project](https://developer.konghq.com/index/ai-gateway/)
- [Documentation / policy](https://developer.konghq.com/plugins/ai-proxy/changelog/)

- [Related case](https://onequill.dev/resources/kong-gemini-streaming-token-accounting)

## LangDB

Category: Specialist gateways. Deployment: Managed, Customer-hosted. Reviewed: 2026-10-07.

AI gateway with observability. Documented ClickHouse TTL deletion runs through asynchronous merges.

Evaluation question: What deletion latency applies to logs, backups and exports?

- [Product / project](https://langdb.ai/)
- [Documentation / policy](https://docs.langdb.ai/enterprise/resources/configuring-data-retention/)



## Leanroute

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Policy dated 1 October 2026 distinguishes database storage from response and semantic caches, which default to 15 minutes; no-persistence is a separate choice.

Evaluation question: Are response caching and semantic caching disabled for sensitive routes?

- [Product / project](https://leanroute.dev/)
- [Documentation / policy](https://leanroute.dev/privacy)



## LiteLLM

Category: Specialist gateways. Deployment: Self-hosted, Managed. Reviewed: 2026-10-07.

Python proxy and SDK distribution are distinct from the vendor's pinned proxy Docker distribution.

Evaluation question: What package or image digest is running, and how are credentials rotated after an incident?

- [Product / project](https://www.litellm.ai/)
- [Documentation / policy](https://docs.litellm.ai/docs/proxy/prod)

- [Related case](https://onequill.dev/resources/litellm-march-2026-package-incident)

## LLM Gateway

Category: Hosted routers. Deployment: Managed, Self-hosted. Reviewed: 2026-10-07.

Metadata-only logging is described as the default; payload logging is an opt-in setting. Upstream retention remains separate.

Evaluation question: Which content logging options and upstream terms are enabled?

- [Product / project](https://llmgateway.io/)
- [Documentation / policy](https://llmgateway.io/legal/privacy)



## LM Studio

Category: Inference engines. Deployment: Self-hosted. Reviewed: 2026-10-07.

Desktop and headless model runtime with an API server. Authentication is optional; network serving and LM Link alter the access and execution boundary.

Evaluation question: Are API-token permissions enforced, which machine executes each model, and what model-loading and eviction behaviour does the application rely on?

- [Product](https://lmstudio.ai/)
- [API-token authentication](https://lmstudio.ai/docs/developer/core/authentication)
- [Network server settings](https://lmstudio.ai/docs/developer/core/server/serve-on-network)
- [LM Link remote execution](https://lmstudio.ai/docs/developer/core/lmlink)
- [Offline operation](https://lmstudio.ai/docs/app/offline)

- [Related case](https://onequill.dev/resources/lm-studio-network-authentication-and-lifecycle)

## Lunar.dev

Category: Specialist gateways. Deployment: Self-hosted, Managed. Reviewed: 2026-10-07.

Gateway and traffic-management offering. Production logs and data-plane placement need their own review.

Evaluation question: Which data leaves the gateway through logs or management integrations?

- [Product / project](https://www.lunar.dev/product/ai-gateway)
- [Documentation / policy](https://docs.lunar.dev/api-gateway/lunar-dev-in-production/lunar-logs)



## Martian

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Hosted routing gateway. Authentication documentation establishes an API boundary, not every enterprise data-control claim.

Evaluation question: What route decision evidence and data-processing terms can be supplied?

- [Product / project](https://withmartian.com/)
- [Documentation / policy](https://gateway-docs.withmartian.com/api-reference/authentication)



## MLflow AI Gateway

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

Current gateway documentation distinguishes development credential-encryption defaults from production key management and SQL-backed rotation.

Evaluation question: Is a production encryption secret configured and is rotation rehearsed?

- [Product / project](https://mlflow.org/docs/latest/genai/governance/ai-gateway)
- [Documentation / policy](https://www.mlflow.org/docs/latest/genai/governance/ai-gateway/api-keys/key-rotation/)



## New API

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

Project with quota and billing functions. Release-candidate version scope matters in the quota-overflow advisory.

Evaluation question: Can credits, reservations and final charges be reconciled under concurrent requests?

- [Product / project](https://docs.newapi.pro/)
- [Documentation / policy](https://docs.newapi.pro/en/docs/guide/wiki/changelog)

- [Related case](https://onequill.dev/resources/new-api-quota-billing-overflow)

## nexos.ai

Category: Specialist gateways. Deployment: Managed. Reviewed: 2026-10-07.

Hosted gateway and AI platform. Procurement needs the service contract, provider list and configured routing policy.

Evaluation question: What geography and retention commitments cover every fallback?

- [Product / project](https://nexos.ai/ai-gateway/)
- [Documentation / policy](https://nexos.ai/legal/privacy-policy/)



## Not Diamond

Category: Related products. Deployment: Managed, Customer-hosted. Reviewed: 2026-10-07.

Routing and model-selection API. It is listed as a related product rather than assumed to implement a full gateway boundary.

Evaluation question: Which controls reside in your calling application and which in the routing service?

- [Product / project](https://docs.notdiamond.ai/docs/what-is-not-diamond)
- [Documentation / policy](https://docs.notdiamond.ai/docs/privacy-security-and-local-deployments)



## Ollama

Category: Inference engines. Deployment: Self-hosted, Managed. Reviewed: 2026-10-07.

Local model runtime and API, with optional cloud-model features. The local API does not require authentication; cloud API credentials and cloud processing are a separate mode.

Evaluation question: Is execution local or cloud, who can reach inference and model-management endpoints, and how do context, parallel requests and model residency fit memory?

- [Project](https://ollama.com/)
- [Local and cloud API authentication](https://docs.ollama.com/api/authentication)
- [Network, cloud and concurrency settings](https://docs.ollama.com/faq)
- [Context and memory](https://docs.ollama.com/context-length)

- [Related case](https://onequill.dev/resources/ollama-model-import-memory-boundary)

## One API

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

Separate project from New API. Model mappings and protocol reconstruction can affect unsupported fields.

Evaluation question: Which request fields survive translation, including tools and usage options?

- [Product / project](https://github.com/songquanpeng/one-api)
- [Documentation / policy](https://github.com/songquanpeng/one-api)



## OneVir

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

OneQuill's own product combines gateway controls and local execution paths. This profile is affiliated and is not an independent security assessment.

Evaluation question: Which controls are enabled in your installed version, and which remain the inference backend's responsibility?

- [Product / project](https://onevir.onequill.dev/)
- [Documentation / policy](https://onevir.onequill.dev/#capabilities)



## OpenRouter

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Provider routing, fallback and endpoint-level privacy controls need to be configured together. ZDR and geography are route properties.

Evaluation question: Can every selected endpoint and fallback satisfy the same privacy constraints?

- [Product / project](https://openrouter.ai/)
- [Documentation / policy](https://openrouter.ai/docs/guides/features/guardrails/overview)



## OpenZiti LLM Gateway

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

Identity-based overlay networking is relevant to gateway access. Platform network features do not establish every application-level safeguard.

Evaluation question: How are service identities issued, revoked and separated between tenants?

- [Product / project](https://blog.openziti.io/ai-secops-why-your-ai-infrastructure-has-a-network-shaped-blind-spot)
- [Documentation / policy](https://openziti.io/docs/learn/introduction/features/)



## Opper

Category: Specialist gateways. Deployment: Managed. Reviewed: 2026-10-07.

Security overview updated 5 October 2026: upstream inference is not EEA-restricted by default. Tracing, backups and provider retention have separate rules.

Evaluation question: Does the selected route meet geography and retention needs, including backups?

- [Product / project](https://opper.ai/)
- [Documentation / policy](https://opper.ai/security-overview)



## orq.ai

Category: Specialist gateways. Deployment: Managed, Customer-hosted. Reviewed: 2026-10-07.

Gateway within an orchestration platform. Vendor security claims and logging retention must be matched to the purchased deployment.

Evaluation question: Which attestations and retention settings cover this service and edition?

- [Product / project](https://orq.ai/platform/ai-gateway)
- [Documentation / policy](https://orq.ai/legal/security)



## Plano

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

Agent proxy project formerly named Arch. Establish the current component boundaries and supported deployment before comparing controls.

Evaluation question: Which tool, model and policy paths pass through Plano?

- [Product / project](https://planoai.dev/)
- [Documentation / policy](https://docs.planoai.dev/get_started/overview.html)



## Portkey / Prisma AIRS AI Gateway

Category: Specialist gateways. Deployment: Self-hosted, Managed. Reviewed: 2026-10-07.

Portkey commercial offerings have moved into Prisma AIRS. The selected advisory applies to the open-source Portkey gateway; it does not establish managed-service exposure.

Evaluation question: Which product, build and custom-host policy are you evaluating?

- [Product / project](https://portkey.ai/features/ai-gateway)
- [Documentation / policy](https://www.paloaltonetworks.com/blog/2026/07/announcing-general-availability-of-prisma-airs-ai-gateway/)

- [Related case](https://onequill.dev/resources/portkey-custom-host-ssrf)

## Requesty

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Gateway logging, caching and upstream retention are separate. EU router processing does not establish the upstream inference location.

Evaluation question: Does this plan and route prohibit training and retention at every layer?

- [Product / project](https://www.requesty.ai/)
- [Documentation / policy](https://www.requesty.ai/privacy)



## SGLang

Category: Inference engines. Deployment: Self-hosted. Reviewed: 2026-10-07.

LLM and multimodal serving framework with distributed workers and prefix-cache reuse. The inference API, administrative controls and worker channels have distinct trust boundaries.

Evaluation question: Which API and admin keys, worker networks, model-loading features and cache isolation are configured in the installed release?

- [Project](https://github.com/sgl-project/sglang)
- [Server arguments and access controls](https://docs.sglang.io/docs/advanced_features/server_arguments)
- [Session-aware radix cache](https://github.com/sgl-project/sglang/blob/main/docs/docs/advanced_features/session_radix_cache.mdx)

- [Related case](https://onequill.dev/resources/sglang-worker-and-management-trust)

## TensorZero

Category: Specialist gateways. Deployment: Self-hosted. Reviewed: 2026-10-07.

Gateway authentication is an explicit operational configuration with PostgreSQL support. Dashboard access is a separate boundary.

Evaluation question: Are both the gateway and dashboard authenticated for this deployment?

- [Product / project](https://www.tensorzero.com/)
- [Documentation / policy](https://www.tensorzero.com/docs/operations/set-up-auth-for-tensorzero)



## Tetrate Agent Router

Category: Specialist gateways. Deployment: Managed, Customer-hosted. Reviewed: 2026-10-07.

Hosted Agent Router Service and enterprise customer-hosted data plane have different data flows. The management plane is separately hosted.

Evaluation question: What crosses between your data plane and the hosted management plane?

- [Product / project](https://www.tetrate.io/)
- [Documentation / policy](https://docs.tetrate.ai/product-architecture/data-flows/)



## Traefik Hub

Category: API platforms. Deployment: Managed, Self-hosted. Reviewed: 2026-10-07.

Hub AI middlewares include Redis-backed shared token limits. Interrupted streams and response-completed events affect usage measurement.

Evaluation question: Does incomplete-stream usage reach the budget ledger?

- [Product / project](https://doc.traefik.io/traefik-hub/ai-gateway/overview)
- [Documentation / policy](https://doc.traefik.io/traefik-hub/ai-gateway/middlewares/token-rate-limit)



## TrueFoundry

Category: Specialist gateways. Deployment: Managed, Customer-hosted. Reviewed: 2026-10-07.

Gateway, traces and deployment platform. Residency-aware fallback is a vendor-described configuration, not proof of your route policy.

Evaluation question: Can the vendor show that all retries and fallback targets satisfy your residency policy?

- [Product / project](https://www.truefoundry.com/ai-gateway)
- [Documentation / policy](https://www.truefoundry.com/docs/ai-gateway/feedback-for-traces)



## Tyk AI Studio / MCP Gateway

Category: API platforms. Deployment: Managed, Self-hosted. Reviewed: 2026-10-07.

AI Studio, chat and MCP offerings are separate capabilities in the Tyk ecosystem.

Evaluation question: Which model-routing and tool-access controls are available in the purchased product?

- [Product / project](https://tyk.io/tyk-mcp-gateway/)
- [Documentation / policy](https://tyk.io/blog/introducing-tyk-ai-studio-welcome-to-the-future-of-ai-governance-powered-by-ai-chat-gateway-and-portal/)



## Vercel AI Gateway

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Gateway content deletion and provider ZDR arrangements have separate scope and eligibility.

Evaluation question: Which providers and plan-specific agreements implement the required ZDR setting?

- [Product / project](https://vercel.com/ai-gateway)
- [Documentation / policy](https://vercel.com/i/secure-ai-gateway)



## vLLM

Category: Inference engines. Deployment: Self-hosted. Reviewed: 2026-10-07.

Inference engine and serving stack. API authentication, internal networks, model loading and GPU scheduling are separate boundaries from a gateway.

Evaluation question: Which engine, features, model format and cluster topology are deployed?

- [Product / project](https://docs.vllm.ai/)
- [Documentation / policy](https://docs.vllm.ai/en/latest/usage/security/)

- [Related case](https://onequill.dev/resources/vllm-kv-transfer-network-isolation)
- [Related case](https://onequill.dev/resources/vllm-model-loading-python-optimisation)
- [Related case](https://onequill.dev/resources/vllm-multimodal-input-validation)
- [Related case](https://onequill.dev/resources/vllm-gguf-gpu-memory-isolation)

## WSO2 API Manager

Category: API platforms. Deployment: Self-hosted, Managed. Reviewed: 2026-10-07.

AI token policies and backend throttling have different enforcement paths. Local per-node limits differ from distributed Traffic Manager policies.

Evaluation question: Which counters are shared across replicas and regions?

- [Product / project](https://apim.docs.wso2.com/en/latest/ai-gateway/rate-limiting/)
- [Documentation / policy](https://apim.docs.wso2.com/en/latest/api-design-manage/design/rate-limiting/protect-backend-services/)



## ZenMux

Category: Hosted routers. Deployment: Managed. Reviewed: 2026-10-07.

Turning off Data Services changes ZenMux logging and related features. It does not independently establish upstream zero retention.

Evaluation question: Which data services are disabled and what provider retention remains?

- [Product / project](https://zenmux.ai/)
- [Documentation / policy](https://zenmux.ai/docs/guide/advanced/data-services.html)


