# AI gateway and inference case library

Product coverage is a researched snapshot, not an exhaustive inventory. Category and deployment mode describe the cited offering; editions, contracts and configuration can change the scope. Selected cases illustrate evaluation questions. Advisory counts are not security rankings.

- [Portkey: custom-host routing and private-network access](https://onequill.dev/resources/portkey-custom-host-ssrf): Security advisory. Open-source Portkey gateway custom-host routing. Reviewed 2026-10-07.
- [Bifrost: management APIs are an execution boundary](https://onequill.dev/resources/bifrost-management-api-boundaries): Security advisory. Reachable management APIs with authentication disabled; stdio MCP registration and remote custom-plugin loading. Reviewed 2026-10-07.
- [Envoy AI Gateway: one message, two interpretations](https://onequill.dev/resources/envoy-ai-gateway-mcp-message-smuggling): Security advisory. MCP JSON-RPC parsing in the Envoy AI Gateway / Agent Router project. Reviewed 2026-10-07.
- [Envoy AI Gateway: enforce limits before buffering](https://onequill.dev/resources/envoy-ai-gateway-mcp-request-limits): Security advisory. MCP POST-body buffering in the external-processing component. Reviewed 2026-10-07.
- [agentgateway: a patched release can still need a policy setting](https://onequill.dev/resources/agentgateway-namespace-isolation): Security advisory. Cross-namespace backend references authored by Kubernetes namespace administrators. Reviewed 2026-10-07.
- [New API: quota arithmetic is part of the trust boundary](https://onequill.dev/resources/new-api-quota-billing-overflow): Security advisory. Quota settlement in affected New API release candidates. Reviewed 2026-10-07.
- [Kong: a streaming usage fix deserves billing regression tests](https://onequill.dev/resources/kong-gemini-streaming-token-accounting): Release fix. Kong AI Proxy plugin handling of Gemini streaming usage. Reviewed 2026-10-07.
- [LiteLLM: package provenance and incident response](https://onequill.dev/resources/litellm-march-2026-package-incident): Incident report. Malicious PyPI distributions, distinct from the vendor's official pinned Proxy Docker images. Reviewed 2026-10-07.
- [GitLab AI Gateway: flow templates and sandbox assumptions](https://onequill.dev/resources/gitlab-ai-gateway-template-sandbox): Security advisory. GitLab Duo Agent Platform flow-template processing. Reviewed 2026-10-07.
- [vLLM: isolate KV-transfer and distributed communication](https://onequill.dev/resources/vllm-kv-transfer-network-isolation): Security advisory. PyNcclPipe KV-cache transfer in the V0 engine. Reviewed 2026-10-07.
- [vLLM: model configuration and Python optimisation](https://onequill.dev/resources/vllm-model-loading-python-optimisation): Security advisory. Model-configuration loading when Python assertion checks are disabled. Reviewed 2026-10-07.
- [vLLM: input features need separate validation and capacity limits](https://onequill.dev/resources/vllm-multimodal-input-validation): Security advisory. Optional prompt embeddings and base64 video/jpeg frame processing; separate findings collected in one feature-validation case. Reviewed 2026-10-07.
- [vLLM: a GGUF kernel defect is distinct from cache policy](https://onequill.dev/resources/vllm-gguf-gpu-memory-isolation): Security advisory. Integer truncation in specific GGUF dequantisation kernels. Reviewed 2026-10-07.
- [SGLang: isolate worker channels and model-management paths](https://onequill.dev/resources/sglang-worker-and-management-trust): Security advisory. Feature-dependent worker serialization, replay tooling and administrative/model-loading interfaces. Reviewed 2026-10-07.
- [Ollama: model import is a separate memory and access boundary](https://onequill.dev/resources/ollama-model-import-memory-boundary): Security advisory. GGUF tensor-size validation during model creation and quantization. Reviewed 2026-10-07.
- [LM Studio: choose authentication, execution location and model lifecycle](https://onequill.dev/resources/lm-studio-network-authentication-and-lifecycle): Documented behaviour. Documented API authentication, network binding, remote model resolution and JIT model residency. Reviewed 2026-10-07.
