# OneQuill public website: full text Canonical HTML, publication dates and source citations remain authoritative. OneQuill develops OneVir. Public information only; no browser automation is required. --- Canonical: https://onequill.dev/about # About OneQuill · Software that does exactly what it says Canonical: https://onequill.dev/about About OneQuill # Software that does exactly what it says. OneQuill is an independent software company in Essen, Germany. We build tools people can run, understand and trust, starting with OneVir, the governed execution layer for private and hybrid enterprise AI. [Meet OneVir →](https://onevir.onequill.dev/) [Contact OneQuill](https://onequill.dev/contact) **01** What we believe ## Precision is a form of respect. Powerful technology should leave people more in control, not less. Every principle below shows up somewhere you can check in our products. 01 ### Control by default Data stays where you put it. In OneVir, models run on your own hardware first; a request leaves only through a provider route you configured, and a single switch or header keeps it local. 02 ### Every decision explained OneVir records why each request went where it did, what policy decided and what it cost. Audit exports never include prompts, responses or keys. 03 ### Engineered, not assembled We prefer one well-built native process to a stack of loosely coupled services. Fewer moving parts mean fewer places for data, cost or policy to slip. 04 ### Claims you can check We publish how performance numbers were measured, label demonstration data as such, compare ourselves against public documentation, and keep the roadmap separate from what ships. **02** Our mark ## A Q and a quill. The ring of the Q stands for the complete system we take responsibility for. The quill stands for the care of writing things down clearly: the documentation, the records and the explanations that make a system trustworthy. Its colour runs from orange through coral and raspberry to a deep berry: warm energy, held with discipline. [Illustration: OneQuill stacked logo] **03** At a glance ## The company. [Company, press & investor enquiries ↗](mailto:contact@onequill.dev) Company | **OneQuill** Based in | Essen, Germany First product | [OneVir](https://onevir.onequill.dev/), the governed execution layer for private and hybrid enterprise AI Licensing | OneVir is source-available under the Business Source License 1.1 ([how it works](https://onequill.dev/licensing)) Contact | [Write an enquiry](mailto:hello@onequill.dev) · [All contact options](https://onequill.dev/contact) Work with us ## Build something you can trust. Talk to us about OneVir, collaborations or partnerships. [Contact OneQuill →](https://onequill.dev/contact) --- Canonical: https://onequill.dev/contact # Contact OneQuill · Enquiries, sales, support and company Canonical: https://onequill.dev/contact Contact OneQuill # Every road leads to Essen. Get in touch for product questions, walkthroughs, partnerships, support or company enquiries. Choose the option below that best fits your request. ## Contact options Enquiries ### Say hello Questions about OneQuill, where to begin, or anything else. [Write an enquiry ↗](mailto:hello@onequill.dev?subject=Enquiry&body=Hello%20OneQuill%20team%2C%0A%0AI%20would%20like%20to%20ask%20about%3A%0A) Sales and partnerships ### Walkthroughs, licences and partners Product demonstrations, evaluations, pricing, commercial licences, integrations and partnerships. [Request a walkthrough ↗](mailto:sales@onequill.dev?subject=OneVir%20walkthrough&body=Hello%20OneQuill%20team%2C%0A%0AOrganisation%3A%0AUse%20case%3A%0APreferred%20next%20step%3A%0A) [Discuss a partnership ↗](mailto:partners@onequill.dev?subject=OneQuill%20partnership&body=Hello%20OneQuill%20partnerships%20team%2C%0A%0AOrganisation%3A%0APartnership%20idea%3A%0A) Support ### OneVir support Help with a deployment, a product issue, or how OneVir works. [Read the support guide](https://onequill.dev/support) for troubleshooting steps. [Get support ↗](mailto:support@onequill.dev?subject=OneVir%20support%20request&body=Hello%20OneQuill%20support%2C%0A%0AOneVir%20version%20and%20operating%20system%3A%0AIssue%20summary%3A%0ASteps%20to%20reproduce%3A%0AExpected%20result%3A%0AActual%20result%3A%0A) [Ask a product question ↗](mailto:help@onequill.dev?subject=OneVir%20question) Company ### Company, legal and privacy Formal correspondence, company information, press and investor enquiries, privacy requests and licence terms. [Contact the company ↗](mailto:contact@onequill.dev?subject=Company%20enquiry) [Press & investor enquiries ↗](mailto:info@onequill.dev?subject=Press%20or%20investor%20enquiry) Where we are ## Essen, Germany. 51.4556° N, 7.0116° E. For formal operator information, see our [Impressum](https://onequill.dev/legal); for how we handle personal data, see [Privacy](https://onequill.dev/privacy). --- Canonical: https://onequill.dev/ # OneQuill · Powerful technology, under your control Canonical: https://onequill.dev/ OneQuill **/** Software engineered in Essen, Germany # Powerful technology. Under your control. OneQuill builds software for enterprise AI. Our flagship, OneVir, brings local inference and external models together under **one boundary for policy, security, cost and audit.** [Explore OneVir →](https://onevir.onequill.dev/) [Talk to OneQuill](https://onequill.dev/contact) FlagshipOneVir Provider types26 RuntimeRust · native BaseEssen · 51.46° N Drag to spin [Scroll](https://onequill.dev/#flagship) OpenAI SDKAnthropic SDKClaude CodeCursorCodex CLIContinueAiderLangChainn8nLibreChatAnythingLLMAmazon BedrockAzure OpenAIGoogle Vertex AIDeepSeekOpenRouter OpenAI SDKAnthropic SDKClaude CodeCursorCodex CLIContinueAiderLangChainn8nLibreChatAnythingLLMAmazon BedrockAzure OpenAIGoogle Vertex AIDeepSeekOpenRouter **01** Flagship product ## OneVir. Governed execution for private and hybrid enterprise AI. Local inference when data must remain under enterprise control. External models where they add value. OneVir applies one unified boundary for policy enforcement, security, cost management and auditability across both. **One native server. OpenAI, Anthropic and Jev compatible.** [See every capability →](https://onevir.onequill.dev/) [Illustration: OneVir Gateway page: external providers switch, model catalogue, routing and local models with per-model device placement.] [Illustration: OneVir Proxy page: three healthy providers with credentials stored inside OneVir, spend this month and content retention settings.] [Illustration: OneVir Models page: published and draft models, requests, tokens and estimated spend.] [Illustration: OneVir Policy Engine: enforcement overview, remote updates and the Stop governed requests control.] [Illustration: OneVir Activity page: total tokens, requests, a year of token activity and daily tokens.] [Illustration: OneVir Performance page in the dark theme: CPU, memory and GPU gauges and live AI request figures.] [Illustration: OneVir AI BOM help: CycloneDX 1.6 JSON, CSV and printable report exports.] Gateway**Local models placed on the best device, published by name** Current interface with illustrative demonstration data. **02** The request path ## Private or hybrid. Governed at execution. Applications connect only to the gateway. Each request is authenticated, inspected and routed before it runs on your hardware or leaves for a provider you approved, and one record of what happened covers both paths. **Your applications**OpenAI and Anthropic SDKs · curl **Coding agents**Claude Code · Cursor · Codex CLI **Workflows and RAG**LangChain · n8n · LibreChat Governed execution - Kill switch - Application key - Policy and guardrails - Aliases and routing - Response cache Self-hosting**On your hardware**GGUF models on GPU, CPU or a MoE hybrid. Prompts stay on the machine. Proxy**Through your providers**26 provider types, with health checks, failover and a budget for each. ## Measured performance 417tok/s Generation for one client on an RTX 4080 SUPER with Vulkan. 7ms Time to first token for a 22-token prompt. 668tok/s Total across 16 concurrent clients with continuous batching. 26types Cloud provider integrations, used only when you allow them. Measured at the client with OneVir's load generator: Qwen2.5-0.5B-Instruct Q4_K_M (a 0.5-billion-parameter model), greedy decoding, streaming, i9-14900K with RTX 4080 SUPER, Windows 11. Larger models are slower. Your models and hardware will give different numbers. **03** What OneVir does ## Execution and governance. One native server. Run local models, route to approved providers, and apply consistent enterprise controls. Policy, security, cost management and auditability cover both execution paths. ### Inference gateway One OpenAI- and Anthropic-compatible endpoint for every app, agent and workflow, with streaming, tool calls and wire translation. ### Self-hosted engine GGUF models on CPU, Vulkan or CUDA, with automatic placement, MoE expert offload, continuous batching and prompt and response caches. ### 26 provider types OpenAI, Anthropic, Azure, Bedrock, Vertex AI, Gemini, Mistral, Groq, xAI, DeepSeek, OpenRouter and any OpenAI-compatible endpoint. ### Guardrails and privacy Guard models catch prompt injection; personal data is redacted locally; tool actions need approval. Draft, simulate, then activate. ### Signed kill switch Security tooling can stop everything, one agent or one model, through a signed, replay-protected contract that survives restarts. ### FinOps and telemetry Budgets reserved before dispatch, spend by model and day, Prometheus metrics and OpenTelemetry traces that never carry prompts. [All fifteen capabilities →](https://onevir.onequill.dev/#capabilities) **04** How we build ## Engineered to be trusted. Four rules shape every product we ship, from the request path to the words on this page. 01 ### Control by default Data stays where you put it. OneVir serves on your own hardware first, and a request leaves only through a route you configured. 02 ### Every decision explained Routing choices, policy verdicts, timing and spend are recorded, so an operator can always answer why something happened. 03 ### Engineered, not assembled One native Rust server with private workers. No Python, Node or Docker to deploy, and a measured engine at its core. 04 ### Claims you can check We publish how numbers were measured, label demonstration data as such, and call the roadmap a roadmap. **05** The company ## A Q and a quill. Precision, written down. OneQuill is an independent software company in Essen, Germany. We turn complex technology into tools that people can run, understand and trust. - **Built for operators.** Clear controls, honest states and consequences explained before you act. - **Built for auditors.** Records, exports and inventories that stand up to questions. - **Built for developers.** Compatible APIs, so adopting our software never means rewriting yours. [About OneQuill →](https://onequill.dev/about) Insights & Guides ## Understand the rules. Build with context. Research and practical help for teams building with enterprise AI. Read the original sources and connect governance requirements to your deployment. [](https://onequill.dev/resources/ai-governance-rules) AI governance / Research reference ### AI governance rules for gateways and local inference A global reference to the EU AI Act, GDPR, ISO 42001, NIST AI RMF, OSFI E-23 and OWASP, with official sources and clearly stated responsibilities. [Read the reference →](https://onequill.dev/resources/ai-governance-rules)[Browse all resources](https://onequill.dev/resources) Start with one workflow ## Your hardware. Your rules. Your first governed workflow. See local serving, routing, guardrails and observability applied to your own use case. Non-commercial and non-production use of OneVir is free. [Request a walkthrough →](mailto:sales@onequill.dev?subject=OneVir%20walkthrough&body=Hello%20OneQuill%20team%2C%0A%0AI%27d%20like%20a%20walkthrough%20of%20OneVir.%0A%0AUse%20case%3A%20%0AHardware%3A%20%0APreferred%20time%3A%20) [How licensing works](https://onequill.dev/licensing) --- Canonical: https://onequill.dev/legal # Impressum · OneQuill Canonical: https://onequill.dev/legal Legal # Impressum Operator and company information. ## OneQuill Heinrich-Brauns-Straße 45355 Essen Germany Email: [contact@onequill.dev](mailto:contact@onequill.dev) **Preparation copy.** The operator's full legal name and complete postal address are awaiting confirmation and will be added before this page is relied on. ## Additional details Registration, authorised representative and tax information will be added where applicable to the confirmed legal operator. ## Responsibility for content We take care that the information on this website is accurate and current. Product descriptions refer to the version being described; demonstration data in product captures is illustrative. Links to external websites are the responsibility of their operators. ## Related - [Privacy information](https://onequill.dev/privacy) - [OneVir licensing](https://onequill.dev/licensing) and the binding [licence text](https://onequill.dev/LICENSE) - [All contact routes](https://onequill.dev/contact) --- Canonical: https://onequill.dev/licensing # OneVir licensing · Free for non-commercial and non-production use · OneQuill Canonical: https://onequill.dev/licensing OneVir licensing · OneQuill # Free for non-commercial use. Free before production. OneVir is licensed by OneQuill under the Business Source License 1.1. Non-commercial use is free, including in production. Development, testing and evaluation are free for everyone when they do not serve live operation. Only commercial production use requires a paid licence. Free ### Non-commercial use Personal, hobby, study, teaching, research and public-interest activities with no commercial purpose, including live production services. Free ### Non-production use Development, testing, CI, staging, evaluation, proofs of concept and demonstrations in any organisation, including commercial businesses, when they do not serve live operation. Paid licence ### Commercial production Live business tools, commercial products, paid client services, shared business gateways or automated jobs supporting commercial activity. [Ask for a commercial licence ↗](mailto:sales@onequill.dev?subject=OneVir%20commercial%20licence&body=Hello%20OneQuill%20team%2C%0A%0AOrganisation%3A%0ACommercial%20production%20use%3A%0A) ## Two questions decide the licence Is the use commercial, and is it in production? A paid licence is required only when both answers are yes, while BUSL-1.1 applies. ### Commercial production: paid licence - A company's live internal tool or shared gateway - A commercial product or paid client application - A revenue-generating hosted service - Automated jobs doing a business's actual work ### Free use - Non-commercial use, including production - Development, testing, CI, staging and evaluation without live operation - Proofs of concept and demonstrations without live operation - One person's interactive workstation serving only them **Production** means live operation: serving real requests or processing data for actual activities. A pilot or public demo that people rely on is production; non-commercial production remains free. **Non-commercial** means the use does not support a revenue-generating business activity, a paid service or other commercial advantage. Business operations and internal business tools are commercial even if end users are not charged directly. An organisation's non-profit status alone does not determine the answer. ## Examples Situation | Paid licence? Your household's daily personal assistant | No: non-commercial A charity's free public-interest service with no commercial purpose | No, including production A university's non-commercial teaching or research service | No, including production A developer's interactive coding assistant on a workstation serving only them | No: excluded from production A commercial team's staging server, CI tests or evaluation without live operation | No: non-production A consultant's demonstration without live operation | No: non-production A company's live shared gateway or employee chatbot | **Yes: commercial production** A commercial product or freelancer's paid client application | **Yes: commercial production** A school's paid commercial service | **Yes: commercial production** ## Commercial licences - **Pricing on request.** Contact [the sales team](mailto:sales@onequill.dev) for a licence covering your commercial production use. - **No licence keys and no usage reporting.** OneVir does not report usage to OneQuill. You count your own commercial production deployments and tell us when the number changes. - **Non-production stays free.** Development, testing, staging and evaluation do not need a paid licence when they do not serve live operation. - **Products and hosted services.** Discuss commercial production licensing for a product or paid service with OneQuill. Copying and redistribution rights remain those in the LICENSE. - **Charities, schools and universities.** Non-commercial activities remain free in production. Paid services and other commercial activities need a licence when they run in production. ## Open source after four years Each version converts to the Apache License 2.0 four years after it is first published. Anyone can then use that version for any purpose, including commercial production. Newer versions keep these terms until their own change date. ## Copies, models and other software - **Copies and changes.** You may copy, modify and redistribute OneVir as long as the licence travels with every copy. The same free-use permissions and commercial production requirement apply to modified versions. - **Names and logos.** The licence grants no rights in the OneVir or OneQuill names and logos. - **Model weights and third-party components** keep their own licences, including llama.cpp (MIT). Optional FFmpeg downloads are GPL-licensed and run as a separate program. Cloud providers apply their own terms. - **Your data and outputs are yours.** The licence gives OneQuill no rights in your prompts, documents or generated output. This page summarises the licence. The [licence text](https://onequill.dev/LICENSE) is binding. Last updated 2 October 2026. Questions about licensing? ## Talk to us before you go live. Ask about a commercial licence, pricing or how the terms apply to your intended use. [Ask for a commercial licence →](mailto:sales@onequill.dev?subject=OneVir%20commercial%20licence&body=Hello%20OneQuill%20team%2C%0A%0AOrganisation%3A%0ACommercial%20production%20use%3A%0A) --- Canonical: https://onequill.dev/privacy # Privacy · OneQuill Canonical: https://onequill.dev/privacy Privacy # Your data, left alone. Contact us with questions or requests about your personal data. [Make a privacy request ↗](mailto:contact@onequill.dev?subject=Privacy%20request) ## This website - **No cookies, no analytics, no advertising trackers.** The site sets no cookies and runs no tracking scripts. - **No third-party requests.** Fonts, images and scripts load from this site only. There are no embeds and no contact forms. - **Nothing stored in your browser.** The site keeps no preferences in local storage. Animations follow your system's reduced-motion setting. Like any website, the server that delivers these pages processes technical connection data, such as your IP address and the requested address, in order to deliver them securely. ## Email Choosing an email link opens your own email application, and nothing is sent until you send it. Messages you send are handled through the relevant OneQuill mailbox and used to answer your request. Please avoid sending passwords, API keys or personal information the request does not need. ## OneVir OneVir runs on hardware you control and does not report usage to OneQuill. Its OpenTelemetry export is off by default and never includes prompts, completions, tool payloads or keys. ## Your rights To ask about, correct or delete personal data we hold about you, write to [contact@onequill.dev](mailto:contact@onequill.dev?subject=Privacy%20request). **Preparation copy.** The full controller identity, hosting and email providers, processing purposes, legal bases, retention periods, recipients and applicable rights are being completed for publication. See also our [Impressum](https://onequill.dev/legal) and [contact routes](https://onequill.dev/contact). --- Canonical: https://onequill.dev/resources # AI gateway and inference education centre · OneQuill Canonical: https://onequill.dev/resources OneQuill / Client education centre # Ask better questions. Build with evidence. Understand AI gateways and inference before you choose, deploy or scale. Six practical guides connect real cases to the questions your team should ask. [Start the learning path ↓](https://onequill.dev/resources#learning-path)[Find your provider](https://onequill.dev/resources/ai-gateway-providers) [Illustration: Six connected evaluation topics surround the client evidence record] Know the component. Understand the condition. Ask for evidence. **Research snapshot / 2026-10-07**[55 provider profiles](https://onequill.dev/resources/ai-gateway-providers)[6 learning guides](https://onequill.dev/resources#learning-path)[16 selected cases](https://onequill.dev/resources/ai-gateway-case-studies) Your learning path ## Six perspectives. One informed evaluation. Start with a business question. Expand the technical detail when needed. [01 / SecurityAI gateway security: protect the boundaries that matterUnderstand gateway authentication, administration, outbound requests and tool policy through source-linked cases and practical evaluation questions.Read the guide →](https://onequill.dev/resources/ai-gateway-security) [02 / PrivacyAI gateway privacy: follow data through every layerMap prompts, logs, caches, providers and fallback before comparing data residency, retention and zero-data-retention claims.Read the guide →](https://onequill.dev/resources/ai-gateway-privacy) [03 / Budgets & scalingAI gateway budgets and scaling: measure the whole requestEvaluate token accounting, concurrent spend reservations, retries, regional counters and inference capacity with practical evidence requests.Read the guide →](https://onequill.dev/resources/ai-gateway-budgets-and-scaling) [04 / ReliabilityAI gateway reliability: define failure before adding fallbackUnderstand buffering limits, stream interruption, circuit breakers, failover and recovery without losing data policy or request meaning.Read the guide →](https://onequill.dev/resources/ai-gateway-reliability) [05 / Supply chainAI gateway supply chain: trust the artifact you actually runEvaluate package provenance, pinned images, model assets, release fixes and incident response using carefully scoped vendor evidence.Read the guide →](https://onequill.dev/resources/ai-gateway-supply-chain) [06 / Inference operationsRunning LLM inference in production: security, isolation and capacityEvaluate vLLM, SGLang, Ollama and LM Studio across API exposure, model loading, execution location, isolation and workload-specific capacity.Read the guide →](https://onequill.dev/resources/llm-inference-security-and-operations) [The marketFind the component you use.Search 55 profiles by name, alias, category and deployment mode.→](https://onequill.dev/resources/ai-gateway-providers)[The evidenceLearn from documented cases.16 examples with prerequisites, vendor responses and practical questions.→](https://onequill.dev/resources/ai-gateway-case-studies) A practical next step ## Take the questions into your review. Six A4 pages. Record the provider, version, configuration, evidence, owner and action for each topic. [Download worksheet PDF ↓](https://onequill.dev/assets/downloads/ai-gateway-evaluation-worksheet.pdf)[Markdown worksheet](https://onequill.dev/assets/downloads/ai-gateway-evaluation-worksheet.md) Keep the governance context ## [AI governance rules for gateways and local inference](https://onequill.dev/resources/ai-governance-rules) The original global reference: laws, standards and frameworks with official sources. - [EU AI Act ↗](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai) - [GDPR ↗](https://commission.europa.eu/law/law-topic/data-protection/information-business-and-organisations/application-gdpr_en) - [ISO 27001 ↗](https://www.iso.org/standard/27001) - [ISO 42001 ↗](https://www.iso.org/standard/42001) - [ISO 23894 ↗](https://www.iso.org/standard/77304.html) [Read the reference →](https://onequill.dev/resources/ai-governance-rules) How to read this research ## Evidence with a defined scope. Product coverage is a researched snapshot, not an exhaustive inventory. Category and deployment mode describe the cited offering; editions, contracts and configuration can change the scope. Selected cases illustrate evaluation questions. Advisory counts are not security rankings. **Security advisory**A published security finding, with prerequisites and remediation evidence. **Incident report**A reported event and response, attributed to its source. **Release fix**A correction recorded in release notes; not automatically a CVE. **Documented behaviour**A design choice, default or limit from official documentation. Published by OneQuill, developer of OneVir. This is a documentary review of selected public sources, not an independent vendor security audit. Published 2026-10-07 · Modified 2026-10-07 · Sources reviewed 2026-10-07. Report corrections through [OneQuill support](https://onequill.dev/support), with the page and primary source. [Download research records (JSON)](https://onequill.dev/assets/downloads/ai-gateway-research.json) · [Publishing method and templates](https://onequill.dev/resources/resource-publishing). Evidence & publishing help ## Explore the supporting records. [Evidence recordOneVir control evidence and deployment responsibilitiesReview six OneVir control topics, the evidence behind each finding, and what to verify for your release and deployment.→](https://onequill.dev/resources/onevir-control-evidence)[Publishing guidePublishing the education centre: templates, evidence and regular updatesA reusable publication workflow for source-linked guides, provider profiles and case studies, with drafts, corrections, review dates and search discovery.→](https://onequill.dev/resources/resource-publishing) --- Canonical: https://onequill.dev/support # OneVir support · OneQuill Canonical: https://onequill.dev/support OneVir support # Help starts with a clear description. Tell us where you got stuck and what you expected to happen. The more precisely you describe it, the faster we can answer. [Email OneVir support →](mailto:support@onequill.dev?subject=OneVir%20support%20request&body=Hello%20OneQuill%20support%2C%0A%0AOneVir%20version%20and%20operating%20system%3A%0AIssue%20summary%3A%0ASteps%20to%20reproduce%3A%0AExpected%20result%3A%0AActual%20result%3A%0A) [Ask a product question](mailto:help@onequill.dev?subject=OneVir%20question) ## Support routes Deployments and issues ### Something is not working Installation problems, errors, unexpected behaviour, performance questions about your own hardware. [Get support ↗](mailto:support@onequill.dev?subject=OneVir%20support%20request&body=Hello%20OneQuill%20support%2C%0A%0AOneVir%20version%20and%20operating%20system%3A%0AIssue%20summary%3A%0ASteps%20to%20reproduce%3A%0AExpected%20result%3A%0AActual%20result%3A%0A) Product questions ### How does OneVir work? Questions about capabilities, configuration choices, providers, policy or licensing for a specific use. [Ask a product question ↗](mailto:help@onequill.dev?subject=OneVir%20question) Insights & Guides ## Read the context. Apply it to your workflow. Our resource library brings together source-linked research, explainers and practical help. Start with the global AI governance reference for gateways and local inference. [Explore the library →](https://onequill.dev/resources)[AI governance reference](https://onequill.dev/resources/ai-governance-rules) **01** What to include ## Six details that save a day. Remove API keys, passwords, tokens and personal content from anything you send. OneVir's decision audit already records decisions without prompts, responses or tool arguments. - **Version and platform.** Your OneVir version, operating system and, for engine issues, the GPU and driver. - **The steps.** What you did, in order, until the problem appeared. - **Expected and actual.** What you expected to happen and what happened instead. - **The message.** The exact error text or a short log excerpt around it. - **The route.** Whether the request ran locally or through a provider, and the model or alias it used. - **Configuration.** The relevant part of your configuration, with every secret removed. **02** Look first ## Answers already inside OneVir. Many questions are answered by records OneVir keeps for you. ### Activity and recent decisions See where each request ran, which rule or alias chose the route, and whether it was served from cache. ### Performance CPU, memory, GPU, disk and network, with OneVir's own share shown separately, sampled every second. ### Built-in API reference Every OneVir instance serves its own typed API reference and OpenAPI document for the exact version you run. Before installing or upgrading ## Back up first. Then upgrade. Follow the installation instructions and requirements published with your release. Keep a backup of your configuration and stored state before an upgrade that changes either. --- Canonical: https://onequill.dev/resources/agentgateway-namespace-isolation # agentgateway: a patched release can still need a policy setting A route that is syntactically valid can still cross an ownership boundary. Credential-backed resources require authorisation for every route and policy reference, not just for the namespace that supplies the frontend. ## What deployment does this concern? **Component and scope:** Cross-namespace backend references authored by Kubernetes namespace administrators. **Affected versions / scope:** < 1.3.0; configuration also matters after upgrade. **Vendor remediation:** 1.3.0 plus AGW_BACKEND_REF_GRANT_MODE=route-and-policy A tenant namespace administrator can author Gateway, HTTPRoute or Policy resources that refer to credential-bearing backends in another namespace. The vendor describes a multi-tenant control-plane issue and no remotely exploitable surface. ## Vendor response and practical action The vendor requires the patched release together with route-and-policy backend-reference grant enforcement. The default route mode alone is insufficient for the described policy-reference condition. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Patch status is a release-and-configuration claim. Record the environment setting, accepted resources and cross-namespace authorisation evidence beside the version number. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Who may create routes and policies in each namespace? - Are route and policy backend references both subject to grants? - Is AGW_BACKEND_REF_GRANT_MODE set to route-and-policy? - Can a tenant use a backend credential owned by another namespace? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway security: protect the boundaries that matter](/resources/ai-gateway-security) for the wider evaluation context. ### Technical detail: evidence and identifier limits An internet caller without control-plane permissions is outside the described prerequisite. Do not present this case as an unauthenticated network attack. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2026-06-29**. These dates do not change merely because this article is rebuilt. ## Primary sources - [GHSA-jwm2-83f3-52xc · vendor advisory](https://github.com/agentgateway/agentgateway/security/advisories/GHSA-jwm2-83f3-52xc) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/ai-gateway-budgets-and-scaling # AI gateway budgets and scaling: measure the whole request A gateway can enforce requests per minute while still overspending a monthly budget. The request count, token estimate, provider bill and usable inference capacity measure different things. Establish the units and boundaries before choosing limits. Begin with a budget policy that names the tenant, provider, model, billing period and currency. Define whether the limit is an estimate at admission, an actual cost after completion or both. Decide what happens when usage is missing rather than silently treating it as free. ## Find the counter's sharing boundary [Azure API Management's token-limit policy](https://learn.microsoft.com/en-us/azure/api-management/llm-token-limit-policy) describes independent counters across gateway, region and workspace boundaries. [WSO2's backend throttling documentation](https://apim.docs.wso2.com/en/latest/api-design-manage/design/rate-limiting/protect-backend-services/) distinguishes local per-node limits from distributed enforcement. [agentgateway's rate-limit documentation](https://agentgateway.dev/docs/standalone/latest/documentation/configuration/resiliency/rate-limits/) distinguishes local in-memory state from a remote shared service. [Traefik Hub's token middleware](https://doc.traefik.io/traefik-hub/ai-gateway/middlewares/token-rate-limit) documents Redis-backed sharing. None of these descriptions alone establishes that your entire multi-region spend is capped. Record the authoritative counter store, update consistency, restart behaviour and failure policy. Ask how admitted work is reconciled when a node exits before settlement. ## Account for concurrent work and arithmetic The [New API quota case](/resources/new-api-quota-billing-overflow) concerns integer overflow in release-candidate quota settlement. It illustrates why a billing boundary includes numeric ranges, sign handling and state transitions. Before dispatch, reserve a defensible bound based on supported input and output limits. After completion, settle actual usage and release unused reservation. Ask how repeated events, cancellation and concurrent admissions are handled. A balance check followed by a separate update can allow several requests to consume the same remaining allowance. ## Streaming changes how usage arrives The [Kong Gemini case](/resources/kong-gemini-streaming-token-accounting) is a release fix for running usage metadata, not a CVE. Cumulative totals must not be added as if each chunk were independent usage. Missing metadata is a different state from a reported zero. A useful demonstration includes a complete stream, an interrupted stream, a retry and a fallback. Compare gateway records with provider-side evidence. Include reasoning, cached input and other billable units when supported by the provider; do not assume every provider uses identical fields. ### Technical detail: capacity and billable attempts Limit logical requests and upstream attempts separately. A single client request may cause more than one billed dispatch. A client disconnect does not necessarily stop upstream work immediately. Define how reservations cover each permitted attempt and when they expire. For self-hosted inference, use the [inference operations guide](/resources/llm-inference-security-and-operations). Record model revision, quantisation, context length, input/output distributions, hardware, parallelism and accepted latency. Requests per second from a different workload are not a useful budget guarantee. ## Questions for your evaluation Complete the budgets page of the [worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf). Keep the counter topology, cost model, concurrency result and provider reconciliation together. Assign ownership for price changes and missing-usage investigations. ## Applying this to OneVir [OneVir documents provider budgets and application limits](https://onevir.onequill.dev/#capabilities). The [implementation evidence record](/resources/onevir-control-evidence#budgets) identifies a focused concurrent-reservation test: eight competing attempts are presented with a two-attempt allowance, and the test asserts two admissions and the matching reserved amount. That reviewed source supports a specific reservation behaviour. It is not a test run performed for this publication or proof that every provider price, streaming protocol and distributed deployment is correct. Request evidence for the configured route and settlement path. --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/ai-gateway-privacy # AI gateway privacy: follow data through every layer A statement such as “we do not store prompts” leaves several questions unanswered. It may describe the gateway's database while excluding its response cache, monitoring service, backups or upstream provider. Write down which component makes the claim and which data it covers. Start with one representative request containing the kinds of content your clients actually use. Trace its prompt, attachments, tool arguments, retrieved context, response and operational metadata. Mark processing location, retention owner, access rights and deletion rule for each copy. ## Distinguish routing location from inference location A gateway in Europe can still route inference elsewhere. [Opper's security overview](https://opper.ai/security-overview), updated 5 October 2026, explicitly says upstream inference is not EEA-restricted by default. Its tracing and backup retention are separate considerations. [Requesty's privacy policy](https://www.requesty.ai/privacy) also needs to be read at both router and upstream-provider scope. For hosted routers such as [OpenRouter](https://openrouter.ai/docs/guides/features/guardrails/overview), evaluate eligible endpoints, provider allowlists and fallback settings together. An approved first choice does not establish that an automatically selected fallback is acceptable. Retain a route configuration or vendor commitment covering the complete permitted set. ## Logs and caches are separate storage choices [Leanroute's privacy policy](https://leanroute.dev/privacy) distinguishes prompt database storage from response and semantic caches, which default to a 15-minute lifetime. A no-persistence option changes that behaviour. This is a documented design choice to evaluate, not a finding that the service is unsafe. [Cloudflare documents gateway logging](https://developers.cloudflare.com/ai-gateway/observability/logging/) with different handling for new customers from 24 September 2026 and earlier accounts. Identify which generation applies to your account before copying a retention limit from a comparison chart. [LangDB's retention documentation](https://docs.langdb.ai/enterprise/resources/configuring-data-retention/) describes ClickHouse TTL cleanup through asynchronous merges. A retention deadline and physical deletion can therefore be different events. Ask how exports and backups are covered. ## Read zero-retention claims at the correct layer [Vercel's security description](https://vercel.com/i/secure-ai-gateway) separates gateway handling from provider ZDR arrangements and eligibility. [ZenMux's Data Services controls](https://zenmux.ai/docs/guide/advanced/data-services.html) change its own logging and related features; upstream processing still needs review. [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers/en/security) also distinguishes its request handling and metadata from provider policies. Record whether a commitment covers prompt content, output, files, embeddings, tool calls, abuse monitoring and metadata. Ask whether enabling debugging, tracing, insurance or support access changes the commitment. Contractual terms and effective settings both belong in the evidence. ### Technical detail: deletion and tenant boundaries A retention change can stop new recording while leaving earlier copies to expire. Request the effective date, deletion job behaviour, backup lifecycle and export inventory. An organisation-wide setting may differ from a project-level setting. For local inference, distinguish intentional prefix-cache sharing from accidental memory disclosure. [The vLLM GGUF case](/resources/vllm-gguf-gpu-memory-isolation) concerns kernel output memory. Cache salting addresses a different mechanism. Neither can be represented by a single generic “private cache” tick box. ## Questions for your evaluation Use the privacy page of the [worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf). Attach a data-flow diagram, provider and subprocessor list, contract scope, retention settings and a deletion example. If evidence is missing, record that gap and an owner instead of inferring a favourable answer. ## Applying this to OneVir [OneVir describes local and configured-provider execution paths](https://onevir.onequill.dev/#path). The [implementation evidence record](/resources/onevir-control-evidence#privacy) identifies retention-policy initialisation and history API documentation: invalid saved retention policy prevents service initialisation, and route retention can prohibit saved chat history. Those observations cover named paths, not every data copy. Evaluate the chosen provider route, local history, files, observability exports and backups. Local execution can reduce upstream disclosure, but the operator still owns storage, access and deletion in the surrounding deployment. --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/ai-gateway-reliability # AI gateway reliability: define failure before adding fallback Adding another provider does not automatically make a service reliable. A fallback may exceed a budget, disclose data to a different region or repeat a tool action. Define an acceptable result and failure response before selecting the routing policy. Build a small failure matrix: upstream timeout, overload, authentication failure, guardrail outage, client disconnect, gateway restart and inference-worker loss. For each event, record whether work was dispatched, whether partial output was returned and whether retry is safe. ## Put limits before resource consumption The [Envoy MCP request-limit case](/resources/envoy-ai-gateway-mcp-request-limits) concerns reading a full body before limits were applied. A rejection after allocation can arrive too late to protect component memory. Choose body-size, decoded-content, concurrency and deadline limits together. Include compression, base64 expansion, image dimensions and video frames when those features are available. Measure which component retains the body and how cancellation releases it. The [vLLM input-validation case](/resources/vllm-multimodal-input-validation) shows why input features need separate bounds. ## Make incomplete output visible [Apache APISIX's AI Proxy documentation](https://apisix.apache.org/docs/apisix/plugins/ai-proxy/) describes timeout behaviour that may close a stream without a DONE event. A network connection ending is therefore not universally proof of successful completion. The caller needs an explicit completion signal, cancellation state and partial-output policy. Decide whether a partially generated answer can be used, must be marked incomplete or must be discarded. Never replay an action-bearing workflow blindly after its visible response is lost. ## Keep fallback inside the approved boundary List every permitted fallback and the conditions for choosing it. Reapply identity, model permissions, data geography, spend admission and tool restrictions immediately before dispatch. Record the route decision so support staff can distinguish “provider failed” from “policy refused.” A circuit breaker can reduce repeated calls to an unhealthy provider. Its thresholds, error classification and recovery trial still need review. A provider health check does not necessarily exercise the same model, quota or tool path as a real request. ### Technical detail: failure domains and recovery Gateway replicas, rate-limit storage, authentication services, policy services and inference workers can fail independently. Write down what state each holds and what happens when it restarts. A local counter may reset, an open circuit may be forgotten and a stream may have no surviving owner. For distributed inference, [vLLM documents Ray and multiprocessing options](https://docs.vllm.ai/en/stable/serving/parallelism_scaling/). [Ray Serve's fault-tolerance guidance](https://docs.ray.io/en/latest/serve/production-guide/fault-tolerance.html) describes recovery mechanisms with the appropriate KubeRay deployment. Avoid assuming either automatic recovery or a mandatory whole-cluster manual restart without inspecting the setup. ## Questions for your evaluation Use the reliability page of the [worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf). Capture rejection behaviour, completion semantics, permitted replay conditions and observed recovery. Choose tests from the failure matrix that directly affect your workflow. ## Applying this to OneVir [OneVir documents ordered fallbacks and provider circuit breakers](https://onevir.onequill.dev/#capabilities). The [implementation evidence record](/resources/onevir-control-evidence#reliability) identifies the provider-health state machine and routing check: down providers are skipped in fallback chains, with cooldown and trial behaviour. This source review does not demonstrate the recovery time of your installation. Ask for a controlled upstream-failure and interrupted-stream demonstration with the configured policy, and assess separately the availability of local inference workers and external dependencies. --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/ai-gateway-security # AI gateway security: protect the boundaries that matter An AI gateway often holds more authority than the application calling it. It may possess provider credentials, reach internal networks and invoke tools. The useful procurement question is whether an ordinary caller can use any of that authority outside its approved purpose. Start with an interface inventory. List the inference API, dashboard, management API, tool registration, model configuration and health endpoints. Record which identity may reach each one, which permissions it carries and where those permissions are checked. Network placement and application authorisation support each other; neither replaces the other. ## Separate application access from administration The [Bifrost management-API case](/resources/bifrost-management-api-boundaries) illustrates why tool registration and native-plugin loading belong to an administrative boundary. Its advisories depend on reachable management APIs with authentication disabled, and the remote-plugin outcome differs between dynamic builds and static Docker images. The [GitLab template case](/resources/gitlab-ai-gateway-template-sandbox) has a different prerequisite: an authenticated Duo Agent Platform user submits a crafted flow template. Treat workflow authoring as a capability with execution consequences. Ask which secrets, filesystem paths and outbound services the runtime can access. A useful demonstration starts with an ordinary application's credential. It should show successful approved inference and rejection of route changes, credential reads, tool installation and administrative exports. Keep the identity and endpoint list with the result. ## Follow the request to its real destination A provider label is not an outbound firewall. Headers, base URLs, redirects, DNS and fallback can change where the request goes. In the [Portkey custom-host case](/resources/portkey-custom-host-ssrf), the scope is the open-source gateway and custom-host routing; it does not establish exposure in every managed service. Document an approved destination set and the point where it is enforced. Decide what happens when a host resolves to a private address, a redirect changes the destination or the approved provider becomes unavailable. The fallback should satisfy the same disclosure and network policy as the original route. ## Preserve the meaning of a tool request The [Envoy MCP parsing case](/resources/envoy-ai-gateway-mcp-message-smuggling) concerns different interpretations of the same protocol message. A policy decision is useful only when it describes the action the destination will actually execute. Ask how ambiguous fields, unsupported content parts and malformed tool arguments are handled. A successful demonstration should connect the parsed request, policy result and forwarded message through a shared request identifier. Rejection behaviour matters alongside compatibility with valid clients. ### Technical detail: release and configuration evidence The [agentgateway namespace case](/resources/agentgateway-namespace-isolation) requires both the patched release and `AGW_BACKEND_REF_GRANT_MODE=route-and-policy`. The vendor describes control-plane authoring permissions, not an unauthenticated internet attack. Record the installed version, effective environment, route and policy reference grants, and the administrative role used for the check. An endpoint matrix should distinguish unauthenticated rejection, authenticated-but-forbidden access and permitted operations. Avoid publishing credentials or exploit payloads as evidence. Preserve enough configuration context for an operator to repeat the relevant check safely. ## Questions for your evaluation Use the security page of the [evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf). Assign an owner to each boundary and attach the installed version, permission matrix, outbound policy and focused regression result. A feature name or certification logo does not establish the behaviour of your particular route. ## Applying this to OneVir [OneVir documents application keys, model grants and policy controls](https://onevir.onequill.dev/#capabilities). Our [implementation evidence record](/resources/onevir-control-evidence#security) identifies the credential-normalisation test in `src/api/mod.rs`: conflicting credentials are rejected and supported credential forms are normalised. That is evidence for a specific parser behaviour, not a complete authentication audit. For an evaluation, request the installed release, enabled application-key scopes and an endpoint access demonstration. A gateway credential parser does not establish the safety of a separately deployed inference engine, plugin or tool runtime. --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/ai-gateway-supply-chain # AI gateway supply chain: trust the artifact you actually run An AI deployment contains more than its gateway binary. Python packages, container layers, native plugins, model configurations, tokenisers and workflow templates can all influence behaviour. Maintain an inventory of the things the service executes or imports. Record the running artifact's digest, provenance, dependency lock and build options. Keep the model revision and configuration hash alongside it. The useful inventory is a deployment record that supports an exposure decision when a supplier publishes an advisory. ## Match the incident to the distribution channel The [LiteLLM March package incident](/resources/litellm-march-2026-package-incident) concerns two PyPI versions in a reported time window. The vendor states its official pinned Proxy Docker distribution was unaffected. Those are different artifact claims. Determine what was installed and executed, not just whether the organisation uses the project. A container that performs unpinned installation at startup can have a different dependency set from the image used in testing. Preserve installation times and artifact evidence during investigation. ## Model and plugin inputs can change execution The [vLLM model-loading case](/resources/vllm-model-loading-python-optimisation) requires malicious model configuration and Python optimisation that removes assertions. It is not a generic prompt attack. A model approval process should account for configuration and code-loading behaviour, not only the weight file's reputation. The [Bifrost management case](/resources/bifrost-management-api-boundaries) distinguishes dynamically linked plugin execution from the static Docker build's fetch behaviour. Build settings and native-code permissions belong in the deployment inventory. ## Upgrade and response are different work A patched artifact repairs a known future path. Incident response considers what may already have happened. Identify which credentials, files and network services were accessible, then decide containment and rotation based on that exposure. Keep release-fix evidence separate from incidents and security advisories. The [Kong accounting case](/resources/kong-gemini-streaming-token-accounting) is a release-note correction; it should inform a regression check without being relabelled as a CVE. A rollback plan also needs a trustworthy artifact. Reverting to an older vulnerable build may restore availability while reintroducing the original exposure. Document the intended safe version and the temporary controls that accompany restoration. ### Technical detail: provenance evidence Request a signed or otherwise verifiable artifact identity, build inputs, dependency inventory and support policy where the vendor supplies them. Confirm what a signature attests: a publisher identity, build process or particular source revision. A valid signature does not establish that code has no defects. For model assets, record the repository revision, files fetched, approval identity and optional remote-code settings. For plugins, record the source, hash, installation permission and sandbox or process boundary. Do not assume that all non-weight files are inert data. ## Questions for your evaluation Use the supply-chain page of the [worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf). Attach an artifact inventory and a short response exercise. Include the upstream notices you rely on, their review dates and any unresolved mapping differences. ## Applying this to OneVir [OneVir's public description](https://onevir.onequill.dev/#capabilities) covers local inference and configured provider routing. The [implementation evidence record](/resources/onevir-control-evidence#supply-chain) identifies a version-locked llama.cpp binding in the reviewed Cargo manifest and records the model-format boundary. A pinned dependency is one supply-chain control, not complete provenance or independent certification. Request the release artifact, dependency and model inventory, update procedure and response ownership for the actual installation. External providers and separately installed backends retain their own supply-chain responsibilities. --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/ai-governance-rules # AI Governance Rules for OneVir Gateway and Inference **Research date: 7 October 2026.** This document reviews the original six instruments and expands them into a global reference for AI gateway, local inference, and connected application controls. Sources are legislation, regulators, government agencies, standards publishers, and the organisations that publish the security frameworks. The distinction between a published obligation and its application to OneVir is explicit: - **Binding law:** enforceable where the jurisdiction, activity, actor, and commencement conditions apply. - **Supervisory expectation:** regulator guidance for a specified regulated sector, such as OSFI E-23. - **Standard or voluntary guidance:** an assurance or risk-management reference; it does not become a worldwide statutory duty merely because OneVir uses AI. - **How OneVir supports this:** an engineering interpretation linking a requirement to relevant product controls. The adjacent responsibility column explains what the operator, application or inference backend must still decide and verify. Sources do not prescribe OneVir configuration names or numerical thresholds. The tables explain how gateway controls can support governance and where accountability sits. They are not a list of discovered OneVir defects. [OneVir's public capabilities](https://onevir.onequill.dev/#capabilities) and our [control evidence record](/resources/onevir-control-evidence) describe named implementation paths and their limits; verify the installed version and enabled settings before treating a control as established. This requirements reference does not certify compliance. Coverage includes major relevant regimes in Europe, North America, Asia, Latin America, Africa, the Middle East, and Australia; it is not an inventory of every country's laws or every US state law. ## 1. Review of the original six instruments The six are useful references, but they have different legal status and scope. They are not six interchangeable certifications or six universal gateway requirements. | Instrument | Verified status and scope | How OneVir supports the requirement | | --- | --- | --- | | **ISO/IEC 27001:2022** | Requirements for an organisation's information security management system, including risk assessment and treatment. ISO describes protection of confidentiality, integrity, and availability. It is an international standard, not an AI-specific statute. [ISO publication](https://www.iso.org/standard/27001). | Support the operator's security controls and evidence. A gateway feature list alone does not establish conformity of an organisation's management system. | | **ISO/IEC 42001:2023** | Requirements for establishing, implementing, maintaining, and continually improving an AI management system. Applies to organisations providing or using AI-based products or services. [ISO publication](https://www.iso.org/standard/42001). | Use approved model/provider choices, assigned ownership and configuration-change evidence within the client's AI management system. Governance, management review and continual improvement remain organisational responsibilities. | | **NIST AI RMF 1.0; NIST AI 600-1** | The AI RMF is voluntary and uses **Govern, Map, Measure, Manage**. The July 2024 Generative AI Profile addresses generative-AI risks including confabulation, harmful bias, privacy, information security, and value-chain integration. [NIST AI RMF](https://www.nist.gov/itl/ai-risk-management-framework), [final Generative AI Profile](https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf). | Define application risks, select relevant gateway controls and retain evaluation and monitoring evidence. The operator owns risk acceptance and review; using the framework does not certify a product. | | **EU AI Act, Regulation (EU) 2024/1689, as amended** | Binding, with distinct prohibited-use, high-risk-system, transparency, and general-purpose AI model obligations. The Commission confirms the AI Omnibus entered into force on **27 July 2026**; Annex III high-risk duties apply from **2 December 2027**, and relevant Annex I product duties from **2 August 2028**. [Commission amendment notice](https://digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force). | Connect the application's approved use and actor-specific duties to access, data handling, evidence and output controls. Section 4 explains the milestones and preparation work; extensions for high-risk systems do not postpone every other AI obligation. | | **California SB 53 — Transparency in Frontier Artificial Intelligence Act** | Enacted in September 2025 and effective **1 January 2026**. Addresses frontier AI developers, safety-framework disclosure, certain critical safety incidents, and whistleblower protections. Duties differ by the statutory developer category. [Governor's signing announcement](https://www.gov.ca.gov/2025/09/29/governor-newsom-signs-sb-53-advancing-californias-world-leading-artificial-intelligence-industry/), [Senate confirmation of effective date](https://sd38.senate.ca.gov/news/summer-recess-activities-and-more-news-senator-blakespear), [2026 government summary](https://www.gov.ca.gov/2026/09/09/governor-newsom-signs-first-in-the-nation-ai-safeguards-to-protect-californians-calls-on-the-federal-government-to-do-its-part/). | Keep model, supplier and version evidence available for the client's assessment and incident process. Routing to a frontier model alone does not make the operator a covered frontier developer; the covered developer owns the applicable public disclosures and reporting. | | **OSFI E-23 — Model Risk Management (2027)** | Final Canadian supervisory guideline published **11 September 2025**, effective **1 May 2027**, for federally regulated financial institutions, including relevant foreign branches. Covers AI/ML and third-party models on a proportional, risk-based basis. [Final guideline](https://www.osfi-bsif.gc.ca/en/guidance/guidance-library/guideline-e-23-model-risk-management-2027). | Support financial clients' model inventories, assigned ownership, independent review, approval, change control, monitoring, and decommissioning evidence. It is not a requirement imposed on every Canadian inference server. | **Evidence limitation for ISO:** the public descriptions establish purpose and scope. The full standards were not accessed, so this document does not claim a verified clause-by-clause ISO control mapping. ## 2. Establish applicability before selecting an enforcement policy OneVir's legal role cannot be determined just from “local inference”, “AI gateway”, or the model name. Record the deployment facts that the cited instruments use: | Fact to establish | Why it changes the applicable duties | | --- | --- | | **Market, establishment, and where outputs are used** | EU AI Act Article 2 covers specified EU actors and certain non-EU providers/deployers whose system output is used in the Union; it also contains exclusions. Personal non-professional use is treated differently from professional deployment. [Article 2](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-2). | | **Role in the value chain** | Model providers, AI-system providers, deployers, and component suppliers have different duties. Rebranding or substantially modifying a high-risk system, or changing its purpose so that it becomes high-risk, can change the responsible actor. [Article 25](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-25), [GPAI provider duties, Article 53](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53). | | **Purpose and effect of the application** | Recruitment, credit, education, biometrics, and health-related uses can trigger specific duties. A generic chat endpoint is insufficient evidence that every request is a legally high-risk application. [Commission risk categories](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai), [Colorado consequential-decision categories](https://www.leg.colorado.gov/bills/SB26-189). | | **Public service versus internal use** | China's generative-AI measures cover services offered to the public in mainland China and expressly exclude specified development/use that does not provide such a public service. The provider definition includes programmable interfaces. [CAC measures, Articles 2 and 22](https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm). | | **Data types, children, and regulated clients** | Data protection and sector rules have independent scope. Establish controller/processor relationships, recipients, and actual data flows. [GDPR application](https://commission.europa.eu/law/law-topic/data-protection/information-business-and-organisations/application-gdpr_en), [HHS cloud guidance](https://www.hhs.gov/hipaa/for-professionals/special-topics/health-information-technology/cloud-computing/index.html), [FTC COPPA guidance](https://www.ftc.gov/business-guidance/resources/complying-coppa-frequently-asked-questions). | **OneVir application:** bind the operator-approved purpose and data restrictions to authenticated applications or principals. A model's classification of a prompt, or a client-supplied label alone, is not proof of legal applicability. This is an implementation consequence of needing deployment context and enforceable authorisation, not a new legal classification rule. [NIST AI RMF](https://www.nist.gov/itl/ai-risk-management-framework), [OWASP object-level authorisation](https://api-security.owasp.org/editions/2023/en/0xa1-broken-object-level-authorization/). ## 3. Additional global laws and regulatory regimes The following are deployment-specific additions to the original six. “In force” does not mean “applicable to every OneVir installation”. Future commencement dates and sector limitations matter. | Jurisdiction / instrument | Published duties and current status | OneVir support and deployment responsibilities | | --- | --- | --- | | **EU / EEA — GDPR** | In-scope processing requires lawfulness, fairness, transparency, purpose limitation, minimisation, accuracy, storage limitation, and security. International transfers need an applicable mechanism. [Commission principles](https://commission.europa.eu/law/law-topic/data-protection/reform/rules-business-and-organisations/principles-gdpr/overview-principles/what-data-can-we-process-and-under-which-conditions_en), [international transfers](https://commission.europa.eu/law/law-topic/data-protection/international-dimension-data-protection/rules-international-data-transfers_en). | Apply restrictions to prompts, responses, files, embeddings, caches, and telemetry containing personal data. GDPR does not universally require all AI requests to stay inside the EU. | | **United Kingdom — UK GDPR / DPA 2018, amended by DUAA 2025** | The ICO confirms all DUAA data-protection provisions were in force by its **19 June 2026** update. Significant automated decisions can use broader lawful bases with safeguards; special-category data remains more restricted. [ICO current explanation](https://ico.org.uk/about-the-ico/what-we-do/legislation-we-cover/data-use-and-access-act-2025/the-data-use-and-access-act-2025-what-does-it-mean-for-organisations/). | Support privacy and decision safeguards, but do not copy EU Article 22 unchanged into a UK policy. The application owns the consequential decision and user process. | | **United States — consumer protection and children's privacy** | FTC enforcement includes deceptive AI capability claims. COPPA covers certain services collecting personal information from children under 13; its amended rule adds retention restrictions and parental-consent requirements for certain third-party disclosures. [FTC AI enforcement](https://www.ftc.gov/industry/technology/artificial-intelligence), [amended-rule announcement](https://www.ftc.gov/news-events/news/press-releases/2025/01/ftc-finalizes-changes-childrens-privacy-rule-limiting-companies-ability-monetize-kids-data), [COPPA scope](https://www.ftc.gov/business-guidance/resources/complying-coppa-frequently-asked-questions). | Support claims with evidence. A child-facing client may need recipient restrictions, deletion, and age/parental-consent processes. COPPA is not an age-verification mandate for every internal inference API. | | **California — CCPA/CPRA and ADMT regulations** | Regulations effective **1 January 2026** include risk assessments, cybersecurity audits, and automated decisionmaking technology rules. Section 7200 sets ADMT compliance for covered significant decisions at **1 January 2027**. Scope and exemptions are specified. [CalPrivacy regulations](https://cppa.ca.gov/regulations/), [operative text, section 7200](https://cppa.ca.gov/regulations/pdf/ccpa_statute_eff_20260101.pdf). | Supply evidence for covered clients' privacy rights and risk assessments. Logs do not replace notices, access/opt-out handling, or applicable exceptions. These duties are distinct from SB 53. | | **Colorado — SB26-189** | Enacted **14 May 2026**, repeals and reenacts the earlier SB24-205 provisions. Covered ADMT duties begin **1 January 2027**: developer documentation and change notification; deployer notices; data rights and meaningful human review after adverse decisions; compliance records for at least **three years**. [Enacted summary](https://www.leg.colorado.gov/bills/SB26-189). | Preserve version/change evidence for covered decision applications. Do not describe the old SB24-205 deadline or impact-assessment regime as the current rule without accounting for its replacement. | | **United States — HIPAA** | Covers relevant covered entities and business associates handling protected health information. HHS explains that maintaining encrypted ePHI can create business-associate status even without the decryption key. Agreements and safeguards remain necessary. [HHS cloud guidance](https://www.hhs.gov/hipaa/for-professionals/special-topics/health-information-technology/cloud-computing/index.html), [contract requirements](https://www.hhs.gov/hipaa/for-professionals/covered-entities/sample-business-associate-agreement-provisions/index.html). | Restrict PHI to permitted recipients and support safeguards, incidents, and return/destruction duties. An inference product is not automatically “HIPAA compliant”; encryption alone does not establish compliance. | | **EU financial services — DORA** | Applicable from **17 January 2025** to financial entities in scope. Covers ICT risk management, incident reporting, testing, and third-party risk; entities maintain ICT contractual-arrangement registers. [EBA application notice](https://eba.europa.eu/activities/direct-supervision-and-oversight/digital-operational-resilience-act/preparation-dora-application), [EBA scope explanation](https://www.eba.europa.eu/publications-and-media/press-releases/eba-amends-its-guidelines-ict-and-security-risk-management-measures-context-dora-application). | Financial clients may require operational/supplier evidence. Running OneVir does not itself make its operator a regulated financial entity or designated critical ICT provider. | | **China — Interim Measures for Generative AI Services (2023)** | Effective **15 August 2023** for covered public services. Includes lawful training inputs, privacy, unlawful-content handling, complaints, and security assessment/algorithm filing for services with public-opinion or social-mobilisation attributes. [Final measures, Articles 2, 7, 11, 14, 15, 17, 22](https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm). | Covered public/API services need specific content and operational controls. Filing/assessment are conditional, not blanket registration for private local models. Other data/security laws require separate scoping. | | **China — generated/synthetic-content labelling measures (2025)** | Effective **1 September 2025**: visible and metadata labelling duties for covered services and a supporting mandatory national standard. Article 9 conditionally permits output without visible labels, with recorded responsibilities and relevant logs for at least six months. Article 10 prohibits malicious label tampering/removal. [Final measures](https://www.cac.gov.cn/2025-03/14/c_1743654684782215.htm), [accompanying standard](https://www.cac.gov.cn/2025-03/14/c_1743654685896173.htm). | Preserve required labels through transformations/exports. Interface disclosure and file metadata have separate roles; a generic API header does not prove satisfaction of all labelling duties. | | **South Korea — AI Basic Act** | In force **22 January 2026**. Article 31 addresses prior notice for high-impact/generative AI products/services and generated-content labelling. MSIT differentiates providers from users and announced at least a one-year investigations/penalties grace period under this provision. [MSIT transparency guidance](https://www.msit.go.kr/eng/bbs/view.do?bbsSeqNo=42&nttSeqNo=1215). | Covered providers need notices/labels. Merely using AI tools for work or creative activities does not automatically trigger the provider duty. Enforcement grace is distinct from the law's commencement. | | **India — DPDP Act 2023 / Rules 2025** | **Phased commencement**: institutional provisions first, specified provisions one year after Gazette publication, and core processing/rights provisions **18 months after publication**. The core provisions remain future at this research date. [Official G.S.R. 843(E)](https://www.meity.gov.in/static/uploads/2025/11/c56ceae6c383460ca69577428d36828b.pdf), [Rules announcement](https://www.pib.gov.in/PressReleasePage.aspx?PRID=2190014). | Prepare for covered clients' notice, lawful processing, security, rights, and processor arrangements. Do not label all processing duties already enforceable in October 2026. Follow the Gazette schedule. | | **Brazil — LGPD / ANPD transfer regulation** | Privacy and automated-decision rights; ANPD Resolution **19/2024** regulates international transfers and contractual mechanisms. The separate AI bill **PL2338/2023** remains pending in the Chamber's status record. [ANPD regulation](https://www.gov.br/anpd/pt-br/acesso-a-informacao/institucional/atos-normativos/regulamentacoes_anpd/resolucao-cd-anpd-no-19-de-23-de-agosto-de-2024), [bill status](https://www.camara.leg.br/proposicoesWeb/fichadetramitacao?idProposicao=2487262). | Enforce authorised destinations and support applicable access/deletion/review processes. Do not present the AI bill as enacted duties or convert the LGPD review right into mandatory human review of every inference. | | **South Africa — POPIA** | Section **71** restricts specified solely automated decisions with legal/substantial effects, subject to exceptions/safeguards; section **72** conditions transfers outside South Africa. [Official Act](https://www.gov.za/sites/default/files/gcis_document/201409/3706726-11act4of2013protectionofpersonalinforcorrect.pdf). | Support permissible transfers and decision-process evidence. The responsible party/application must implement the statutory safeguards; routing metadata is insufficient. | | **Saudi Arabia — PDPL and implementing/transfer regulations** | SDAIA's controller/processor guide covers lawful processing, rights, security, accountability, and cross-border rules. Transfer rules are distinct from AI ethics principles. [Controller/processor guide](https://dgp.sdaia.gov.sa/wps/portal/pdp/knowledgecenter/details/PDPLCP), [laws and regulations](https://sdaia.gov.sa/en/SDAIA/about/Pages/RegulationsAndPolicies.aspx). | Apply covered clients' privacy/transfer restrictions. Do not infer blanket Saudi-only residency or equate an ethics document with all PDPL duties. | | **Canada — privacy obligations and regulator AI guidance** | Regulators' GenAI principles address legal authority, appropriate purposes, necessity/proportionality, openness, accountability, safeguards, accuracy, and individual access. They explain privacy expectations, not a new omnibus AI statute. [Joint regulator principles](https://www.priv.gc.ca/en/privacy-topics/technology/artificial-intelligence/gd_principles_ai/). | Support the client's applicable federal/provincial privacy law. Select that law and sector separately; E-23 adds financial-supervision expectations. | **California transparency update:** the Governor announced further AI Transparency Act amendments, including AB2713 and SB1000, on **30 September 2026**. Their consolidated operative provisions were not verified in this research; no exact additional threshold, gateway watermarking mandate, or effective date is asserted. A California public generative-media deployment needs that specific follow-up. [Official enactment announcement](https://www.gov.ca.gov/2026/09/30/californias-nation-leading-ai-framework-just-got-stronger-governor-newsom-signs-more-first-in-the-nation-worker-protections-and-more/). ## 4. EU AI Act: keep the different obligations separate ### 4.1 Current timeline | Obligation category | Application milestone | What clients should prepare | | --- | --- | --- | | Original prohibited practices | **2 February 2025**. Definitions and exceptions determine scope. | Review the application's intended use and identify prohibited practices with the responsible owner. Configure approved restrictions; a general prompt-safety score cannot establish legal classification. | | GPAI model-provider rules | **2 August 2025**, with transition arrangements for earlier models. | Establish whether you provide a GPAI model or consume one. Obtain the relevant supplier documentation and assign any model-provider duties to the correct actor. | | Article 50 transparency | **2 August 2026**. Different duties cover interaction notices, synthetic-output marking and deployer disclosure. | Decide which user notices and output markings your application must present. Verify that routing, formatting and streaming preserve the required information. | | New specified-content prohibition and certain Article 50(2) transitions | **2 December 2026**, as set out in the current Commission timeline. | Assess the newly prohibited system purposes and the transitional marking deadline for eligible systems placed on the market before 2 August 2026. Document applicability and the required change. | | Annex III high-risk system rules | **2 December 2027**. Applies to the high-risk systems in scope. | Classify the intended use and build the risk-management, documentation, logging and human-oversight process before deployment. Gateway controls support selected technical measures within that process. | | Relevant Annex I high-risk product rules | **2 August 2028**. Concerns high-risk AI embedded in the regulated products in scope. | Coordinate AI evidence with the product's applicable conformity and safety process. Record who owns the integrated system assessment and supplier evidence. | Sources: [Commission AI Act overview](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai), [current implementation timeline](https://ai-act-service-desk.ec.europa.eu/en/ai-act/eu-ai-act-implementation-timeline), [Omnibus notice](https://digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force). Check transitional rules and exclusions for the actual deployment. The Omnibus also changed the earlier AI-literacy provision; do not carry forward old summaries unchanged. ### 4.2 Specific duties influencing gateway/inference design | Published requirement | How OneVir supports this / who remains responsible | | --- | --- | | **Prohibited practices:** Article 5 defines particular practices, including harmful manipulation/exploitation, social scoring, and specified biometric uses. [Article 5](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-5). | Support restrictions on identified prohibited applications. Intent, thresholds, and exceptions need application context; a gateway cannot conclusively classify all legal uses from text alone. | | **High-risk risk management:** Article 9 requires an iterative documented process and testing for intended purpose and foreseeable misuse. [Article 9](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-9). | Preserve evaluation/deployment evidence and support restrictions derived from assessed risks. A model benchmark alone does not validate the complete application. | | **Human oversight:** Article 14 includes understanding limitations, avoiding automation bias, disregarding/overriding/reversing output, and intervention or safe interruption. [Article 14](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-14). | Cancellation/suspension can support oversight. The application must supply the competent human, understandable information, and meaningful decision authority. | | **Accuracy, robustness, cybersecurity:** Article 15 includes fault resilience and appropriate protection against poisoning, adversarial inputs, and confidentiality attacks. [Article 15](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-15). | Support measured performance, controlled updates, and security. The article does not specify universal accuracy percentages or prompt-filter thresholds. | | **High-risk logging:** Articles 19 and 26(6) require retention of automatically generated logs under the actor's control for an appropriate period of at least six months, unless applicable law provides otherwise, particularly data-protection law. [Article 19](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-19), [Article 26](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-26). | Provide scoped retention/export. This does not require keeping every prompt/output forever or six-month retention for all non-high-risk chat. | | **Transparency:** Article 50 distinguishes AI-interaction notices, provider duties for machine-readable synthetic-output marking/detection, and deployer disclosures for deepfakes/public-interest text. It includes exceptions and accessibility requirements. [Article 50](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-50). | Preserve provenance/markings through transformations and support client notices. Metadata alone does not provide required visible, accessible disclosure. | | **GPAI model providers:** Article 53 covers technical/downstream documentation, copyright policy, and a public training-content summary, with specified open-source exceptions. [Article 53](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53). | Keep supplier documentation available. Do not automatically assign a developer's training-summary obligation to an operator merely forwarding inference requests. | ## 5. Additional published standards and voluntary guidance These references provide relevant practices even when a particular AI statute does not apply. Their inclusion does not make them mandatory worldwide. | Reference | Verified scope and useful controls | | --- | --- | | **ISO/IEC 23894:2023** | AI risk-management guidance for organisations developing, deploying, or using AI. Complements the management-system standards above. [ISO abstract](https://www.iso.org/standard/77304.html). | | **OWASP LLM Top 10, 2025 edition** | Prompt injection, sensitive-information disclosure, supply chain, poisoning, improper output handling, excessive agency, system-prompt leakage, vector/embedding weaknesses, misinformation, and unbounded consumption. A security risk reference. [Current list](https://genai.owasp.org/llm-top-10/). | | **OWASP API Security Top 10, 2023 edition** | Conventional API risks: authentication, object/function authorisation, resource consumption, SSRF, misconfiguration, and unsafe upstream API consumption. [Publisher's list](https://api-security.owasp.org/editions/2023/en/0x11-t10/). | | **NCSC/CISA and international partners — Guidelines for Secure AI System Development** | Secure design, development, deployment, and operation/maintenance; explicitly relevant to hosted models and external APIs. [Guidelines](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development). | | **Singapore — GenAI and Agentic AI governance frameworks** | GenAI guidance addresses accountability, data, deployment, incidents, testing, security, and provenance. The Agentic AI framework launched **22 January 2026** covers responsible agent deployment and human accountability. [GenAI dimensions](https://www.imda.gov.sg/-/media/imda/files/news-and-events/media-room/media-releases/2024/05/annex-a-nine-dimensions-of-the-model-ai-governance-framework-for-generative-ai.pdf), [Agentic AI publication](https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf). | | **Japan — AI Guidelines for Business, version 1.2** | METI published version 1.2 in **March 2026**. Separate this guidance from Japan's AI promotion law, fully effective **1 September 2025**, which establishes governmental measures and cooperation responsibilities. Neither is the EU high-risk regime. [Current guidelines](https://www.meti.go.jp/shingikai/mono_info_service/ai_shakai_jisso/20260331_report.html), [Cabinet Office law outline](https://www8.cao.go.jp/cstp/ai/ai_hou_gaiyou_en.pdf). | | **Australia — Guidance for AI Adoption** | Current guidance evolves the older Voluntary AI Safety Standard into six essential practices. Its implementation guidance covers accountability, risk, oversight, testing/monitoring, records, and supply-chain governance. [Replacement explanation](https://www.industry.gov.au/publications/voluntary-ai-safety-standard), [implementation guidance](https://www.ai.gov.au/staying-safe-and-responsible/essential-ai-practices/guidance-ai-adoption-implementation-guidance). | ## 6. Published control families and their application to OneVir Each source requires an outcome for a defined deployment or recommends a security practice. **OneVir application** translates that evidence into product responsibilities; it does not claim a regulator specified this exact implementation. ### 6.1 Request admission, identity, and data movement | Control and published basis | How OneVir supports this | Operator / application responsibility | | --- | --- | --- | | **Authenticate callers and protect authentication flows.** [OWASP API2](https://api-security.owasp.org/editions/2023/en/0xa2-broken-authentication/). | Authenticate inference/admin access as required by the deployment; protect credentials and apply revocation to subsequent access. | The model must not authenticate users through conversation. | | **Protect network transport and sensitive storage.** OWASP recommends HTTPS for secure REST services and risk-appropriate cryptographic storage/key management. [REST security](https://cheatsheetseries.owasp.org/cheatsheets/REST_Security_Cheat_Sheet.html), [cryptographic storage](https://cheatsheetseries.owasp.org/cheatsheets/Cryptographic_Storage_Cheat_Sheet.html). | Protect network API credentials/payloads and upstream connections; secure retained sensitive data and keep encryption keys protected. | Encryption supports security; it does not establish lawful processing, recipient approval, or complete compliance. | | **Authorise objects and privileged functions.** [OWASP API1](https://api-security.owasp.org/editions/2023/en/0xa1-broken-object-level-authorization/), [API5](https://api-security.owasp.org/editions/2023/en/0xa5-broken-function-level-authorization/). | Scope model/provider access, stored files/conversations, cache/memory objects, and administration to the authenticated principal. | A valid key does not imply access to every tenant's objects. | | **Minimise and protect personal/sensitive information.** [GDPR principles](https://commission.europa.eu/law/law-topic/data-protection/reform/rules-business-and-organisations/principles-gdpr/overview-principles/what-data-can-we-process-and-under-which-conditions_en), [OWASP LLM02](https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/). | Apply approved restrictions before external disclosure and during storage/logging; use appropriate sanitisation/redaction and access control. | Redaction does not prove anonymisation, lawful processing, or contractual approval. | | **Control external recipients and transfers.** [Commission transfers](https://commission.europa.eu/law/law-topic/data-protection/international-dimension-data-protection/rules-international-data-transfers_en), [NCSC secure design](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-design). | Restrict eligible providers/regions to those approved for the client's data/purpose, including fallback and retry destinations. | Verify contracts, subprocessors, training use, retention, and processing locations with supplier evidence. A hostname alone does not establish them. | | **Prevent server-side request forgery.** Validate destinations and restrict network access for input-derived server-side requests. [OWASP API7](https://api-security.owasp.org/editions/2023/en/0xa7-server-side-request-forgery/). | Constrain remote resource/model fetching, media URLs, and permitted upstream endpoints where these paths exist. | Deployment network controls also matter; validating URL syntax is insufficient. | | **Bound resource/financial consumption.** Published mitigations include size limits, quotas, timeouts, resource allocation, monitoring, and queue limits. [OWASP LLM10](https://genai.owasp.org/llmrisk/llm102025-unbounded-consumption/). | Bound input/output size, concurrency/queues, duration, controlled agent iterations, and external spend. Include retries in aggregate limits. | Derive numbers from assessed capacity/client policy, not invented regulatory limits. | ### 6.2 Inference, retrieval, generated outputs, and tools | Control and published basis | How OneVir supports this | Operator / application responsibility | | --- | --- | --- | | **Mitigate direct/indirect prompt injection.** RAG/fine-tuning do not fully remove the risk. [OWASP LLM01](https://genai.owasp.org/llmrisk/llm01-prompt-injection/). | Treat user inputs, retrieved documents, files, and tool results as untrusted content; retain enforceable boundaries outside the model. | A prompt or detection model is not guaranteed prevention. | | **Keep secrets/security decisions outside system prompts.** Prompts can leak and must not be relied upon for authorisation. [OWASP LLM07](https://genai.owasp.org/llmrisk/llm072025-system-prompt-leakage/). | Keep provider credentials and privileged policy enforcement in gateway/runtime security mechanisms, separate from model text. | “Never reveal this secret” is not secret protection. | | **Limit tool functionality, permissions, and autonomy.** Execute in user context, mediate downstream authorisation, and require approval for high-impact actions as appropriate. [OWASP LLM06](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/). | Where OneVir hosts/mediates tools, restrict operations, validate calls/arguments, and preserve user context. | Tool-call JSON does not enforce an external agent's execution. The executor must enforce these controls. | | **Validate and safely handle outputs.** Generated data can cause injection in downstream browsers, interpreters, databases, and commands. [OWASP LLM05](https://genai.owasp.org/llmrisk/llm052025-improper-output-handling/). | Validate structured outputs/arguments where consumed; encode/sanitise generated content in OneVir's UI and consumers. | Valid JSON does not make embedded commands, URLs, SQL, or HTML safe. | | **Enforce permission-aware retrieval and prevent cross-context leakage.** Fine-grained permissions, partitioning, source validation, and monitoring are published mitigations. [OWASP LLM08](https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/). | Preserve scope for embeddings, retrieval, memory, and semantic reuse. Similarity is not permission to reuse a stored answer. | Embeddings are not automatically anonymous; access/deletion depend on their content/use. | | **Assess misinformation and fitness for purpose.** [OWASP LLM09](https://genai.owasp.org/llmrisk/llm092025-misinformation/), [EU Article 9](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-9). | Retain model/version/evaluation evidence and communicate supported-use limitations. | The application owns factual verification and consequential decisions. Fluency or a generated explanation is not validation evidence. | | **Apply the relevant content/disclosure policy.** EU and China sources specify different scopes and duties. [EU Article 50](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-50), [China service measures](https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm). | Support the selected deployment's restrictions, refusal/escalation, and output labelling. | Do not invent one worldwide prohibited-topic list; restrictions, exceptions, and publication duties differ. | ### 6.3 Model lifecycle, evidence, and operational response | Control and published basis | How OneVir supports this | Operator / application responsibility | | --- | --- | --- | | **Protect the model/software supply chain.** Risks include untrusted models/components, licence restrictions, weak provenance, and outdated dependencies. [OWASP LLM03](https://genai.owasp.org/llmrisk/llm032025-supply-chain/). | Track origin, revision, integrity evidence, licence/permitted use; protect import/update paths and inference dependencies. | Download availability or an open licence does not prove safety or permission for every use. | | **Isolate untrusted imports and protect inference assets.** NCSC recommends scanning/isolation when importing third-party models/weights and tracking/protecting models, data, prompts, and logs. [Secure design](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-design), [secure development](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-development). | Restrict access to worker/model assets and use suitable isolation for model import/loading; retain provenance and a way to restore known-good assets. | A file-format choice alone does not establish that an untrusted model or inference dependency is safe. | | **Guard against poisoning.** Validate sources and integrity; assess data/model changes. [OWASP LLM04](https://genai.owasp.org/llmrisk/llm042025-data-and-model-poisoning/). | Validate imported models and retrieval/memory ingestion under OneVir's control; investigate unexpected behaviour after changes. | A gateway cannot reconstruct supplier training provenance from outputs. | | **Inventory, validate, approve, and monitor models.** E-23 expects risk-proportional governance, independent review, change control, and monitoring. [OSFI E-23, sections B–D](https://www.osfi-bsif.gc.ca/en/guidance/guidance-library/guideline-e-23-model-risk-management-2027). | Supply version/change/evaluation/operation evidence; regulated clients may need review of model/provider/routing changes. | Health checks or generic benchmarks do not independently validate a financial model. | | **Log and monitor without unnecessary data exposure.** [NCSC operation/maintenance](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-operation-and-maintenance), [GDPR principles](https://commission.europa.eu/law/law-topic/data-protection/reform/rules-business-and-organisations/principles-gdpr/overview-principles/what-data-can-we-process-and-under-which-conditions_en). | Keep protected evidence of access, selection, policy decisions, changes, failures, and incidents; minimise raw content and provide authorised export/deletion. | Required duration/content vary. EU high-risk logs and Colorado compliance records are different record classes. | | **Maintain incident response and controlled release.** Guidance calls for plans, security evaluation, limitations, secure defaults, and customer audit information. [NCSC secure deployment](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-deployment). | Support model/provider isolation/suspension, evidence preservation, remediation, and controlled update activation. | Reporting recipients/deadlines are regime-specific; reporting and customer communications remain assigned organisational duties. | ## 7. Policy must survive routing, caching, and streaming Sources do not mandate a particular pipeline. These are **engineering consequences** of applying their authorisation, privacy, output, and supplier restrictions to OneVir's serving paths: 1. **A fallback is another disclosure.** Re-evaluate whether the next provider is permitted before sending data. A primary-provider failure does not remove transfer/recipient restrictions. [Transfer rules](https://commission.europa.eu/law/law-topic/data-protection/international-dimension-data-protection/rules-international-data-transfers_en). 2. **A cache hit is another access to stored data.** Authorise the requester and reused material; preserve applicable retention/deletion. Similarity is not permission. This applies the published authorisation and leakage principles to caching. [OWASP API1](https://api-security.owasp.org/editions/2023/en/0xa1-broken-object-level-authorization/), [OWASP LLM08](https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/). 3. **Released stream content has already been disclosed.** If selected policy requires inspection before disclosure, a check after delivery cannot satisfy it. Choose buffering, staged release, or another control according to assessed risk; no source mandates one universal strategy. [OWASP LLM02](https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/). 4. **Tool execution needs its own enforcement point.** Preserve user scope and required approval at hand-off to the actual tool. [OWASP LLM06 complete mediation](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/). 5. **Logs, caches, and exports are also processing.** Local inference does not remove privacy duties for other stores/telemetry recipients containing personal data. [GDPR processing scope](https://commission.europa.eu/law/law-topic/data-protection/information-business-and-organisations/application-gdpr_en). These identify where to apply a selected policy consistently. They do not require every customer to enable every filter, all data to stay local, or imply that a gateway alone makes a deployment compliant. ## 8. Responsibilities of the operator and consuming application The cited regimes also require decisions/processes outside the inference request: - **Legal basis, purpose, notices, rights, and contracts:** the responsible controller/operator establishes these and directs processors appropriately. [GDPR roles](https://commission.europa.eu/law/law-topic/data-protection/information-business-and-organisations/application-gdpr_en), [EDPB rights and processor assistance](https://www.edpb.europa.eu/sme/be-compliant/respect-individuals-rights_en). - **Meaningful oversight and recourse:** the client supplies human authority and the user process where required. Cancellation alone is insufficient. [EU Article 14](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-14), [Colorado duties](https://www.leg.colorado.gov/bills/SB26-189). - **Governance and independent review:** policies, owners, management review, validation, and approval are organisational activities. [ISO 42001](https://www.iso.org/standard/42001), [OSFI E-23](https://www.osfi-bsif.gc.ca/en/guidance/guidance-library/guideline-e-23-model-risk-management-2027). - **Supplier facts and disclosures:** obtain applicable documentation and verify operational/contractual facts. API compatibility does not establish provenance, processing location, or supplier compliance. [EU Article 53](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53), [NCSC due diligence](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development/guidelines/secure-design). For implementation work, establish deployment facts, select cited obligations/practices, and inspect/test the affected OneVir paths. This research does not assign implementation status, invent acceptance thresholds, or assert certification. Refresh sources for new jurisdictions, regulated uses, or providers and before relying on future commencement dates. --- Canonical: https://onequill.dev/resources/bifrost-management-api-boundaries # Bifrost: management APIs are an execution boundary These capabilities have greater authority than model inference. Registering a process or loading native code can cross directly into the host execution boundary. The vendor distinguishes the published static Docker builds: plugin.Open fails there, while the remote fetch still creates a blind-SSRF concern. ## What deployment does this concern? **Component and scope:** Reachable management APIs with authentication disabled; stdio MCP registration and remote custom-plugin loading. **Affected versions / scope:** Stdio advisory: < 1.5.27. Remote-plugin advisory: < 1.6.3. **Vendor remediation:** Both advisories identify 2.1.0. The management API must be reachable with authentication disabled. Stdio MCP registration can start processes as the service user. The remote-plugin code-execution path additionally requires a dynamically linked build produced with DYNAMIC=1. ## Vendor response and practical action The vendor's 2.1.0 notes corroborate authentication protections for stdio registration and private-address handling for remote plugins. Use the supported patched release, authenticate administration and keep management access separate from model callers. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn An API key that permits inference should not automatically permit plugin installation or process registration. Test administration with the credentials used by an ordinary client, including after a proxy or ingress is added. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Can an application caller reach MCP registration or plugin installation? - Is the deployed artifact static or dynamically linked? - Which identity and filesystem permissions apply to spawned processes? - Does outbound fetching reject private destinations and redirects? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway security: protect the boundaries that matter](/resources/ai-gateway-security) for the wider evaluation context. ### Technical detail: evidence and identifier limits Two mechanisms are discussed in one case because they share an administrative trust boundary. The remote-plugin RCE finding must not be applied to every static image. - Version ranges belong to separate advisories; do not combine them into one affected interval. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2026-09-23**. These dates do not change merely because this article is rebuilt. ## Primary sources - [Stdio MCP advisory · GHSA-86gf-xh3g-rvxq](https://github.com/maximhq/bifrost/security/advisories/GHSA-86gf-xh3g-rvxq) - [Remote plugin advisory · GHSA-2qp8-4xgm-fw6g](https://github.com/maximhq/bifrost/security/advisories/GHSA-2qp8-4xgm-fw6g) - [Bifrost 2.1.0 changelog](https://docs.getbifrost.ai/changelogs/v2.1.0) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/envoy-ai-gateway-mcp-message-smuggling # Envoy AI Gateway: one message, two interpretations If the policy stage checks one interpretation and the destination acts on another, an authorised-looking envelope can conceal an action the policy did not approve. Clients need assurance that tool name, arguments and the forwarded message are bound to the same interpretation. ## What deployment does this concern? **Component and scope:** MCP JSON-RPC parsing in the Envoy AI Gateway / Agent Router project. **Affected versions / scope:** < 0.6.0 **Vendor remediation:** 0.6.0 The affected MCP path interprets a message differently across parsing and forwarding stages. The finding concerns protocol interpretation, not a generic vulnerability in every Envoy deployment. ## Vendor response and practical action The vendor identifies 0.6.0 as patched. Update the affected project and check the deployed MCP path, including policy plugins and any intermediate normalisation. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Protocol compatibility includes security semantics. Evaluate ambiguous input handling and rejection behaviour as well as successful tool calls. A parser repair should be verified on the exact path that enforces tool policy. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Does policy inspect exactly the message sent to the tool? - Are ambiguous or duplicate protocol fields rejected consistently? - Which gateway release and parser are on the MCP route? - Can the operator show a regression result without sharing exploit payloads? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway security: protect the boundaries that matter](/resources/ai-gateway-security) for the wider evaluation context. ### Technical detail: evidence and identifier limits The project repository now uses Agent Router naming. This case does not transfer the finding to the whole Envoy proxy ecosystem. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2026-05-13**. These dates do not change merely because this article is rebuilt. ## Primary sources - [GHSA-4gph-2hhr-5mwg · vendor advisory](https://github.com/theagentrouter/agent-router/security/advisories/GHSA-4gph-2hhr-5mwg) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/envoy-ai-gateway-mcp-request-limits # Envoy AI Gateway: enforce limits before buffering Request limits applied after allocation protect downstream work but may leave the gateway itself exposed to memory exhaustion. Availability depends on where a limit runs, how many bodies can be buffered concurrently and whether cancellation releases the allocated resources. ## What deployment does this concern? **Component and scope:** MCP POST-body buffering in the external-processing component. **Affected versions / scope:** 0.4.0 through 0.7.0 **Vendor remediation:** 1.0.0 The affected path reads the full MCP POST body before applying the relevant limits. An authenticated caller may still be able to send a body large enough to exhaust component memory. ## Vendor response and practical action The vendor identifies 1.0.0 as patched. Upgrade and configure request-size, concurrency and timeout controls at the earliest ingress layer and the processing component. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn A small per-request limit is not a complete capacity model. Multiply buffering by admitted concurrency and account for protocol expansion, such as decoding, before choosing a memory budget. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Where is the body-size limit enforced relative to allocation? - What is the maximum simultaneously buffered input? - Do authenticated tenants share the same processing memory? - How are rejection, cancellation and recovery observed? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway reliability: define failure before adding fallback](/resources/ai-gateway-reliability) for the wider evaluation context. ### Technical detail: evidence and identifier limits This is a documented MCP component issue. It is not evidence that all large LLM requests or every Envoy data plane will fail. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2026-09-26**. These dates do not change merely because this article is rebuilt. ## Primary sources - [GHSA-43xg-mvg9-qwpq · vendor advisory](https://github.com/theagentrouter/agent-router/security/advisories/GHSA-43xg-mvg9-qwpq) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/gitlab-ai-gateway-template-sandbox # GitLab AI Gateway: flow templates and sandbox assumptions The vendor describes a template sandbox escape. Authoring a workflow is therefore a security-sensitive capability even when the user does not hold host-administration rights. ## What deployment does this concern? **Component and scope:** GitLab Duo Agent Platform flow-template processing. **Affected versions / scope:** 18.1.6 to < 19.2.4; 19.3 to < 19.3.2; 19.4 to < 19.4.1 **Vendor remediation:** 19.2.4, 19.3.2 and 19.4.1 An authenticated user with access to the Duo Agent Platform can submit a crafted flow template to the affected template-processing path. This is an application-specific prerequisite. ## Vendor response and practical action The vendor released branch-specific patches and states its hosted gateways were already fixed. Self-hosted operators retain their upgrade responsibility and should use the patch appropriate to their release branch. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Ask what an author can cause a workflow runtime to execute, read and contact. A template language or sandbox needs a defined permission boundary and a plan for patching that runtime. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Who can create and submit flow templates? - Which AI Gateway branch and patch are installed? - What host permissions and secrets are available to the runtime? - Does a hosted-service patch cover your self-hosted components? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway security: protect the boundaries that matter](/resources/ai-gateway-security) for the wider evaluation context. ### Technical detail: evidence and identifier limits This case describes GitLab's application-specific gateway. It is not a generic finding against interchangeable AI model gateways. - Vendor identifier: CVE-2026-90970. Keep branch-specific intervals separate. Evidence label: **security advisory**. Source-review date: **2026-10-07**. A publication date is not stated in the linked release notice. Review and article dates do not change merely because this article is rebuilt. ## Primary sources - [GitLab AI Gateway patch release notice](https://docs.gitlab.com/releases/patches/other-patches/patch-release-gitlab-ai-gateway-19-4-1-released/) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/kong-gemini-streaming-token-accounting # Kong: a streaming usage fix deserves billing regression tests Reliable accounting requires understanding whether a value is cumulative, incremental or absent. Summing every reported total can overstate a charge; treating absence as zero can erase evidence of prior work. ## What deployment does this concern? **Component and scope:** Kong AI Proxy plugin handling of Gemini streaming usage. **Affected versions / scope:** The cited release notes describe the fix; no security-advisory affected interval is supplied here. **Vendor remediation:** 3.16.0.0 includes the stated fix. Streaming Gemini responses can repeat running usage metadata or omit metadata in some chunks. The release notes address overcounting repeated cumulative values and resetting usage when metadata is missing. ## Vendor response and practical action The vendor's 3.16.0.0 changelog records the correction. Confirm which AI Proxy plugin is deployed and reproduce the usage pattern with provider-side accounting evidence. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Use release notes as evaluation evidence without promoting every correction to a CVE. A billing test should include complete streams, disconnected streams, missing final events and retries. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Is each usage field incremental or cumulative? - What remains recorded if the client disconnects? - Which plugin version processes the Gemini response? - How are gateway measurements compared with the provider bill? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway budgets and scaling: measure the whole request](/resources/ai-gateway-budgets-and-scaling) for the wider evaluation context. ### Technical detail: evidence and identifier limits This is labelled a release fix. The cited entry does not establish an independent vulnerability or apply automatically to every newer Kong AI offering. Evidence label: **release fix**. Source-review date: **2026-10-07**. Source publication or event date: **2026-09-15**. These dates do not change merely because this article is rebuilt. ## Primary sources - [Kong AI Proxy plugin changelog](https://developer.konghq.com/plugins/ai-proxy/changelog/) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/litellm-march-2026-package-incident # LiteLLM: package provenance and incident response A dependency update is a change to executable code with the authority of the installation environment. Exposure assessment needs the actual artifact and install history, not just the top-level application name. ## What deployment does this concern? **Component and scope:** Malicious PyPI distributions, distinct from the vendor's official pinned Proxy Docker images. **Affected versions / scope:** PyPI LiteLLM 1.82.7 and 1.82.8 during the reported publication window. **Vendor remediation:** Vendor describes a clean 1.83 release and pipeline v2 on 30 March; incident response also requires containment and credential review. The organisation installed or executed one of the affected PyPI versions. The vendor reports a roughly 40-minute window beginning at 10:39 UTC on 24 March. The vendor states its official Proxy Docker distribution was unaffected because it pinned requirements. ## Vendor response and practical action The vendor removed the affected packages, investigated the incident and rebuilt the release pipeline. Its account of a connection to a Trivy-related compromise is expressed as a belief, not independently established attribution. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn An upgrade removes a malicious artifact from the future path; it does not establish that previously accessible secrets are safe. Preserve artifact evidence, isolate affected environments and rotate credentials according to the incident investigation. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - What package versions and image digests were installed during the window? - Which secrets and network permissions were available to that environment? - Can builds be reproduced from a trusted, pinned dependency set? - Who owns containment, rotation and verification before service resumes? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway supply chain: trust the artifact you actually run](/resources/ai-gateway-supply-chain) for the wider evaluation context. ### Technical detail: evidence and identifier limits This article summarises the vendor's incident account. It does not reproduce malicious code, assert every LiteLLM installation was compromised or equate a package incident with the security of all deployments. Evidence label: **incident report**. Source-review date: **2026-10-07**. Source publication or event date: **2026-03-24**. These dates do not change merely because this article is rebuilt. ## Primary sources - [LiteLLM March 2026 security update](https://docs.litellm.ai/blog/security-update-march-2026) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/llm-inference-security-and-operations # Running LLM inference in production: security, isolation and capacity An AI gateway and an inference engine occupy different parts of the request path. The gateway can authorise callers and choose a route; the engine loads a model, validates inputs and schedules GPU work. Clients should evaluate both components and assign an owner to the boundary between them. This guide covers **vLLM, SGLang, Ollama and LM Studio**. Their roles and deployment choices differ. Use the official documentation and scoped cases to evaluate the configuration you actually run, alongside the gateway that routes requests to it. ## Choose the right evaluation boundary | Component | Deployment boundary to establish | Capacity evidence to request | | --- | --- | --- | | [vLLM](/resources/ai-gateway-providers#vllm) | Public inference endpoints, distributed channels, enabled input features and model-loading configuration. | Queue time, preemption, cache reuse, TTFT and token latency for the model and topology. | | [SGLang](/resources/ai-gateway-providers#sglang) | Inference API, administrative operations, distributed workers and approved model/adapter sources. | Cache-hit and miss workloads, session lifecycle, memory pressure and worker recovery. | | [Ollama](/resources/ai-gateway-providers#ollama) | Local API versus cloud execution, network access and model-management permissions. | Model residency, context allocation, parallel requests, queue rejection and CPU/GPU placement. | | [LM Studio](/resources/ai-gateway-providers#lm-studio) | API-token permissions, server binding and the device selected by LM Link. | JIT versus explicit model loading, idle TTL, Auto-Evict and cold/warm request latency. | The table is an evaluation framework. It does not establish integration support, relative speed or the security of a product from its name. Record the release, model, enabled features and effective configuration with each answer. ## SGLang: separate API, administration and worker trust [Current server arguments](https://docs.sglang.io/docs/advanced_features/server_arguments) distinguish `--api-key` from `--admin-api-key` for designated control endpoints. Both are optional configuration fields. Establish which endpoints the installed release protects, and isolate worker communication independently. Model revisions and any use of `--trust-remote-code` also need approval. The [SGLang case](/resources/sglang-worker-and-management-trust) separates the March serialization/replay findings from a later July coordinated notice. Their prerequisites and remediation records differ; the case preserves conflicting version information rather than presenting one universal fixed release. [Session-aware radix caching](https://github.com/sgl-project/sglang/blob/main/docs/docs/advanced_features/session_radix_cache.mdx) uses session references to influence eviction. Closing a session removes its references without immediately freeing all reusable KV; references can still be evicted under pressure. Evaluate cache behaviour and tenant permissions separately. A `session_id` used for cache lifecycle is not proof of an authorised tenant boundary. ## Ollama: distinguish local service and cloud execution [Ollama documents no authentication for the local API](https://docs.ollama.com/api/authentication), while access to its cloud API uses credentials. Its [deployment FAQ](https://docs.ollama.com/faq) describes loopback binding by default, network configuration and a local-only setting such as `OLLAMA_NO_CLOUD=1`. Identify the selected model and execution mode; a local API address alone does not establish the entire processing route. Separate prompt submission from model creation, imports and exports. The [GGUF model-import case](/resources/ollama-model-import-memory-boundary) explains a reported tensor-validation issue and why a registry patch field needs artifact evidence. Use `ollama ps` and the configured context, parallelism, queue and `keep_alive` settings in the capacity review. The [context guidance](https://docs.ollama.com/context-length) explains that larger context requires more memory; current documentation can evolve, so record the effective value rather than assume one default applies to every release and machine. Test model switching and overload as well as repeated requests to a warm model. ## LM Studio: verify access and the execution device [Native API-token authentication](https://lmstudio.ai/docs/developer/core/authentication) is documented for 0.4.0 or newer and can be required with selected permissions. The vendor recommends authentication when [serving beyond localhost](https://lmstudio.ai/docs/developer/core/server/serve-on-network). Confirm the effective setting on the API and SDK paths your application uses. With [LM Link](https://lmstudio.ai/docs/developer/core/lmlink), a localhost request can be served by a linked remote machine. Record that device and its approved data boundary. The vendor's [offline-operation guidance](https://lmstudio.ai/docs/app/offline) describes local use of downloaded models; model discovery, runtime downloads and enabled integrations deserve their own network review. The [LM Studio configuration case](/resources/lm-studio-network-authentication-and-lifecycle) is labelled **documented behaviour**. It covers access, execution location and model residency. [Idle TTL and Auto-Evict](https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict) affect JIT-loaded models differently from explicit loads, so measure first-load latency and model-switching behaviour under the settings you will use. ## vLLM: inventory public and internal interfaces [Current vLLM security documentation](https://docs.vllm.ai/en/latest/usage/security/) describes API-key protection for `/v1`, `/v2`, `/inference` and `/cohere` prefixes. Other inference, operational, development and profiling endpoints have different exposure conditions. Record which endpoints exist under your enabled features and how ingress policy controls them. The same guidance warns about insecure defaults in several inter-node communication paths. Map the actual backend and topology rather than describe every deployment as a single ZeroMQ channel. The [V0 PyNcclPipe case](/resources/vllm-kv-transfer-network-isolation) is specifically about KV-cache transfer and has a historical version range and fix. ## Approve model assets and input features separately The [Python optimisation case](/resources/vllm-model-loading-python-optimisation) requires loading malicious model configuration while Python uses `-O` or `PYTHONOPTIMIZE=1`. These settings differ from vLLM's performance flags. Model approval, runtime configuration and request authorisation are separate controls. The [input-validation case](/resources/vllm-multimodal-input-validation) collects prompt-embedding validation and video-frame limits with their separate evidence. Optional features deserve separate limits and validation. The concurrency follow-up establishes an invariant-check bypass; its reproduction did not establish a live-server crash or code execution. ## Distinguish tenant-isolation mechanisms [vLLM documents optional per-request cache salting](https://docs.vllm.ai/en/latest/design/prefix_caching/) to separate prefix-cache reuse between trust groups and reduce timing inference concerns. That is distinct from process isolation, GPU allocator behaviour and kernel correctness. The [GGUF case](/resources/vllm-gguf-gpu-memory-isolation) concerns integer truncation in specific dequantisation kernels, leaving portions of an output tensor uninitialised. The current vendor advisory lists 0.24.0 as patched. It is not a general finding that every user's KV cache is left unwiped, and cache salting is not its remediation. ## Measure cache pressure and scheduling tradeoffs [The current V1 tuning guidance](https://docs.vllm.ai/en/v0.31.0/configuration/optimization/) describes decode-prioritised scheduling and chunked prefill enabled where possible. Under KV-cache pressure, requests can be preempted and later recomputed. The effect on latency and throughput depends on the workload. Measure time to first token, inter-token latency, end-to-end latency, queue time, preemption and memory under your input/output mix. Vary admitted concurrency and token batching deliberately. Avoid universal claims such as “concurrency drops to zero” or a fixed CPU request-per-second ceiling. Cache-aware routing introduces another tradeoff. Production Stack documents [prefix-aware routing](https://docs.vllm.ai/projects/production-stack/en/latest/use_cases/prefix-aware-routing.html) and [load-aware routing](https://docs.vllm.ai/projects/production-stack/en/latest/use_cases/loadaware-routing.html). Reusing a warm prefix can reduce computation, but concentrating traffic on a warm worker can increase queue pressure. ### Technical detail: distributed recovery and startup [vLLM's parallelism documentation](https://docs.vllm.ai/en/stable/serving/parallelism_scaling/) supports multi-node Ray and multiprocessing options. Ray is not universally required. [Ray Serve documents fault tolerance](https://docs.ray.io/en/latest/serve/production-guide/fault-tolerance.html) when the deployment and KubeRay recovery mechanisms are configured appropriately. Record tensor, pipeline and data-parallel layout; node and process failure domains; and who restores workers or the head service. Evaluate tokenisation and API CPU work separately from GPU inference. For startup, record weight storage, download caching, compilation caches, memory profiling and warm readiness. Publish numeric startup or throughput claims only with the hardware, software and reproducible workload. ## Questions for your evaluation Use the inference page of the [worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf). Attach an endpoint inventory, topology, model and feature configuration, isolation evidence and a workload-specific capacity report. Record unsupported features and untested outcomes explicitly. ## Applying this to OneVir [OneVir describes separate local and provider execution paths](https://onevir.onequill.dev/#path). Our [implementation evidence record](/resources/onevir-control-evidence#inference) identifies local GGUF execution through llama.cpp bindings, a separate inference implementation from vLLM. When evaluating a separately configured upstream such as vLLM, SGLang, Ollama or LM Studio, request backend evidence alongside the OneVir route configuration. Gateway access and budget policy do not establish the runtime's patch status, execution location, internal-network isolation or GPU correctness. This guide does not claim that OneVir includes these runtimes or supports every endpoint they offer. A local engine choice also needs its own model-format, memory-isolation and capacity evaluation. --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/lm-studio-network-authentication-and-lifecycle # LM Studio: choose authentication, execution location and model lifecycle The application’s API address does not fully describe who can call the server or where model execution occurs. Model switching can also introduce cold-load latency and memory pressure. ## Which deployment does this concern? Authentication is optional by default. Enabling network serving makes the API reachable beyond localhost. With LM Link, a localhost API request can be served by a model on a linked remote machine. **Applies to:** Configuration-dependent behaviour. Native API-token authentication is documented for LM Studio 0.4.0 or newer; this case does not identify a vulnerable version range. ## Vendor controls and evaluation The vendor provides native API tokens and permissions, recommends authentication for network binds, documents LM Link routing, and exposes idle TTL and Auto-Evict controls. **Configuration to validate:** Enable required authentication and scoped token permissions for shared access; approve bind addresses and remote devices; configure model residency intentionally. ## What clients can learn Verify effective server settings, token permissions and execution location. Measure first-load and repeated-request behaviour under the model lifecycle settings your application actually uses. Keep an endpoint inventory, the effective configuration and the running artifact together. A configuration or model change should trigger a review of the boundary it changes. Assign an owner to the evidence and to any required action. ## Questions for your provider - Do requests without a valid token fail on the deployed API? - Which inference, model-management and tool permissions does each token have? - Does LM Link resolve the model to an approved device? - How do JIT loading, explicit loads, idle TTL and Auto-Evict affect latency and memory? Use the [inference guide](/resources/llm-inference-security-and-operations) and [evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record the answer in your deployment context. ### Technical detail: source scope and identifiers This is documented configuration behaviour, not a security advisory, CVE or an observed compromise. The offline documentation describes local operation; enabled remote links and integrations require their own data-flow review. - No CVE is assigned by this case. Evidence type: documented behaviour. - The vendor distinguishes JIT-loaded models from explicitly loaded models. Auto-Evict does not apply indiscriminately to every model in memory. ## Primary and supporting records - [Native authentication and token permissions](https://lmstudio.ai/docs/developer/core/authentication) - [Serving on a network](https://lmstudio.ai/docs/developer/core/server/serve-on-network) - [LM Link execution location](https://lmstudio.ai/docs/developer/core/lmlink) - [Idle TTL and Auto-Evict](https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict) - [Offline operations and runtime updates](https://lmstudio.ai/docs/app/offline) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/new-api-quota-billing-overflow # New API: quota arithmetic is part of the trust boundary An admission check and final bill can disagree even when both work in ordinary examples. Numeric bounds, rounding and the sign of the final ledger entry deserve the same scrutiny as authentication. ## What deployment does this concern? **Component and scope:** Quota settlement in affected New API release candidates. **Affected versions / scope:** ≤ 1.0.0-rc.17 **Vendor remediation:** ≥ 1.0.0-rc.18 The caller must satisfy normal pre-consumption requirements through a positive balance or valid subscription. Integer overflow during settlement can then turn a charge into a credit. Self-registration with free credits increases exposure when configured. ## Vendor response and practical action The vendor reports exploitation observed on 6 July and an emergency fix on 7 July, with the advisory published on 9 July. rc.18 is the stated patch; rc.19 adds logging context. Preserve the release-candidate qualifier. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Evaluate the ledger as a state transition: reserve, dispatch, settle or release. Concurrency, very large usage values, cancellation and repeated settlement should not create spendable credit. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - What upper bounds apply to tokens, prices and multiplication? - Is the reservation atomic under competing requests? - Can a failed or repeated settlement increase a user's balance? - How are anomalous credits reconciled with provider invoices? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway budgets and scaling: measure the whole request](/resources/ai-gateway-budgets-and-scaling) for the wider evaluation context. ### Technical detail: evidence and identifier limits The finding belongs to New API. It must not be assigned to the separate One API project merely because the projects share naming conventions. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2026-07-09**. These dates do not change merely because this article is rebuilt. ## Primary sources - [GHSA-8r8v-xf7q-rcpr · vendor advisory](https://github.com/QuantumNous/new-api/security/advisories/GHSA-8r8v-xf7q-rcpr) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/ollama-model-import-memory-boundary # Ollama: model import is a separate memory and access boundary Model-management permissions can be more sensitive than inference permissions. A shared service should establish who may import, create, quantize and export models, alongside who may submit prompts. ## Which deployment does this concern? The record concerns attacker-controlled GGUF tensor offsets and sizes processed during model creation/quantization. It is not a finding that an ordinary prompt in every local Ollama installation exposes memory. **Affected scope:** The reviewed registry identifies versions before 0.17.1. Exploitation depends on supplying a malformed GGUF to the affected model-creation path; the reported export scenario also uses model push. ## Response and remediation evidence Upstream PR #14406 validates expected tensor sizes during model creation. The registry identifies 0.17.1 as patched. The linked release is dated 24 February while the PR merge is dated 25 February; those records alone do not prove that a particular 0.17.1 artifact contains the change. **Remediation record:** Registry patched-version field: 0.17.1. Confirm inclusion of the upstream tensor-size fix in the deployed artifact. ## What clients can learn Patch the affected path with artifact-level evidence, restrict model-management operations and keep the service on approved networks. A gateway access policy supports admission but cannot repair an inference runtime’s model loader. Keep an endpoint inventory, the effective configuration and the running artifact together. A configuration or model change should trigger a review of the boundary it changes. Assign an owner to the evidence and to any required action. ## Questions for your provider - Can application users invoke create, import, quantization or push operations? - What installed artifact and commit establish the tensor-size fix? - Does the local service remain on loopback or have an authenticated network boundary? - Are model sources and export destinations approved separately? Use the [inference guide](/resources/llm-inference-security-and-operations) and [evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record the answer in your deployment context. ### Technical detail: source scope and identifiers The registry’s version field is attributed to the registry. This documentary review did not reproduce the reported memory disclosure or verify every release artifact. - Finding identifier: GHSA-x8qc-fggm-mpqg; CVE mapping: CVE-2026-7482. - Preserve the registry’s 0.17.1 patch statement alongside the upstream release/merge chronology. Request build or backport evidence rather than widening an affected range or declaring the artifact verified. ## Primary and supporting records - [GHSA-x8qc-fggm-mpqg / CVE-2026-7482 · reviewed registry record](https://github.com/advisories/GHSA-x8qc-fggm-mpqg) - [Upstream tensor-size validation, PR #14406](https://github.com/ollama/ollama/pull/14406) - [Upstream v0.17.1 release](https://github.com/ollama/ollama/releases/tag/v0.17.1) - [Local versus cloud API authentication](https://docs.ollama.com/api/authentication) - [Network binding and deployment settings](https://docs.ollama.com/faq) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/onevir-control-evidence # OneVir control evidence and deployment responsibilities Review six OneVir control topics, the evidence behind each finding, and what to verify for your release and deployment. ## Source snapshot - Published: 2026-10-07. Article updated: 2026-10-07. Source reviewed: 2026-10-07. - Author: OneQuill Research. Affiliation: OneQuill develops OneVir. - Source package: OneVir 0.2.43 (working tree; not a verified release artifact). - Base revision: `2164d6209a7f66e3c00913e430916d5595dd6c36`. - Snapshot state: Working tree with uncommitted sources. The base revision alone does not identify the reviewed files. - Snapshot captured: 2026-10-07T14:02:36Z. - Method: Documentary review of named source paths and test assertions; referenced tests were not executed for this publication. - Release applicability: Not established; request evidence for the installed release and effective configuration. This is not an independent audit, certification or deployment guarantee. The [source snapshot JSON](https://onequill.dev/assets/downloads/onevir-source-snapshot.json) records SHA-256 fingerprints for the seven named files without publishing source contents or private configuration. ## At a glance | Control and finding | Evidence status | Verify in your deployment | | --- | --- | --- | | [Security](https://onequill.dev/resources/onevir-control-evidence#security): The named test asserts rejection of conflicting credentials and handling of supported bearer and WebSocket formats. | Test source reviewed; referenced test not run | Endpoint coverage, key scopes, administrative permissions and transport security. | | [Privacy](https://onequill.dev/resources/onevir-control-evidence#privacy): Initialisation rejects an invalid saved retention policy; named history paths describe retention and image restrictions. | Implementation and API documentation reviewed | Effective retention settings and data flows through providers, caches, files, exports and backups. | | [Budgets](https://onequill.dev/resources/onevir-control-evidence#budgets): The named test asserts two admissions from eight competing attempts against a two-attempt allowance. | Test source reviewed; referenced test not run | Configured prices, settlement, incomplete streams and counters across enabled routes and nodes. | | [Reliability](https://onequill.dev/resources/onevir-control-evidence#reliability): Named source paths describe provider health states, cooldown and trial behaviour, and health-aware fallback selection. | Implementation and source documentation reviewed | Recovery times, partial streams, provider failure behaviour and restart durability. | | [Supply chain](https://onequill.dev/resources/onevir-control-evidence#supply-chain): The source manifest pins llama-cpp-sys-2 to =0.1.157; this identifies one declared dependency. | Dependency declaration reviewed | Delivered artifact digest, complete dependency inventory, model revisions and vulnerability assessment. | | [Inference](https://onequill.dev/resources/onevir-control-evidence#inference): Named sources distinguish GGUF execution on llama.cpp from exported accelerator assets. | Implementation and engine boundary reviewed | Separate inference services, patch levels, endpoint exposure, execution location and memory isolation. | ## Responsibilities: who supplies the evidence Use this division of work to plan your evaluation. It describes evidence to request and checks to assign; it does not establish contractual or legal responsibility. | Participant | Evidence to request or record | What to verify for the deployment | | --- | --- | --- | | OneQuill / OneVir implementation | Release identifiers, the relevant source snapshot, control scope, dependency declarations and available test or demonstration records. | Whether the supplied record covers the installed build and the paths being used. | | Deployment operator | Installed artifact digest, effective configuration, endpoint and permission matrix, provider and model inventory, and an accountable owner for each control. | Retention and data flows, route accounting, failure recovery, access boundaries and changes to the configuration. | | Connected provider or inference service | Service and model versions, data-handling terms and settings, patch information, endpoint exposure and service-specific test evidence. | Where each request executes and what the upstream stores, bills, retries or isolates. | Ask OneQuill for implementation and release evidence. Record your own configuration and results. Request service-specific evidence from every connected provider. A source review does not replace those deployment checks. ## Security: credential interpretation **Reviewed finding:** The named test asserts rejection of conflicting bearer and API-key credentials, normalisation of supported bearer syntax, and separate handling of WebSocket credentials. **Evidence status:** Test source reviewed. The referenced test was not run for this publication. **Verify in your deployment:** Obtain the installed version, key scopes and an endpoint and permission matrix. Check application and administrative access separately, including transport security. **Limit:** This finding covers credential interpretation. It does not establish endpoint coverage, administrative permissions or complete authentication correctness. ### Technical detail: credential test source The source snapshot identifies this test: ~~~text File: src/api/mod.rs Test: access_auth_normalizes_supported_credentials_and_rejects_ambiguity ~~~ Its assertions distinguish supported bearer syntax, matching and conflicting credentials, unsupported authentication syntax and the WebSocket credential path. These are assertions in test source, not published execution results. Product context: [OneVir capabilities](https://onevir.onequill.dev/#capabilities). Traceability: [source snapshot and file fingerprints](https://onequill.dev/assets/downloads/onevir-source-snapshot.json). ## Privacy: retention and saved history **Reviewed finding:** The named initialisation path refuses to start with an invalid saved retention policy. Named history API paths describe restrictions on saved conversations and images. **Evidence status:** Implementation and API documentation reviewed. Retention and deletion were not demonstrated for a deployment. **Verify in your deployment:** Record the effective retention settings and trace data through providers, caches, files, exports and backups. Request demonstrations for the routes and storage locations you enable. **Limit:** Initialisation and named history paths do not prove deletion across every storage location or external provider. ### Technical detail: retention and history paths ~~~text Initialisation: src/api/mod.rs History API: src/api/memory.rs ~~~ The initialisation path rejects an invalid saved retention policy. The history API documentation describes rejection when route retention prohibits saved history and restrictions on saved images under a privacy policy. API documentation is evidence of described behaviour; it does not establish every deployed outcome. Product context: [OneVir request paths](https://onevir.onequill.dev/#path). Traceability: [source snapshot and file fingerprints](https://onequill.dev/assets/downloads/onevir-source-snapshot.json). ## Budgets: concurrent spend reservation **Reviewed finding:** The named test presents eight competing attempts to a shared runtime with a two-attempt allowance. It asserts two admissions and the matching reserved amount. **Evidence status:** Test source reviewed. The referenced test was not run for this publication. **Verify in your deployment:** Request concurrent reservation and final settlement evidence using configured provider prices and enabled routes. Include incomplete streams and any counters shared across nodes. **Limit:** This is a specific concurrent reservation scenario. It does not prove every price, stream outcome, settlement path or multi-node counter. ### Technical detail: concurrent reservation test source ~~~text File: src/governance/deployment_runtime_tests.rs Test: deployment_runtime_provider_spend_reservation_is_atomic_under_competing_attempts ~~~ The test launches eight competing attempts, counts admissions and asserts the reserved amount against a two-attempt allowance. The source snapshot records this file as untracked in the reviewed working tree. Ask for its inclusion and executed result in the release evidence you receive. Product context: [OneVir capabilities](https://onevir.onequill.dev/#capabilities). Traceability: [source snapshot and file fingerprints](https://onequill.dev/assets/downloads/onevir-source-snapshot.json). ## Reliability: provider health and fallback **Reviewed finding:** Named source paths describe healthy, degraded, down and refused provider states, cooldown and trial behaviour, and health-aware fallback selection. **Evidence status:** Implementation and source documentation reviewed. Recovery service levels were not established for a deployment. **Verify in your deployment:** Demonstrate provider failure and recovery with your streaming workloads. Record recovery time, partial output handling, fallback decisions and behaviour after a process restart. **Limit:** Provider-health state and selection do not establish complete stream handling, a recovery-time guarantee or restart durability. ### Technical detail: provider health and selection paths ~~~text Provider health: src/proxy/health.rs Fallback selection: src/api/openai.rs ~~~ The health source documents states, cooldown and trial behaviour. The fallback path consults provider health when other endpoints remain to try. Effective provider settings and workload demonstrations are still needed to establish application behaviour. Product context: [OneVir capabilities](https://onevir.onequill.dev/#capabilities). Traceability: [source snapshot and file fingerprints](https://onequill.dev/assets/downloads/onevir-source-snapshot.json). ## Supply chain: a named dependency pin **Reviewed finding:** The reviewed source manifest declares an exact llama-cpp-sys-2 version of =0.1.157 and documents the intent to use one llama.cpp binding version. **Evidence status:** Dependency declaration reviewed. Delivered artifact provenance was not established. **Verify in your deployment:** Obtain the delivered artifact's digest, build provenance, complete dependency inventory and exact model revisions. Assess vulnerabilities for those actual versions and artifacts. **Limit:** One declared dependency pin is not a complete software inventory, proof of the delivered binary's contents or evidence that the dependency has no vulnerabilities. ### Technical detail: dependency declaration ~~~text Manifest: Cargo.toml Dependency: llama-cpp-sys-2 =0.1.157 ~~~ The source package version is 0.2.43 in the captured working tree. This identifies the source package; it does not verify that a particular release artifact was built from these files. The source snapshot supplies the manifest fingerprint and base revision. Product context: [OneVir capabilities](https://onevir.onequill.dev/#capabilities). Traceability: [source snapshot and file fingerprints](https://onequill.dev/assets/downloads/onevir-source-snapshot.json). ## Inference: engine and upstream boundaries **Reviewed finding:** Named sources identify llama.cpp bindings and distinguish GGUF execution from exported accelerator assets. **Evidence status:** Implementation and engine boundary reviewed. Connected inference services were not assessed. **Verify in your deployment:** Inventory every inference engine, model revision and endpoint. Obtain patch, execution-location and isolation evidence for each separately connected service. **Limit:** A finding about a separate engine cannot automatically be applied to this implementation. Local execution also does not establish the behaviour or isolation of an upstream service. ### Technical detail: model format and engine paths ~~~text Bindings: Cargo.toml Accelerator manifest: accelerators/make_manifest.py ~~~ The manifest identifies llama.cpp bindings. The accelerator manifest builder distinguishes GGUF files, which run on llama.cpp, from exported accelerator assets. A separately configured upstream such as vLLM, SGLang, Ollama or LM Studio needs its own evidence for endpoints, patch levels, execution location and memory isolation. Product context: [OneVir request paths](https://onevir.onequill.dev/#path). Traceability: [source snapshot and file fingerprints](https://onequill.dev/assets/downloads/onevir-source-snapshot.json). ## Keeping this record useful The downloadable source snapshot records the base revision, working-tree status and SHA-256 fingerprints of the seven named files. Because some reviewed files contain uncommitted changes, the base revision alone does not reproduce this review. Request the matching source snapshot or release-specific records when evaluating an installation. Recheck the relevant evidence after changes to the release, effective configuration, providers, inference engines or model revisions. Keep article modification dates separate from source-review dates, and record an owner and follow-up date for unresolved questions in the [evaluation checklist](https://onequill.dev/assets/downloads/onevir-evaluation-checklist.md). ## Request release evidence Share your version, deployment mode, enabled providers and controls to evaluate with [OneQuill sales](mailto:sales@onequill.dev?subject=OneVir%20release%20evidence%20request&body=Hello%20OneQuill%20team%2C%0A%0AI%20would%20like%20evidence%20for%20my%20OneVir%20release%20and%20deployment.%0A%0AOneVir%20version%20%2F%20build%3A%0ADeployment%20mode%20(local%2C%20hybrid%20or%20connected%20providers)%3A%0AEnabled%20providers%20and%20inference%20engines%3A%0AControls%20to%20evaluate%3A%0AInstalled%20artifact%20digest%2C%20if%20available%3A%0A%0ARequested%20evidence%3A%20release%20and%20configuration%20scope%2C%20endpoint%20permissions%2C%20retention%20and%20data%20flows%2C%20budget%20accounting%2C%20failure%20recovery%2C%20dependency%20inventory%20and%20inference%20boundaries.%0A%0AReference%3A%20https%3A%2F%2Fonequill.dev%2Fresources%2Fonevir-control-evidence%0ASource%20review%3A%202026-10-07%3B%20source%20package%200.2.43%20(working%20tree%2C%20release%20applicability%20not%20established).%0A%0AThank%20you.). Use the [evaluation checklist](https://onequill.dev/assets/downloads/onevir-evaluation-checklist.md) to record evidence and decisions. --- Canonical: https://onequill.dev/resources/portkey-custom-host-ssrf # Portkey: custom-host routing and private-network access The documented SSRF lets the gateway make requests to destinations that should not be selectable by an application caller. The business concern is the gateway's network authority: credentials and access intended for a trusted operator can become reachable through an untrusted request. ## What deployment does this concern? **Component and scope:** Open-source Portkey gateway custom-host routing. **Affected versions / scope:** < 1.14.0 **Vendor remediation:** 1.14.0 An application caller can influence a custom upstream host on an affected open-source gateway. A useful attack also depends on the gateway being able to reach a private destination or metadata service. ## Vendor response and practical action The vendor identifies 1.14.0 as patched. Upgrade the affected distribution, restrict custom-host selection and enforce outbound network rules that also cover resolved addresses and redirects. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn A provider name in a route is only the start of the decision. Check the actual host and network destination at dispatch time. A tenant should not gain the gateway's access to internal services by changing an endpoint header. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Can an application key set a custom host or base URL? - Are private, loopback and metadata destinations blocked after DNS resolution? - Do redirects and fallback endpoints pass the same destination checks? - Which installed build and regression evidence establish remediation? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [AI gateway security: protect the boundaries that matter](/resources/ai-gateway-security) for the wider evaluation context. ### Technical detail: evidence and identifier limits The advisory does not establish that every managed Portkey or Prisma AIRS deployment was affected. A product-family name is insufficient evidence of exposure. - Primary identifier: GHSA-hhh5-2cvx-vmfp. Vendor mapping: CVE-2025-66405. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2025-12-01**. These dates do not change merely because this article is rebuilt. ## Primary sources - [GHSA-hhh5-2cvx-vmfp · vendor advisory](https://github.com/Portkey-AI/gateway/security/advisories/GHSA-hhh5-2cvx-vmfp) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/resource-publishing # Publishing the education centre The centre is a static, manually curated publication. Its source of truth is `content/resources.json`, Markdown under `content/articles/`, and `content/gateway-research.json`. Generated HTML lives under `www/`; edit sources and rebuild, not generated articles. ## Regular publication 1. Copy a template from `content/templates/`, or run `npm run content:new -- article-slug "Article title"`. New articles begin as drafts. 2. Research the exact product, edition, artifact, version, feature and deployment. Prefer vendor advisories, official documentation, release notes and incident notices. Save URLs and reviewed dates in the research record. 3. Write original client-facing analysis: business consequence, prerequisites, vendor response, limits and questions. Keep exploit payloads and unsupported accusations out of the article. 4. Complete the manifest, including topic, related cases/guides, summary and takeaways. Set status to published only when ready. Drafts never enter HTML, Markdown, feeds, structured data or public research downloads. 5. Run `npm run content:build`, then `node scripts/check-resources.mjs`. Review the affected article and filters in a browser. Check the specific links and download that changed. 6. Update publication dates only for first publication; update modified dates for substantive edits. Rechecking a source updates its reviewed date without inventing an article edit. 7. Publish the main site with `npm run deploy:main`. Verify the affected public URLs after deployment. 8. Run `npm run search:notify -- https://onequill.dev/resources/article-slug` for the changed canonical pages, including updated collections. With no URLs, the command submits the current main sitemap's URLs. Preview the list first with `--dry-run`. The command checks that the verification file is live and records the receipt locally; submission does not confirm indexing. ## Evidence policy Use one of security-advisory, incident-report, release-fix or documented-behaviour. A release correction is not automatically a security advisory. A vendor claim is attributed to the vendor, not represented as an independent certification. Record source-specific version and identifier differences in identifierNotes. Where CVE mappings conflict, retain the primary GHSA as the finding identifier. Never merge uncertain ranges or count uncertain mappings as separate flaws. Scope findings to their prerequisites; do not infer managed-service exposure from an open-source package notice. Provider profiles are alphabetical and descriptive. Selected case totals do not measure vendor safety. The directory is a researched snapshot rather than an exhaustive inventory of everything available. OneQuill develops OneVir. Preserve this affiliation and the evidence limits in every relevant article. Local source review does not equal a release audit or a freshly executed test. ## Corrections Edit the affected research record and article together. Add a dated corrections entry to the article manifest (date and text), explain the changed identifier, scope or recommendation, and retain the appropriate original publication date. Build and inspect only the affected resource behaviour. To report a correction, use the [public support route](https://onequill.dev/support) and identify the page, primary source and proposed correction. Source updates can change applicability after the last review. ## Search discovery The builder generates HTML available without JavaScript, canonical metadata, Article and collection schema, sitemap entries, RSS, public Markdown and llms.txt links. Social images are original diagrams. Reader links open formatted HTML. Links labelled **Download Markdown**, **Markdown worksheet** or **research records (JSON)** intentionally provide reusable files. Agents can start with [the document index](https://onequill.dev/agents.json), [full website text](https://onequill.dev/llms-full.txt) or [the discovery guide](https://onequill.dev/llms.txt), without browser automation. The main [sitemap](https://onequill.dev/sitemap.xml) is listed in robots.txt and submitted through the verified Google Search Console and Bing Webmaster Tools properties. Resubmit it after a substantial publication when appropriate. Google recommends [Search Console, its API or robots.txt for sitemap discovery](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap). The local `search:notify` command uses [IndexNow](https://www.indexnow.org/documentation) to notify Bing and other participating engines of changed canonical URLs after deployment. HTTP 200 means received; HTTP 202 means received with key validation pending. This command does not submit to Google. Keep publication dates, modification dates and source-review dates distinct, and assess actual impressions and indexing in the search consoles. [Google's AI search guidance](https://developers.google.com/search/docs/appearance/ai-features) says no special AI optimisation is required and inclusion is not guaranteed. Maintain useful, accessible, original and well-sourced content; do not promise traffic or use fabricated author expertise or review scores. --- Canonical: https://onequill.dev/resources/sglang-worker-and-management-trust # SGLang: isolate worker channels and model-management paths A protected public chat route can coexist with execution-sensitive worker or management interfaces. The useful boundary is which actor can reach each interface and supply executable or serialization-sensitive material. ## Which deployment does this concern? CVE-2026-3059 and CVE-2026-3060 concern unsafe deserialization in enabled worker paths reachable by an attacker. CVE-2026-3989 instead requires replaying a malicious dump. The later July findings concern different endpoints and optional settings. **Affected scope:** The March worker issues require multimodal generation or encoder parallel disaggregation and reachable affected channels. The July notice describes separate feature/configuration conditions without one unified version range. ## Response and remediation evidence CERT/CC’s 7 April update identifies 0.5.10 for the March findings; upstream PR #20904 records the replay-dump change. The July notice reported no patches at its publication and no vendor statement in that note. Current documentation describes API and admin keys, but documentation alone does not establish remediation of every historical path. **Remediation record:** March CERT/CC update: 0.5.10, with conflicting CVE-2026-3059 metadata. A fixed release for the July group is not established by the cited July notice. ## What clients can learn Keep worker channels isolated, separate inference from administration, approve model and adapter sources, and verify the exact release and enabled endpoint inventory. Treat patch statements for different findings independently. Keep an endpoint inventory, the effective configuration and the running artifact together. A configuration or model change should trigger a review of the boundary it changes. Assign an owner to the evidence and to any required action. ## Questions for your provider - Can an untrusted caller reach the worker or disaggregation ports? - Which model-update, adapter, dumper and replay paths are enabled? - Do API and admin credentials protect the intended endpoint matrix? - Which artifact or commit proves each selected finding is remediated? Use the [inference guide](/resources/llm-inference-security-and-operations) and [evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record the answer in your deployment context. ### Technical detail: source scope and identifiers The March and July notices describe different findings. Their conditions and patch status cannot be merged into a claim about every SGLang deployment or the current release. - CERT/CC VU#665416 covers CVE-2026-3059, CVE-2026-3060 and CVE-2026-3989. Its April update calls 0.5.10 a remediation, while the CVE-2026-3059 record has listed 0.5.10 as affected; retain both statements and request artifact-specific confirmation. - The separate 30 July VU#281278 covers CVE-2026-15969, CVE-2026-15971, CVE-2026-15974, CVE-2026-15976, CVE-2026-15977 and CVE-2026-15978. No current fixed version is inferred from that historical notice. ## Primary and supporting records - [CERT/CC VU#665416 · March findings and April update](https://kb.cert.org/vuls/id/665416) - [CERT/CC VU#281278 · separate July findings](https://kb.cert.org/vuls/id/281278) - [CVE-2026-3059 · source-specific version record](https://www.cve.org/CVERecord?id=CVE-2026-3059) - [Upstream replay-dump fix, PR #20904](https://github.com/sgl-project/sglang/pull/20904) - [Current API and administration settings](https://docs.sglang.io/docs/advanced_features/server_arguments) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/vllm-gguf-gpu-memory-isolation # vLLM: a GGUF kernel defect is distinct from cache policy In a shared GPU environment, uninitialised output can expose data from earlier GPU tensors. This is a kernel-memory correctness concern, distinct from whether a prefix cache is intentionally shared across tenants. ## What deployment does this concern? **Component and scope:** Integer truncation in specific GGUF dequantisation kernels. **Affected versions / scope:** Vendor header: ≥ 0.5.5; affected kernel path and tensor shape are prerequisites. **Vendor remediation:** Current vendor patched-version field: ≥ 0.24.0 The model and execution path use the affected GGUF dequantisation kernels with dimensions that encounter integer truncation. The defect can leave part of an allocated output tensor uninitialised. ## Vendor response and practical action The current vendor advisory lists 0.24.0 as patched. Upgrade the affected runtime and evaluate model-format and tenant-isolation choices. Do not replace the current primary patch statement with an older registry value. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Per-request cache salting addresses cache reuse and timing inference between trust groups. It does not repair this kernel defect or establish process and GPU memory isolation by itself. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Are GGUF models and the affected kernels used? - Which tenants share a process, GPU and memory allocator? - Is the deployed version consistent with the current vendor patch statement? - What isolation evidence covers kernels separately from prefix-cache policy? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [Running LLM inference in production: security, isolation and capacity](/resources/llm-inference-security-and-operations) for the wider evaluation context. ### Technical detail: evidence and identifier limits The advisory is not a general finding that every user's KV cache is left unwiped. It concerns particular dequantisation kernels and partially uninitialised output tensors. - Primary identifier: GHSA-5jv2-g5wq-cmr4; vendor mapping CVE-2026-53923. - An older registry patch value differs; the current primary advisory's ≥ 0.24.0 statement is retained. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2026-06-11**. These dates do not change merely because this article is rebuilt. ## Primary sources - [GHSA-5jv2-g5wq-cmr4 · vendor advisory](https://github.com/vllm-project/vllm/security/advisories/GHSA-5jv2-g5wq-cmr4) - [vLLM prefix-cache isolation](https://docs.vllm.ai/en/latest/design/prefix_caching/) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/vllm-kv-transfer-network-isolation # vLLM: isolate KV-transfer and distributed communication An internal cluster channel can carry execution-sensitive messages without being the public inference API. Restricting /v1 access alone does not control a service bound on a different address or inter-node port. ## What deployment does this concern? **Component and scope:** PyNcclPipe KV-cache transfer in the V0 engine. **Affected versions / scope:** 0.6.5 through 0.8.4 **Vendor remediation:** 0.8.5 The affected V0 PyNcclPipe transfer path receives untrusted serialised network messages. Network reachability to the relevant communication service and use of that backend define exposure; this is not a malicious-model-loading case. ## Vendor response and practical action The vendor identifies 0.8.5 as patched and discusses binding and network isolation. Current security guidance separately warns about insecure defaults in several distributed communication paths. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Map public API, model-management and cluster networks independently. Keep worker communication on trusted isolated networks and verify bind addresses, firewall rules and the backend actually in use. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Is the deployed engine V0 and is PyNcclPipe KV transfer enabled? - Which addresses and ports are reachable from untrusted networks? - What authentication and encryption apply to each distributed channel? - Which release and topology changes establish remediation? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [Running LLM inference in production: security, isolation and capacity](/resources/llm-inference-security-and-operations) for the wider evaluation context. ### Technical detail: evidence and identifier limits The current documentation must be read for the chosen distributed backend. It does not justify saying every vLLM deployment sends plaintext prompts, weights and outputs through ZeroMQ. - Primary identifier: GHSA-hjq4-87xh-g4fv; vendor mapping CVE-2025-47277. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2025-05-20**. These dates do not change merely because this article is rebuilt. ## Primary sources - [GHSA-hjq4-87xh-g4fv · vendor advisory](https://github.com/vllm-project/vllm/security/advisories/GHSA-hjq4-87xh-g4fv) - [vLLM security guidance](https://docs.vllm.ai/en/latest/usage/security/) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/vllm-model-loading-python-optimisation # vLLM: model configuration and Python optimisation A model bundle is part of the software supply chain. Configuration can influence executable behaviour, so authorisation to add a model is more powerful than permission to send an ordinary inference prompt. ## What deployment does this concern? **Component and scope:** Model-configuration loading when Python assertion checks are disabled. **Affected versions / scope:** < 0.22.0 **Vendor remediation:** ≥ 0.22.0 The operator loads malicious model configuration while Python runs with python -O or PYTHONOPTIMIZE=1. The advisory describes an activation-function import path whose assertion checks disappear in that mode. ## Vendor response and practical action The vendor identifies 0.22.0 as patched. Upgrade the affected loader, establish model provenance and review the Python execution environment before importing model assets. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Keep runtime optimisation and trust validation separate. vLLM's own performance or compilation levels are not the same setting as Python's -O flag. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Who can approve and load model configurations? - Is python -O or PYTHONOPTIMIZE used in the service environment? - Are model revision and configuration hashes recorded? - Which validation remains active under every supported runtime mode? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [Running LLM inference in production: security, isolation and capacity](/resources/llm-inference-security-and-operations) for the wider evaluation context. ### Technical detail: evidence and identifier limits The prerequisite is malicious model configuration plus Python optimisation. The advisory is not evidence that every ordinary inference request can execute code. - Primary identifier: GHSA-q8gq-377p-jq3r; vendor mapping CVE-2026-41523. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2026-06-14**. These dates do not change merely because this article is rebuilt. ## Primary sources - [GHSA-q8gq-377p-jq3r · vendor advisory](https://github.com/vllm-project/vllm/security/advisories/GHSA-q8gq-377p-jq3r) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/vllm-multimodal-input-validation # vLLM: input features need separate validation and capacity limits Schema validation and resource budgeting are different controls. A structurally accepted tensor can still need concurrency-safe checks, while a frame list can consume excessive decode and processing work before inference starts. ## What deployment does this concern? **Component and scope:** Optional prompt embeddings and base64 video/jpeg frame processing; separate findings collected in one feature-validation case. **Affected versions / scope:** Original embedding header: ≥ 0.10.2 and < 0.11.1. Follow-up header: ≥ 0.21.0. Video header: ≥ 0.7.0. These are source-specific statements. **Vendor remediation:** Original embeddings: 0.13.0. Concurrency follow-up: ≥ 0.26.0. Video frames: 0.19.0. Prompt-embedding advisories depend on enabling the feature; the follow-up explicitly requires --enable-prompt-embeds, which is off by default. The video issue concerns base64 video/jpeg frames rather than every binary video loader. ## Vendor response and practical action The upstream advisories list separate fixes. The follow-up demonstrates a concurrency-related invariant-check bypass but explicitly does not establish a live-server crash, unsafe conversion, GPU memory corruption or code execution. Treat the vendor notice as the starting point for an applicability decision. Identify the installed artifact and configuration, document whether the prerequisite exists, and assign an owner to any required change. A public advisory does not establish that your installation was exposed or that a managed service shares the same condition. ## What clients can learn Maintain an input-feature inventory. Limit decoded dimensions, frame counts, aggregate input size and concurrent work; disable unsupported input types rather than rely on the text-only request limit. A useful evaluation result connects a named control to evidence from the actual deployment. Keep the provider's statement, your effective configuration and a relevant demonstration together. If the result depends on a feature being disabled or a network being isolated, retain that fact with the version number so a later change triggers review. ## Questions to take to your provider - Are prompt embeddings and video input enabled on the public route? - Which limits apply after base64 decoding and tensor construction? - Are invariant checks isolated between concurrent requests? - What evidence confirms each separate patched path? Use the [six-page evaluation worksheet](/assets/downloads/ai-gateway-evaluation-worksheet.pdf) to record evidence, ownership and actions. Continue with [Running LLM inference in production: security, isolation and capacity](/resources/llm-inference-security-and-operations) for the wider evaluation context. ### Technical detail: evidence and identifier limits These findings share an evaluation theme but are not one vulnerability. The follow-up's unproven outcomes must not be reported as demonstrated RCE. - The original embedding advisory header lists ≥ 0.10.2 and < 0.11.1 while naming 0.13.0 as patched; do not silently widen the interval. - The upstream video advisory uses CVE-2026-34755. Red Hat describes the frame issue under CVE-2026-5497. The relationship is not counted as two independently confirmed flaws. Primary finding identifier remains GHSA-pq5c-rjhq-qp7p. Evidence label: **security advisory**. Source-review date: **2026-10-07**. Source publication or event date: **2026-07-27**. These dates do not change merely because this article is rebuilt. ## Primary sources - [Prompt embedding advisory · GHSA-mcmc-2m55-j8jj](https://github.com/vllm-project/vllm/security/advisories/GHSA-mcmc-2m55-j8jj) - [Concurrency follow-up · GHSA-pr7f-p5mw-fc87](https://github.com/vllm-project/vllm/security/advisories/GHSA-pr7f-p5mw-fc87) - [Video frames · GHSA-pq5c-rjhq-qp7p](https://github.com/vllm-project/vllm/security/advisories/GHSA-pq5c-rjhq-qp7p) - [Red Hat video-frame record · CVE-2026-5497](https://access.redhat.com/security/cve/cve-2026-5497) --- Published: 2026-10-07. Modified: 2026-10-07. Sources reviewed: 2026-10-07. Author: OneQuill Research. Affiliation: OneQuill develops OneVir. Documentary review; selected case totals are not vendor security rankings. --- Canonical: https://onequill.dev/resources/ai-gateway-providers # AI gateway provider directory: products, scope and evaluation questions · OneQuill Canonical: https://onequill.dev/resources/ai-gateway-providers Provider directory # Find your gateway. Know its role. Alphabetical profiles of specialist gateways, API platforms, cloud gateways, hosted routers, related products and inference engines. Deployment modes describe the cited offerings, not every available edition. Sources reviewed 2026-10-07 · By OneQuill Research Search providers CategoryAll categoriesSpecialist gatewaysHosted routersAPI platformsCloud gatewaysRelated productsInference enginesDeployment / evidenceAll optionsManagedSelf-hostedCustomer-hostedClear filters 55 profiles Specialist gateways ## agentgateway [#](https://onequill.dev/resources/ai-gateway-providers#agentgateway) - Self-hosted Rust gateway for AI and MCP. Local rate limits are held in memory; a remote rate-limit service supplies a different sharing boundary. [Product / project ↗](https://agentgateway.dev/)[Documentation / policy ↗](https://agentgateway.dev/docs/standalone/latest/documentation/configuration/resiliency/rate-limits/) ### Evaluation question & research context **Who can author policies and reference credentials across namespaces?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [agentgateway: a patched release can still need a policy setting](https://onequill.dev/resources/agentgateway-namespace-isolation) Hosted routers ## AI/ML API [#](https://onequill.dev/resources/ai-gateway-providers#aimlapi) - Managed Hosted model API. Service key management does not replace review of provider processing and retention terms. [Product / project ↗](https://aimlapi.com/)[Documentation / policy ↗](https://docs.aimlapi.com/api-references/service-endpoints/api-key-management) ### Evaluation question & research context **Can keys be scoped and revoked without interrupting all applications?** Also searched as: AIMLAPI, AI ML API. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## AIHubMix [#](https://onequill.dev/resources/ai-gateway-providers#aihubmix) - Managed Hosted model aggregation. Review the service policy alongside the actual selected upstream. [Product / project ↗](https://aihubmix.com/developers)[Documentation / policy ↗](https://docs.aihubmix.com/en/terms-and-privacy/Privacy) ### Evaluation question & research context **What provider, geography and content-retention terms govern this model?** Also searched as: AI Hub Mix. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. API platforms ## Apache APISIX [#](https://onequill.dev/resources/ai-gateway-providers#apisix) - Self-hosted AI plugins extend the API gateway. Documented streaming timeouts can close a stream without a DONE event. [Product / project ↗](https://apisix.apache.org/ai-gateway/)[Documentation / policy ↗](https://apisix.apache.org/docs/apisix/plugins/ai-proxy/) ### Evaluation question & research context **How does the client detect an incomplete stream and avoid unsafe replay?** Also searched as: APISIX. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Cloud gateways ## Apigee / Model Armor [#](https://onequill.dev/resources/ai-gateway-providers#apigee) - Managed Model Armor is a configured integration. Inspection limits, supported regions and service latency affect coverage. [Product / project ↗](https://docs.cloud.google.com/model-armor/model-armor-apigee-integration)[Documentation / policy ↗](https://docs.cloud.google.com/apigee/docs/api-platform/tutorials/using-model-armor-policies) ### Evaluation question & research context **What is the explicit fail-open or fail-closed policy for inspection errors?** Also searched as: Google Cloud, GCP. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Cloud gateways ## AWS Bedrock AgentCore Gateway [#](https://onequill.dev/resources/ai-gateway-providers#aws-agentcore) - Managed AgentCore Gateway now documents inference targets as well as tool integration. Policy-language and target support still have limits. [Product / project ↗](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway.html)[Documentation / policy ↗](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway-targets-inference.html) ### Evaluation question & research context **Which inference and tool targets are in scope for each policy?** Also searched as: AWS, Amazon, Bedrock. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Cloud gateways ## Azure API Management [#](https://onequill.dev/resources/ai-gateway-providers#azure-apim) - Managed - Self-hosted Token-limit counters are independent across specified gateway, region and workspace boundaries. A newer AI gateway preview has separate limitations. [Product / project ↗](https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities)[Documentation / policy ↗](https://learn.microsoft.com/en-us/azure/api-management/llm-token-limit-policy) ### Evaluation question & research context **Is your budget global, or is it enforced independently per region?** Also searched as: Microsoft, Azure, APIM. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## Bifrost [#](https://onequill.dev/resources/ai-gateway-providers#bifrost) - Self-hosted - Managed Gateway with MCP and custom-plugin administration. The September advisories distinguish dynamically linked builds from published static Docker images. [Product / project ↗](https://www.getbifrost.ai/)[Documentation / policy ↗](https://docs.getbifrost.ai/) ### Evaluation question & research context **Is the management API authenticated and isolated from application callers?** Also searched as: Maxim. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [Bifrost: management APIs are an execution boundary](https://onequill.dev/resources/bifrost-management-api-boundaries) Specialist gateways ## Braintrust AI Proxy [#](https://onequill.dev/resources/ai-gateway-providers#braintrust) - Managed - Customer-hosted Proxy and observability sit within Braintrust's architecture. Data-plane hosting and logging choices need to be assessed together. [Product / project ↗](https://www.braintrust.dev/)[Documentation / policy ↗](https://www.braintrust.dev/docs/platform/architecture) ### Evaluation question & research context **Where are request logs and evaluation datasets stored?** Also searched as: Braintrust proxy. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Cloud gateways ## Cloudflare AI Gateway [#](https://onequill.dev/resources/ai-gateway-providers#cloudflare) - Managed Account-level AI Gateway permissions can cover all gateways, including stored BYOK credentials. Logging behaviour changed for new customers from 24 September 2026. [Product / project ↗](https://developers.cloudflare.com/ai-gateway/)[Documentation / policy ↗](https://developers.cloudflare.com/ai-gateway/configuration/authentication/) ### Evaluation question & research context **How are tenants separated and which logging generation applies to your account?** Also searched as: Cloudflare. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## CometAPI [#](https://onequill.dev/resources/ai-gateway-providers#cometapi) - Managed Hosted model aggregator. Published privacy claims need a deployment-specific contractual scope. [Product / project ↗](https://www.cometapi.com/about/)[Documentation / policy ↗](https://www.cometapi.com/privacy-policy/) ### Evaluation question & research context **Which upstreams receive content and which contractual exceptions apply?** Also searched as: Comet API. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Cloud gateways ## Databricks AI Gateway [#](https://onequill.dev/resources/ai-gateway-providers#databricks) - Managed Unity-oriented gateway services and legacy Mosaic endpoints have different controls. Inference-table payload capture is a separate configuration. [Product / project ↗](https://docs.databricks.com/aws/en/ai-gateway/)[Documentation / policy ↗](https://docs.databricks.com/gcp/en/ai-gateway/model-services) ### Evaluation question & research context **Which service generation and payload-capture settings are enabled?** Also searched as: Unity AI Gateway, Mosaic. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## Eden AI [#](https://onequill.dev/resources/ai-gateway-providers#eden) - Managed Hosted aggregation across AI services. The data-processing agreement and provider exceptions matter alongside security descriptions. [Product / project ↗](https://www.edenai.co/)[Documentation / policy ↗](https://www.edenai.co/dpa) ### Evaluation question & research context **Does the DPA cover each selected service and subprocessor?** Also searched as: Eden AI. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## Envoy AI Gateway / Agent Router [#](https://onequill.dev/resources/ai-gateway-providers#envoy-ai-gateway) - Self-hosted Project naming and repository have moved towards Agent Router. Advisories retain the Envoy AI Gateway name and component scope. [Product / project ↗](https://aigateway.envoyproxy.io/)[Documentation / policy ↗](https://github.com/theagentrouter/agent-router) ### Evaluation question & research context **Which release and MCP request-processing path are actually deployed?** Also searched as: Envoy, Agent Router. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [Envoy AI Gateway: one message, two interpretations](https://onequill.dev/resources/envoy-ai-gateway-mcp-message-smuggling) - [Envoy AI Gateway: enforce limits before buffering](https://onequill.dev/resources/envoy-ai-gateway-mcp-request-limits) API platforms ## F5 NGINX Gateway Fabric [#](https://onequill.dev/resources/ai-gateway-providers#f5-nginx) - Self-hosted AI Guardrails integration is a configured stack feature, not a property of every NGINX proxy. Verify policy acceptance and request-size handling. [Product / project ↗](https://docs.nginx.com/nginx-gateway-fabric/)[Documentation / policy ↗](https://docs.nginx.com/nginx-gateway-fabric/how-to/f5-ai-guardrails/) ### Evaluation question & research context **Which guardrail integration and enforcement status are active?** Also searched as: NGINX, F5. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## FastRouter [#](https://onequill.dev/resources/ai-gateway-providers#fastrouter) - Managed Hosted router with routing and privacy descriptions. Confirm the effective configuration rather than infer it from feature names. [Product / project ↗](https://fastrouter.ai/features)[Documentation / policy ↗](https://fastrouter.ai/privacy) ### Evaluation question & research context **What happens to data and cost when a route falls back?** Also searched as: Fast Router. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Related products ## GitLab AI Gateway [#](https://onequill.dev/resources/ai-gateway-providers#gitlab) - Managed - Self-hosted Application-specific gateway for GitLab Duo, not a general interchangeable model router. Hosted and self-hosted patch responsibilities differ. [Product / project ↗](https://docs.gitlab.com/administration/gitlab_duo/)[Documentation / policy ↗](https://docs.gitlab.com/releases/patches/other-patches/patch-release-gitlab-ai-gateway-19-4-1-released/) ### Evaluation question & research context **Which Duo access and flow-template capabilities are exposed?** Also searched as: GitLab Duo. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [GitLab AI Gateway: flow templates and sandbox assumptions](https://onequill.dev/resources/gitlab-ai-gateway-template-sandbox) API platforms ## Gravitee [#](https://onequill.dev/resources/ai-gateway-providers#gravitee) - Self-hosted - Managed AI policies and masking are edition and flow-order dependent. Payload logging introduces its own memory and retention needs. [Product / project ↗](https://www.gravitee.io/platform/ai-gateway)[Documentation / policy ↗](https://documentation.gravitee.io/apim/create-and-configure-apis/apply-policies/policy-reference/data-logging-masking) ### Evaluation question & research context **Is masking applied before every logger and export in this edition?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## Helicone [#](https://onequill.dev/resources/ai-gateway-providers#helicone) - Managed - Self-hosted Observability can capture requests automatically. Retention and self-hosting depend on the selected service and configuration. [Product / project ↗](https://www.helicone.ai/)[Documentation / policy ↗](https://docs.helicone.ai/getting-started/quick-start) ### Evaluation question & research context **Which content fields are logged, for how long, and under whose account?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. API platforms ## Higress [#](https://onequill.dev/resources/ai-gateway-providers#higress) - Self-hosted Gateway with AI provider and security-guard plugins. Plugin selection and policy order define the effective behaviour. [Product / project ↗](https://higress.io/en/ai-gateway)[Documentation / policy ↗](https://github.com/higress-group/higress/security) ### Evaluation question & research context **Which request and response paths are covered by the selected plugins?** Also searched as: Alibaba Higress. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## Hugging Face Inference Providers [#](https://onequill.dev/resources/ai-gateway-providers#huggingface) - Managed Hugging Face's own request handling and metadata retention differ from upstream provider policies. Dedicated Inference Endpoints are a separate service. [Product / project ↗](https://huggingface.co/docs/inference-providers/en/index)[Documentation / policy ↗](https://huggingface.co/docs/inference-providers/en/security) ### Evaluation question & research context **Which provider receives the request, and are its retention terms approved?** Also searched as: HF, HuggingFace. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. API platforms ## IBM API Connect [#](https://onequill.dev/resources/ai-gateway-providers#ibm-api-connect) - Managed - Self-hosted AI support and limits vary between SaaS and software versions, including watsonx integrations. [Product / project ↗](https://www.ibm.com/docs/en/api-connect/cloud/saas?topic=applications-using-ai-gateway-support-watsonxai-apis)[Documentation / policy ↗](https://www.ibm.com/docs/en/api-connect/software/12.1.1?topic=overview-known-limitations) ### Evaluation question & research context **Which AI features and limitations apply to your deployment and release?** Also searched as: IBM, API Connect. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. API platforms ## kgateway [#](https://onequill.dev/resources/ai-gateway-providers#kgateway) - Self-hosted Kubernetes gateway with AI extensions. Envoy-based documentation is versioned; prompt guards and rate limits need configured policies. [Product / project ↗](https://kgateway.dev/)[Documentation / policy ↗](https://kgateway.dev/docs/envoy/2.1.x/ai/about/) ### Evaluation question & research context **What happens if an external guard or rate-limit service is unavailable?** Also searched as: K Gateway. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. API platforms ## Kong AI Gateway [#](https://onequill.dev/resources/ai-gateway-providers#kong) - Managed - Self-hosted Traditional AI Proxy plugins and newer AI Gateway offerings have distinct version and deployment scopes. [Product / project ↗](https://developer.konghq.com/index/ai-gateway/)[Documentation / policy ↗](https://developer.konghq.com/plugins/ai-proxy/changelog/) ### Evaluation question & research context **How is streaming usage reconciled against provider invoices?** Also searched as: Kong, AI Proxy. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [Kong: a streaming usage fix deserves billing regression tests](https://onequill.dev/resources/kong-gemini-streaming-token-accounting) Specialist gateways ## LangDB [#](https://onequill.dev/resources/ai-gateway-providers#langdb) - Managed - Customer-hosted AI gateway with observability. Documented ClickHouse TTL deletion runs through asynchronous merges. [Product / project ↗](https://langdb.ai/)[Documentation / policy ↗](https://docs.langdb.ai/enterprise/resources/configuring-data-retention/) ### Evaluation question & research context **What deletion latency applies to logs, backups and exports?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## Leanroute [#](https://onequill.dev/resources/ai-gateway-providers#leanroute) - Managed Policy dated 1 October 2026 distinguishes database storage from response and semantic caches, which default to 15 minutes; no-persistence is a separate choice. [Product / project ↗](https://leanroute.dev/)[Documentation / policy ↗](https://leanroute.dev/privacy) ### Evaluation question & research context **Are response caching and semantic caching disabled for sensitive routes?** Also searched as: Lean Route. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## LiteLLM [#](https://onequill.dev/resources/ai-gateway-providers#litellm) - Self-hosted - Managed Python proxy and SDK distribution are distinct from the vendor's pinned proxy Docker distribution. [Product / project ↗](https://www.litellm.ai/)[Documentation / policy ↗](https://docs.litellm.ai/docs/proxy/prod) ### Evaluation question & research context **What package or image digest is running, and how are credentials rotated after an incident?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [LiteLLM: package provenance and incident response](https://onequill.dev/resources/litellm-march-2026-package-incident) Hosted routers ## LLM Gateway [#](https://onequill.dev/resources/ai-gateway-providers#llmgateway) - Managed - Self-hosted Metadata-only logging is described as the default; payload logging is an opt-in setting. Upstream retention remains separate. [Product / project ↗](https://llmgateway.io/)[Documentation / policy ↗](https://llmgateway.io/legal/privacy) ### Evaluation question & research context **Which content logging options and upstream terms are enabled?** Also searched as: LLMGateway, theopenco. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Inference engines ## LM Studio [#](https://onequill.dev/resources/ai-gateway-providers#lm-studio) - Self-hosted Desktop and headless model runtime with an API server. Authentication is optional; network serving and LM Link alter the access and execution boundary. [Product ↗](https://lmstudio.ai/)[API-token authentication ↗](https://lmstudio.ai/docs/developer/core/authentication)[Network server settings ↗](https://lmstudio.ai/docs/developer/core/server/serve-on-network)[LM Link remote execution ↗](https://lmstudio.ai/docs/developer/core/lmlink)[Offline operation ↗](https://lmstudio.ai/docs/app/offline) ### Evaluation question & research context **Are API-token permissions enforced, which machine executes each model, and what model-loading and eviction behaviour does the application rely on?** Also searched as: LMStudio, LM Studio llmster, llmster, lms. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [LM Studio: choose authentication, execution location and model lifecycle](https://onequill.dev/resources/lm-studio-network-authentication-and-lifecycle) Specialist gateways ## Lunar.dev [#](https://onequill.dev/resources/ai-gateway-providers#lunar) - Self-hosted - Managed Gateway and traffic-management offering. Production logs and data-plane placement need their own review. [Product / project ↗](https://www.lunar.dev/product/ai-gateway)[Documentation / policy ↗](https://docs.lunar.dev/api-gateway/lunar-dev-in-production/lunar-logs) ### Evaluation question & research context **Which data leaves the gateway through logs or management integrations?** Also searched as: Lunar. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## Martian [#](https://onequill.dev/resources/ai-gateway-providers#martian) - Managed Hosted routing gateway. Authentication documentation establishes an API boundary, not every enterprise data-control claim. [Product / project ↗](https://withmartian.com/)[Documentation / policy ↗](https://gateway-docs.withmartian.com/api-reference/authentication) ### Evaluation question & research context **What route decision evidence and data-processing terms can be supplied?** Also searched as: With Martian. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## MLflow AI Gateway [#](https://onequill.dev/resources/ai-gateway-providers#mlflow) - Self-hosted Current gateway documentation distinguishes development credential-encryption defaults from production key management and SQL-backed rotation. [Product / project ↗](https://mlflow.org/docs/latest/genai/governance/ai-gateway)[Documentation / policy ↗](https://www.mlflow.org/docs/latest/genai/governance/ai-gateway/api-keys/key-rotation/) ### Evaluation question & research context **Is a production encryption secret configured and is rotation rehearsed?** Also searched as: MLflow. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## New API [#](https://onequill.dev/resources/ai-gateway-providers#new-api) - Self-hosted Project with quota and billing functions. Release-candidate version scope matters in the quota-overflow advisory. [Product / project ↗](https://docs.newapi.pro/)[Documentation / policy ↗](https://docs.newapi.pro/en/docs/guide/wiki/changelog) ### Evaluation question & research context **Can credits, reservations and final charges be reconciled under concurrent requests?** Also searched as: QuantumNous, new-api. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [New API: quota arithmetic is part of the trust boundary](https://onequill.dev/resources/new-api-quota-billing-overflow) Specialist gateways ## nexos.ai [#](https://onequill.dev/resources/ai-gateway-providers#nexos) - Managed Hosted gateway and AI platform. Procurement needs the service contract, provider list and configured routing policy. [Product / project ↗](https://nexos.ai/ai-gateway/)[Documentation / policy ↗](https://nexos.ai/legal/privacy-policy/) ### Evaluation question & research context **What geography and retention commitments cover every fallback?** Also searched as: Nexos. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Related products ## Not Diamond [#](https://onequill.dev/resources/ai-gateway-providers#not-diamond) - Managed - Customer-hosted Routing and model-selection API. It is listed as a related product rather than assumed to implement a full gateway boundary. [Product / project ↗](https://docs.notdiamond.ai/docs/what-is-not-diamond)[Documentation / policy ↗](https://docs.notdiamond.ai/docs/privacy-security-and-local-deployments) ### Evaluation question & research context **Which controls reside in your calling application and which in the routing service?** Also searched as: NotDiamond. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Inference engines ## Ollama [#](https://onequill.dev/resources/ai-gateway-providers#ollama) - Self-hosted - Managed Local model runtime and API, with optional cloud-model features. The local API does not require authentication; cloud API credentials and cloud processing are a separate mode. [Project ↗](https://ollama.com/)[Local and cloud API authentication ↗](https://docs.ollama.com/api/authentication)[Network, cloud and concurrency settings ↗](https://docs.ollama.com/faq)[Context and memory ↗](https://docs.ollama.com/context-length) ### Evaluation question & research context **Is execution local or cloud, who can reach inference and model-management endpoints, and how do context, parallel requests and model residency fit memory?** Also searched as: Ollama local server, Ollama Cloud. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [Ollama: model import is a separate memory and access boundary](https://onequill.dev/resources/ollama-model-import-memory-boundary) Specialist gateways ## One API [#](https://onequill.dev/resources/ai-gateway-providers#one-api) - Self-hosted Separate project from New API. Model mappings and protocol reconstruction can affect unsupported fields. [Product / project ↗](https://github.com/songquanpeng/one-api)[Documentation / policy ↗](https://github.com/songquanpeng/one-api) ### Evaluation question & research context **Which request fields survive translation, including tools and usage options?** Also searched as: songquanpeng. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## OneVir [#](https://onequill.dev/resources/ai-gateway-providers#onevir) - Self-hosted OneQuill's own product combines gateway controls and local execution paths. This profile is affiliated and is not an independent security assessment. [Product / project ↗](https://onevir.onequill.dev/)[Documentation / policy ↗](https://onevir.onequill.dev/#capabilities) ### Evaluation question & research context **Which controls are enabled in your installed version, and which remain the inference backend's responsibility?** Also searched as: OneQuill. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Affiliation: OneQuill develops OneVir. [Own implementation evidence and limits](https://onequill.dev/resources/onevir-control-evidence). Hosted routers ## OpenRouter [#](https://onequill.dev/resources/ai-gateway-providers#openrouter) - Managed Provider routing, fallback and endpoint-level privacy controls need to be configured together. ZDR and geography are route properties. [Product / project ↗](https://openrouter.ai/)[Documentation / policy ↗](https://openrouter.ai/docs/guides/features/guardrails/overview) ### Evaluation question & research context **Can every selected endpoint and fallback satisfy the same privacy constraints?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## OpenZiti LLM Gateway [#](https://onequill.dev/resources/ai-gateway-providers#openziti) - Self-hosted Identity-based overlay networking is relevant to gateway access. Platform network features do not establish every application-level safeguard. [Product / project ↗](https://blog.openziti.io/ai-secops-why-your-ai-infrastructure-has-a-network-shaped-blind-spot)[Documentation / policy ↗](https://openziti.io/docs/learn/introduction/features/) ### Evaluation question & research context **How are service identities issued, revoked and separated between tenants?** Also searched as: Ziti. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## Opper [#](https://onequill.dev/resources/ai-gateway-providers#opper) - Managed Security overview updated 5 October 2026: upstream inference is not EEA-restricted by default. Tracing, backups and provider retention have separate rules. [Product / project ↗](https://opper.ai/)[Documentation / policy ↗](https://opper.ai/security-overview) ### Evaluation question & research context **Does the selected route meet geography and retention needs, including backups?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## orq.ai [#](https://onequill.dev/resources/ai-gateway-providers#orq) - Managed - Customer-hosted Gateway within an orchestration platform. Vendor security claims and logging retention must be matched to the purchased deployment. [Product / project ↗](https://orq.ai/platform/ai-gateway)[Documentation / policy ↗](https://orq.ai/legal/security) ### Evaluation question & research context **Which attestations and retention settings cover this service and edition?** Also searched as: Orq. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## Plano [#](https://onequill.dev/resources/ai-gateway-providers#plano) - Self-hosted Agent proxy project formerly named Arch. Establish the current component boundaries and supported deployment before comparing controls. [Product / project ↗](https://planoai.dev/)[Documentation / policy ↗](https://docs.planoai.dev/get_started/overview.html) ### Evaluation question & research context **Which tool, model and policy paths pass through Plano?** Also searched as: Arch, Arch Gateway. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## Portkey / Prisma AIRS AI Gateway [#](https://onequill.dev/resources/ai-gateway-providers#portkey) - Self-hosted - Managed Portkey commercial offerings have moved into Prisma AIRS. The selected advisory applies to the open-source Portkey gateway; it does not establish managed-service exposure. [Product / project ↗](https://portkey.ai/features/ai-gateway)[Documentation / policy ↗](https://www.paloaltonetworks.com/blog/2026/07/announcing-general-availability-of-prisma-airs-ai-gateway/) ### Evaluation question & research context **Which product, build and custom-host policy are you evaluating?** Also searched as: Portkey AI, Prisma AIRS. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [Portkey: custom-host routing and private-network access](https://onequill.dev/resources/portkey-custom-host-ssrf) Hosted routers ## Requesty [#](https://onequill.dev/resources/ai-gateway-providers#requesty) - Managed Gateway logging, caching and upstream retention are separate. EU router processing does not establish the upstream inference location. [Product / project ↗](https://www.requesty.ai/)[Documentation / policy ↗](https://www.requesty.ai/privacy) ### Evaluation question & research context **Does this plan and route prohibit training and retention at every layer?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Inference engines ## SGLang [#](https://onequill.dev/resources/ai-gateway-providers#sglang) - Self-hosted LLM and multimodal serving framework with distributed workers and prefix-cache reuse. The inference API, administrative controls and worker channels have distinct trust boundaries. [Project ↗](https://github.com/sgl-project/sglang)[Server arguments and access controls ↗](https://docs.sglang.io/docs/advanced_features/server_arguments)[Session-aware radix cache ↗](https://github.com/sgl-project/sglang/blob/main/docs/docs/advanced_features/session_radix_cache.mdx) ### Evaluation question & research context **Which API and admin keys, worker networks, model-loading features and cache isolation are configured in the installed release?** Also searched as: SGLang Runtime, SGLang serving framework, sgl-project. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [SGLang: isolate worker channels and model-management paths](https://onequill.dev/resources/sglang-worker-and-management-trust) Specialist gateways ## TensorZero [#](https://onequill.dev/resources/ai-gateway-providers#tensorzero) - Self-hosted Gateway authentication is an explicit operational configuration with PostgreSQL support. Dashboard access is a separate boundary. [Product / project ↗](https://www.tensorzero.com/)[Documentation / policy ↗](https://www.tensorzero.com/docs/operations/set-up-auth-for-tensorzero) ### Evaluation question & research context **Are both the gateway and dashboard authenticated for this deployment?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## Tetrate Agent Router [#](https://onequill.dev/resources/ai-gateway-providers#tetrate) - Managed - Customer-hosted Hosted Agent Router Service and enterprise customer-hosted data plane have different data flows. The management plane is separately hosted. [Product / project ↗](https://www.tetrate.io/)[Documentation / policy ↗](https://docs.tetrate.ai/product-architecture/data-flows/) ### Evaluation question & research context **What crosses between your data plane and the hosted management plane?** Also searched as: Agent Router Service, TARS, Tetrate Enterprise. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. API platforms ## Traefik Hub [#](https://onequill.dev/resources/ai-gateway-providers#traefik) - Managed - Self-hosted Hub AI middlewares include Redis-backed shared token limits. Interrupted streams and response-completed events affect usage measurement. [Product / project ↗](https://doc.traefik.io/traefik-hub/ai-gateway/overview)[Documentation / policy ↗](https://doc.traefik.io/traefik-hub/ai-gateway/middlewares/token-rate-limit) ### Evaluation question & research context **Does incomplete-stream usage reach the budget ledger?** Also searched as: Traefik. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Specialist gateways ## TrueFoundry [#](https://onequill.dev/resources/ai-gateway-providers#truefoundry) - Managed - Customer-hosted Gateway, traces and deployment platform. Residency-aware fallback is a vendor-described configuration, not proof of your route policy. [Product / project ↗](https://www.truefoundry.com/ai-gateway)[Documentation / policy ↗](https://www.truefoundry.com/docs/ai-gateway/feedback-for-traces) ### Evaluation question & research context **Can the vendor show that all retries and fallback targets satisfy your residency policy?** Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. API platforms ## Tyk AI Studio / MCP Gateway [#](https://onequill.dev/resources/ai-gateway-providers#tyk) - Managed - Self-hosted AI Studio, chat and MCP offerings are separate capabilities in the Tyk ecosystem. [Product / project ↗](https://tyk.io/tyk-mcp-gateway/)[Documentation / policy ↗](https://tyk.io/blog/introducing-tyk-ai-studio-welcome-to-the-future-of-ai-governance-powered-by-ai-chat-gateway-and-portal/) ### Evaluation question & research context **Which model-routing and tool-access controls are available in the purchased product?** Also searched as: Tyk AI Studio, Tyk MCP. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## Vercel AI Gateway [#](https://onequill.dev/resources/ai-gateway-providers#vercel) - Managed Gateway content deletion and provider ZDR arrangements have separate scope and eligibility. [Product / project ↗](https://vercel.com/ai-gateway)[Documentation / policy ↗](https://vercel.com/i/secure-ai-gateway) ### Evaluation question & research context **Which providers and plan-specific agreements implement the required ZDR setting?** Also searched as: Vercel. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Inference engines ## vLLM [#](https://onequill.dev/resources/ai-gateway-providers#vllm) - Self-hosted Inference engine and serving stack. API authentication, internal networks, model loading and GPU scheduling are separate boundaries from a gateway. [Product / project ↗](https://docs.vllm.ai/)[Documentation / policy ↗](https://docs.vllm.ai/en/latest/usage/security/) ### Evaluation question & research context **Which engine, features, model format and cluster topology are deployed?** Also searched as: vLLM Production Stack. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. - [vLLM: isolate KV-transfer and distributed communication](https://onequill.dev/resources/vllm-kv-transfer-network-isolation) - [vLLM: model configuration and Python optimisation](https://onequill.dev/resources/vllm-model-loading-python-optimisation) - [vLLM: input features need separate validation and capacity limits](https://onequill.dev/resources/vllm-multimodal-input-validation) - [vLLM: a GGUF kernel defect is distinct from cache policy](https://onequill.dev/resources/vllm-gguf-gpu-memory-isolation) API platforms ## WSO2 API Manager [#](https://onequill.dev/resources/ai-gateway-providers#wso2) - Self-hosted - Managed AI token policies and backend throttling have different enforcement paths. Local per-node limits differ from distributed Traffic Manager policies. [Product / project ↗](https://apim.docs.wso2.com/en/latest/ai-gateway/rate-limiting/)[Documentation / policy ↗](https://apim.docs.wso2.com/en/latest/api-design-manage/design/rate-limiting/protect-backend-services/) ### Evaluation question & research context **Which counters are shared across replicas and regions?** Also searched as: WSO2, APIM. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. Hosted routers ## ZenMux [#](https://onequill.dev/resources/ai-gateway-providers#zenmux) - Managed Turning off Data Services changes ZenMux logging and related features. It does not independently establish upstream zero retention. [Product / project ↗](https://zenmux.ai/)[Documentation / policy ↗](https://zenmux.ai/docs/guide/advanced/data-services.html) ### Evaluation question & research context **Which data services are disabled and what provider retention remains?** Also searched as: Zen Mux. Evidence: documented behaviour · Sources reviewed 2026-10-07. These links establish the cited scope, not an independent audit of every feature. No case was selected for this profile. This does not establish the presence or absence of vulnerabilities. ### No matches. Change a filter or clear your search to see more records. How to read this research ## Evidence with a defined scope. Product coverage is a researched snapshot, not an exhaustive inventory. Category and deployment mode describe the cited offering; editions, contracts and configuration can change the scope. Selected cases illustrate evaluation questions. Advisory counts are not security rankings. **Security advisory**A published security finding, with prerequisites and remediation evidence. **Incident report**A reported event and response, attributed to its source. **Release fix**A correction recorded in release notes; not automatically a CVE. **Documented behaviour**A design choice, default or limit from official documentation. Published by OneQuill, developer of OneVir. This is a documentary review of selected public sources, not an independent vendor security audit. Published 2026-10-07 · Modified 2026-10-07 · Sources reviewed 2026-10-07. Report corrections through [OneQuill support](https://onequill.dev/support), with the page and primary source. [Download research records (JSON)](https://onequill.dev/assets/downloads/ai-gateway-research.json) · [Publishing method and templates](https://onequill.dev/resources/resource-publishing). A practical next step ## Take the questions into your review. Six A4 pages. Record the provider, version, configuration, evidence, owner and action for each topic. [Download worksheet PDF ↓](https://onequill.dev/assets/downloads/ai-gateway-evaluation-worksheet.pdf)[Markdown worksheet](https://onequill.dev/assets/downloads/ai-gateway-evaluation-worksheet.md) --- Canonical: https://onequill.dev/resources/ai-gateway-case-studies # AI gateway and inference case studies: conditions and remediation · OneQuill Canonical: https://onequill.dev/resources/ai-gateway-case-studies Case library # Real cases. Useful questions. Selected advisories, incidents, release fixes and documented behaviours explain deployment conditions and responses. Counts reflect editorial selection, not vendor security rankings. Sources reviewed 2026-10-07 · By OneQuill Research Search cases CategoryAll categoriesSecurityReliabilityBudgets & scalingSupply chainInference operationsDeployment / evidenceAll optionsSecurity advisoryRelease fixIncident reportDocumented behaviourClear filters ProviderAll providersagentgatewayBifrostEnvoy AI Gateway / Agent RouterGitLab AI GatewayKong AI GatewayLiteLLMLM StudioNew APIOllamaPortkey / Prisma AIRS AI GatewaySGLangvLLM16 cases Security advisoryPortkey / Prisma AIRS AI Gateway ## [Portkey: custom-host routing and private-network access](https://onequill.dev/resources/portkey-custom-host-ssrf) Open-source Portkey gateway custom-host routing **Remediation record:** 1.14.0 Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/portkey-custom-host-ssrf) Security advisoryBifrost ## [Bifrost: management APIs are an execution boundary](https://onequill.dev/resources/bifrost-management-api-boundaries) Reachable management APIs with authentication disabled; stdio MCP registration and remote custom-plugin loading **Remediation record:** Both advisories identify 2.1.0. Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/bifrost-management-api-boundaries) Security advisoryEnvoy AI Gateway / Agent Router ## [Envoy AI Gateway: one message, two interpretations](https://onequill.dev/resources/envoy-ai-gateway-mcp-message-smuggling) MCP JSON-RPC parsing in the Envoy AI Gateway / Agent Router project **Remediation record:** 0.6.0 Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/envoy-ai-gateway-mcp-message-smuggling) Security advisoryEnvoy AI Gateway / Agent Router ## [Envoy AI Gateway: enforce limits before buffering](https://onequill.dev/resources/envoy-ai-gateway-mcp-request-limits) MCP POST-body buffering in the external-processing component **Remediation record:** 1.0.0 Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/envoy-ai-gateway-mcp-request-limits) Security advisoryagentgateway ## [agentgateway: a patched release can still need a policy setting](https://onequill.dev/resources/agentgateway-namespace-isolation) Cross-namespace backend references authored by Kubernetes namespace administrators **Remediation record:** 1.3.0 plus AGW_BACKEND_REF_GRANT_MODE=route-and-policy Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/agentgateway-namespace-isolation) Security advisoryNew API ## [New API: quota arithmetic is part of the trust boundary](https://onequill.dev/resources/new-api-quota-billing-overflow) Quota settlement in affected New API release candidates **Remediation record:** ≥ 1.0.0-rc.18 Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/new-api-quota-billing-overflow) Release fixKong AI Gateway ## [Kong: a streaming usage fix deserves billing regression tests](https://onequill.dev/resources/kong-gemini-streaming-token-accounting) Kong AI Proxy plugin handling of Gemini streaming usage **Remediation record:** 3.16.0.0 includes the stated fix. Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/kong-gemini-streaming-token-accounting) Incident reportLiteLLM ## [LiteLLM: package provenance and incident response](https://onequill.dev/resources/litellm-march-2026-package-incident) Malicious PyPI distributions, distinct from the vendor's official pinned Proxy Docker images **Remediation record:** Vendor describes a clean 1.83 release and pipeline v2 on 30 March; incident response also requires containment and credential review. Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/litellm-march-2026-package-incident) Security advisoryGitLab AI Gateway ## [GitLab AI Gateway: flow templates and sandbox assumptions](https://onequill.dev/resources/gitlab-ai-gateway-template-sandbox) GitLab Duo Agent Platform flow-template processing **Remediation record:** 19.2.4, 19.3.2 and 19.4.1 Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/gitlab-ai-gateway-template-sandbox) Security advisoryvLLM ## [vLLM: isolate KV-transfer and distributed communication](https://onequill.dev/resources/vllm-kv-transfer-network-isolation) PyNcclPipe KV-cache transfer in the V0 engine **Remediation record:** 0.8.5 Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/vllm-kv-transfer-network-isolation) Security advisoryvLLM ## [vLLM: model configuration and Python optimisation](https://onequill.dev/resources/vllm-model-loading-python-optimisation) Model-configuration loading when Python assertion checks are disabled **Remediation record:** ≥ 0.22.0 Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/vllm-model-loading-python-optimisation) Security advisoryvLLM ## [vLLM: input features need separate validation and capacity limits](https://onequill.dev/resources/vllm-multimodal-input-validation) Optional prompt embeddings and base64 video/jpeg frame processing; separate findings collected in one feature-validation case **Remediation record:** Original embeddings: 0.13.0. Concurrency follow-up: ≥ 0.26.0. Video frames: 0.19.0. Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/vllm-multimodal-input-validation) Security advisoryvLLM ## [vLLM: a GGUF kernel defect is distinct from cache policy](https://onequill.dev/resources/vllm-gguf-gpu-memory-isolation) Integer truncation in specific GGUF dequantisation kernels **Remediation record:** Current vendor patched-version field: ≥ 0.24.0 Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/vllm-gguf-gpu-memory-isolation) Security advisorySGLang ## [SGLang: isolate worker channels and model-management paths](https://onequill.dev/resources/sglang-worker-and-management-trust) Feature-dependent worker serialization, replay tooling and administrative/model-loading interfaces **Remediation record:** March CERT/CC update: 0.5.10, with conflicting CVE-2026-3059 metadata. A fixed release for the July group is not established by the cited July notice. Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/sglang-worker-and-management-trust) Security advisoryOllama ## [Ollama: model import is a separate memory and access boundary](https://onequill.dev/resources/ollama-model-import-memory-boundary) GGUF tensor-size validation during model creation and quantization **Remediation record:** Registry patched-version field: 0.17.1. Confirm inclusion of the upstream tensor-size fix in the deployed artifact. Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/ollama-model-import-memory-boundary) Documented behaviourLM Studio ## [LM Studio: choose authentication, execution location and model lifecycle](https://onequill.dev/resources/lm-studio-network-authentication-and-lifecycle) Documented API authentication, network binding, remote model resolution and JIT model residency **Deployment control:** Enable required authentication and scoped token permissions for shared access; approve bind addresses and remote devices; configure model residency intentionally. Reviewed 2026-10-07[Read case →](https://onequill.dev/resources/lm-studio-network-authentication-and-lifecycle) ### No matches. Change a filter or clear your search to see more records. How to read this research ## Evidence with a defined scope. Product coverage is a researched snapshot, not an exhaustive inventory. Category and deployment mode describe the cited offering; editions, contracts and configuration can change the scope. Selected cases illustrate evaluation questions. Advisory counts are not security rankings. **Security advisory**A published security finding, with prerequisites and remediation evidence. **Incident report**A reported event and response, attributed to its source. **Release fix**A correction recorded in release notes; not automatically a CVE. **Documented behaviour**A design choice, default or limit from official documentation. Published by OneQuill, developer of OneVir. This is a documentary review of selected public sources, not an independent vendor security audit. Published 2026-10-07 · Modified 2026-10-07 · Sources reviewed 2026-10-07. Report corrections through [OneQuill support](https://onequill.dev/support), with the page and primary source. [Download research records (JSON)](https://onequill.dev/assets/downloads/ai-gateway-research.json) · [Publishing method and templates](https://onequill.dev/resources/resource-publishing). A practical next step ## Take the questions into your review. Six A4 pages. Record the provider, version, configuration, evidence, owner and action for each topic. [Download worksheet PDF ↓](https://onequill.dev/assets/downloads/ai-gateway-evaluation-worksheet.pdf)[Markdown worksheet](https://onequill.dev/assets/downloads/ai-gateway-evaluation-worksheet.md)