# AI gateways and inference: client evaluation worksheet

Published: 2026-10-07. OneQuill Research. OneQuill develops OneVir.
Record evidence as verified, partial, missing or not applicable, with a reason. This worksheet supports evaluation and does not certify a product.

## 1. Gateway security

Evaluate the authority of each interface: application inference, administration, tool execution and outbound networking. A single API-key setting rarely describes all four.

- Provider / product:
- Version / artifact:
- Deployment / topology:
- Enabled features / configuration:
- Reviewer / date:

### Which identities can invoke inference, change routes and install tools?

- Evidence / source / result:
- Owner:
- Action / due date:

### Can an application request reach a private or unapproved destination?

- Evidence / source / result:
- Owner:
- Action / due date:

### Is tool policy applied to the exact forwarded message?

- Evidence / source / result:
- Owner:
- Action / due date:

### Which patch and configuration evidence establishes the boundary?

- Evidence / source / result:
- Owner:
- Action / due date:

[Guide and sources](https://onequill.dev/resources/ai-gateway-security)

---

## 2. Privacy and data handling

Privacy is a property of the complete route and its retained copies. Gateway processing, upstream inference, observability and backups can have different terms.

- Provider / product:
- Version / artifact:
- Deployment / topology:
- Enabled features / configuration:
- Reviewer / date:

### Where do prompts, outputs, files and metadata travel?

- Evidence / source / result:
- Owner:
- Action / due date:

### Do all upstream endpoints and fallback routes meet approved terms?

- Evidence / source / result:
- Owner:
- Action / due date:

### What remains in logs, caches, exports and backups, and for how long?

- Evidence / source / result:
- Owner:
- Action / due date:

### How are deletion and retention changes verified?

- Evidence / source / result:
- Owner:
- Action / due date:

[Guide and sources](https://onequill.dev/resources/ai-gateway-privacy)

---

## 3. Budgets and scaling

A rate limit is not a global spend ceiling. Budgets require well-defined counters, concurrent admission, final settlement and a plan for incomplete usage.

- Provider / product:
- Version / artifact:
- Deployment / topology:
- Enabled features / configuration:
- Reviewer / date:

### Which counters are shared across replicas, regions and tenants?

- Evidence / source / result:
- Owner:
- Action / due date:

### Is maximum admitted spend reserved atomically before dispatch?

- Evidence / source / result:
- Owner:
- Action / due date:

### How are retries, partial streams and missing usage settled?

- Evidence / source / result:
- Owner:
- Action / due date:

### Which workload and latency target establish usable capacity?

- Evidence / source / result:
- Owner:
- Action / due date:

[Guide and sources](https://onequill.dev/resources/ai-gateway-budgets-and-scaling)

---

## 4. Reliability

A reliable gateway has a defined response to overload, partial output and upstream failure. Fallback must preserve policy and the client's understanding of the result.

- Provider / product:
- Version / artifact:
- Deployment / topology:
- Enabled features / configuration:
- Reviewer / date:

### Where are size, concurrency and deadline limits enforced?

- Evidence / source / result:
- Owner:
- Action / due date:

### How does the client distinguish complete and partial output?

- Evidence / source / result:
- Owner:
- Action / due date:

### Do retries and fallback preserve privacy, cost and tool policy?

- Evidence / source / result:
- Owner:
- Action / due date:

### What recovery evidence covers gateway, dependency and worker failure?

- Evidence / source / result:
- Owner:
- Action / due date:

[Guide and sources](https://onequill.dev/resources/ai-gateway-reliability)

---

## 5. Supply chain

Security claims should identify the artifact, dependency set and model assets actually executed. A project name and a version label alone do not establish provenance.

- Provider / product:
- Version / artifact:
- Deployment / topology:
- Enabled features / configuration:
- Reviewer / date:

### Can the running artifact be traced to an approved build and digest?

- Evidence / source / result:
- Owner:
- Action / due date:

### Are dependencies, plugins and model assets pinned and reviewed?

- Evidence / source / result:
- Owner:
- Action / due date:

### Which distribution channels and configurations does each advisory cover?

- Evidence / source / result:
- Owner:
- Action / due date:

### Who owns incident containment, secret rotation and safe restoration?

- Evidence / source / result:
- Owner:
- Action / due date:

[Guide and sources](https://onequill.dev/resources/ai-gateway-supply-chain)

---

## 6. Inference security and operations

An inference engine has its own API, internal networks, model loader, GPU allocator and scheduler. Gateway policy supports access control but cannot repair a vulnerable kernel or guarantee backend capacity.

- Provider / product:
- Version / artifact:
- Deployment / topology:
- Enabled features / configuration:
- Reviewer / date:

### Which API endpoints, input features and internal ports are exposed?

- Evidence / source / result:
- Owner:
- Action / due date:

### Who approves model assets, and how are tenants isolated?

- Evidence / source / result:
- Owner:
- Action / due date:

### What workload demonstrates acceptable TTFT, token latency and cache pressure?

- Evidence / source / result:
- Owner:
- Action / due date:

### How do worker recovery and cold or warm starts behave on this topology?

- Evidence / source / result:
- Owner:
- Action / due date:

[Guide and sources](https://onequill.dev/resources/llm-inference-security-and-operations)

---
