A gateway can enforce requests per minute while still overspending a monthly budget. The request count, token estimate, provider bill and usable inference capacity measure different things. Establish the units and boundaries before choosing limits.
Begin with a budget policy that names the tenant, provider, model, billing period and currency. Define whether the limit is an estimate at admission, an actual cost after completion or both. Decide what happens when usage is missing rather than silently treating it as free.
Find the counter's sharing boundary
Azure API Management's token-limit policy describes independent counters across gateway, region and workspace boundaries. WSO2's backend throttling documentation distinguishes local per-node limits from distributed enforcement.
agentgateway's rate-limit documentation distinguishes local in-memory state from a remote shared service. Traefik Hub's token middleware documents Redis-backed sharing. None of these descriptions alone establishes that your entire multi-region spend is capped.
Record the authoritative counter store, update consistency, restart behaviour and failure policy. Ask how admitted work is reconciled when a node exits before settlement.
Account for concurrent work and arithmetic
The New API quota case concerns integer overflow in release-candidate quota settlement. It illustrates why a billing boundary includes numeric ranges, sign handling and state transitions.
Before dispatch, reserve a defensible bound based on supported input and output limits. After completion, settle actual usage and release unused reservation. Ask how repeated events, cancellation and concurrent admissions are handled. A balance check followed by a separate update can allow several requests to consume the same remaining allowance.
Streaming changes how usage arrives
The Kong Gemini case is a release fix for running usage metadata, not a CVE. Cumulative totals must not be added as if each chunk were independent usage. Missing metadata is a different state from a reported zero.
A useful demonstration includes a complete stream, an interrupted stream, a retry and a fallback. Compare gateway records with provider-side evidence. Include reasoning, cached input and other billable units when supported by the provider; do not assume every provider uses identical fields.
Technical detail: capacity and billable attempts
Limit logical requests and upstream attempts separately. A single client request may cause more than one billed dispatch. A client disconnect does not necessarily stop upstream work immediately. Define how reservations cover each permitted attempt and when they expire.
For self-hosted inference, use the inference operations guide. Record model revision, quantisation, context length, input/output distributions, hardware, parallelism and accepted latency. Requests per second from a different workload are not a useful budget guarantee.
Questions for your evaluation
Complete the budgets page of the worksheet. Keep the counter topology, cost model, concurrency result and provider reconciliation together. Assign ownership for price changes and missing-usage investigations.
Applying this to OneVir
OneVir documents provider budgets and application limits. The implementation evidence record identifies a focused concurrent-reservation test: eight competing attempts are presented with a two-attempt allowance, and the test asserts two admissions and the matching reserved amount.
That reviewed source supports a specific reservation behaviour. It is not a test run performed for this publication or proof that every provider price, streaming protocol and distributed deployment is correct. Request evidence for the configured route and settlement path.