Adding another provider does not automatically make a service reliable. A fallback may exceed a budget, disclose data to a different region or repeat a tool action. Define an acceptable result and failure response before selecting the routing policy.
Build a small failure matrix: upstream timeout, overload, authentication failure, guardrail outage, client disconnect, gateway restart and inference-worker loss. For each event, record whether work was dispatched, whether partial output was returned and whether retry is safe.
Put limits before resource consumption
The Envoy MCP request-limit case concerns reading a full body before limits were applied. A rejection after allocation can arrive too late to protect component memory.
Choose body-size, decoded-content, concurrency and deadline limits together. Include compression, base64 expansion, image dimensions and video frames when those features are available. Measure which component retains the body and how cancellation releases it. The vLLM input-validation case shows why input features need separate bounds.
Make incomplete output visible
Apache APISIX's AI Proxy documentation describes timeout behaviour that may close a stream without a DONE event. A network connection ending is therefore not universally proof of successful completion.
The caller needs an explicit completion signal, cancellation state and partial-output policy. Decide whether a partially generated answer can be used, must be marked incomplete or must be discarded. Never replay an action-bearing workflow blindly after its visible response is lost.
Keep fallback inside the approved boundary
List every permitted fallback and the conditions for choosing it. Reapply identity, model permissions, data geography, spend admission and tool restrictions immediately before dispatch. Record the route decision so support staff can distinguish “provider failed” from “policy refused.”
A circuit breaker can reduce repeated calls to an unhealthy provider. Its thresholds, error classification and recovery trial still need review. A provider health check does not necessarily exercise the same model, quota or tool path as a real request.
Technical detail: failure domains and recovery
Gateway replicas, rate-limit storage, authentication services, policy services and inference workers can fail independently. Write down what state each holds and what happens when it restarts. A local counter may reset, an open circuit may be forgotten and a stream may have no surviving owner.
For distributed inference, vLLM documents Ray and multiprocessing options. Ray Serve's fault-tolerance guidance describes recovery mechanisms with the appropriate KubeRay deployment. Avoid assuming either automatic recovery or a mandatory whole-cluster manual restart without inspecting the setup.
Questions for your evaluation
Use the reliability page of the worksheet. Capture rejection behaviour, completion semantics, permitted replay conditions and observed recovery. Choose tests from the failure matrix that directly affect your workflow.
Applying this to OneVir
OneVir documents ordered fallbacks and provider circuit breakers. The implementation evidence record identifies the provider-health state machine and routing check: down providers are skipped in fallback chains, with cooldown and trial behaviour.
This source review does not demonstrate the recovery time of your installation. Ask for a controlled upstream-failure and interrupted-stream demonstration with the configured policy, and assess separately the availability of local inference workers and external dependencies.