Every model, routed to the fastest and cheapest gateway.
Hero Run runs one catalog across many inference gateways. When several serve the same model, we route to the cheapest and fail over to the next on error. Hero goes further: it reads your prompt and picks a right-sized model. These numbers come from real traffic and update continuously.
Gateway speed
Throughput and time-to-first-token (TTFT), measured from live streamed runs.
| Gateway | Runs | tok/s | TTFT |
|---|---|---|---|
| Loading… | |||
What Hero routes
Hero scores each prompt and sends it to a right-sized tier: simple work to a fast cheap model, hard reasoning to a frontier model, at one flat price.
TTFT is time to the first streamed token, measured on real runs; tok/s is effective throughput including latency. Sample sizes vary by gateway (see Runs). Numbers reflect the models actually routed through each gateway, not a fixed benchmark model.