Skip to content

Rankings

Fast because of how it routes. Cheap because we take nothing per token.

Two things decide what an inference request costs you in time and money: which route it takes, and what the gateway adds on top. SurgeX races the first and adds nothing to the second.

0%

Platform fee per token

you pay the provider’s published rate

Most gateways take a percentage of every token you spend, forever. SurgeX takes its fee once, when you add credit, and never touches the per token price again. At any real volume that difference compounds.

2

Routes raced per request

the second starts before the first fails

A gateway that waits for an error has already cost you the timeout. SurgeX starts a second route the moment the first misses its own observed median for first token, and kills the loser before it bills you anything.

0

Switches after the first byte

never, by design

Failover happens before a single character reaches you. Splicing two providers into one answer corrupts it, and a corrupted answer is worse than a slow one, so once the stream starts SurgeX commits to it.

Fastest routes

Median time to first token, measured on this deployment’s own traffic. A route appears here once it has served enough requests to have a median, and not before.

Latency fills in from live traffic.

These are measured numbers, not vendor claims, so the table stays empty until this deployment has served enough requests for a median to mean something.

Cheapest routes

Each provider’s own published rate per million tokens. This is the price you pay: SurgeX adds nothing to it.

#ModelProviderInput / MOutput / M
1GPT-4.1 nanoopenai$0.10$0.40
2GPT-4.1 miniopenai$0.40$1.60
3o4-miniopenai$1.10$4.40
4GPT-4.1openai$2.00$8.00