What the data shows

- Finding 1 - Location matters more than the route- Direct was faster from Singapore and Mumbai. OpenRouter was often faster from Montreal. - Singapore and Mumbai: direct 621–664 ms, routes 782–1,348 ms

- Montreal: routes 422–692 ms, direct 732 ms

- Finding 2 - DeepSeek's direct API was more consistent- DeepSeek's API varied less between locations: its slowest city was 36% above its fastest, against 89–219% through OpenRouter. - City spread: 36% direct, 89–219% via OpenRouter

- Slow days: 0 direct, 1 via OpenRouter

- Finding 3 - Most of the wait isn't the network- Connection setup was only 11–19 ms. Most latency came after reaching the provider. - Network: 11–19 ms

- First token: 462–883 ms

Time to first token, day by day

Does your location matter?

DeepSeek direct was faster from Singapore and Mumbai in our tests.

Where does the wait happen?

How much of the wait a closer server could remove, and how much is the provider’s own queue and model.

Direct vs OpenRouter

9–22 September 2026 · six test locations

How we test

We send the same streaming prompt from six cities to DeepSeek directly and through OpenRouter, then measure time to first visible token.

Same request on both sides: thinking off, temperature 0. The Fireworks and BaseTen routes are left out: one throttled us, the other never returned a visible token.

Full testing methodology →Good to know

- Network setup is the hop to OpenRouter's edge (11–19 ms from every city); the hop from OpenRouter to the provider's machines is counted as waiting for the model, which is why Mumbai still waits longer than Montreal.

- Direct sends deepseek-v4-flash, which DeepSeek has served with V4.1 Flash since 10 September (V4 Flash is retired), while the OpenRouter routes pin deepseek-v4-flash-0731, so the two sides have not been the same model since that day.

- Each OpenRouter request pins one provider with fallbacks off, and a city's run is discarded if any response came from another provider.

- The prompt is "Say 'ok' and nothing else." (max_tokens 256, temperature 0, thinking off), sent daily between 03:30 and 06:46 UTC, 5 requests per city after one discarded warm-up.

- We time the first visible token only; tokens per second, long prompts and thinking on are not measured.

How does your API compare?

Test your endpoint from the same 6 cities used here and see your response time next to these numbers.