Home › Guides › Direct vs gateway

Calling Veo through Google's platform or through a gateway: what changes in operations

Updated 2026-10-02

This is not a "which is better" page. Both routes call the same family of models. What differs is who you authenticate with, who holds your quota, what a bill looks like and what happens when a backend is slow. If you are deciding how to wire Veo into a product, these are the operational questions that tend to matter after launch rather than at the prototype stage. Where this page describes Google's side, it is deliberately general, because details change; the checklist at the end tells you what to confirm in Google's own documentation.

Authentication and credentials

Direct access to a large cloud platform usually means cloud-native identity: a service account or workload identity, short-lived tokens, and permissions granted through roles. That is excellent for security teams and awkward for a script on a laptop or a service hosted on another cloud, where you must decide how credentials get there and how they rotate.

Through VideoRouter the credential is a bearer token (Authorization: Bearer llmr_sk_live_...). Keys can carry their own monthly spend cap, scopes and a model allow-list, so you can hand a narrow key to one service. A leaked key is revoked and replaced without touching any cloud IAM. The tradeoff is that you are trusting a bearer secret, so store it like a database password.

Quotas and rate limits

Direct, quotas are generally defined per project and region and are managed in that platform's console, often with a request-and-wait process for increases. You plan capacity against one vendor's limits for one model family.

Through a gateway, you hit two kinds of limit: your key's own RPM/TPM and spend cap, which you control, and upstream capacity, which the gateway can route around by trying another host that serves the same model. A request that exceeds your key limit returns 429 with a Retry-After header; an exhausted balance returns 402. Neither requires a support ticket.

Billing and invoicing

TopicDirect to GoogleThrough a gateway
Who invoices youGoogle, on the billing account attached to your projectVideoRouter, from a prepaid credit balance
UnitPer Google's pricing for the model; confirm in Google's docsPer requested second, billed once at creation; polling is free
Failed jobsCheck Google's termsJobs that fail upstream are not billed
Spend limitsBudgets and alerts you configure on the accountPer-key monthly caps enforced at request time
Platform feeNone beyond Google's rate2% on image and video usage

The fee row is real: if you only ever use Google's own endpoint at volume, the gateway adds a cost and you should weigh it against what you gain in the other rows. If you also use other models, one prepaid balance and one invoice trail replace several.

Failover and availability

Veo is served by Google and, per the live catalog, by additional hosts. Calling Google's platform directly gives you Google's availability and nothing else: if a region has an incident or you are throttled, your job fails or waits. Through the gateway, an unpinned request goes to the cheapest healthy host and can fall back to another host serving the same checkpoint. A suffix such as google/veo-3.1-fast/<host> is a soft preference that can still fall back; a hard pin requires provider.only plus allow_fallbacks: false. The documentation page on provider selection spells out the exact request shape.

Fallback has a cost: different hosts can differ on supported inputs, resolutions and price, and an unsupported resolution is ignored rather than rejected. If you need byte-for-byte consistent behaviour, hard-pin one host and accept the loss of failover.

One key versus many integrations

Most products that use Veo also use something else: another video model for a different style, an image model for start frames, a speech model for narration. Direct integration means one SDK, auth flow and error model per vendor. The gateway's error envelope is OpenAI-style, {"error": {"message", "type", "code"}}, for every model, and the job lifecycle (queued, in_progress, completed or failed) is identical. Swapping Veo for another model is a change to the model string:

curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{"model": "google/veo-3.1-fast", "prompt": "a paper boat on a rain-soaked street", "duration_secs": 5}'

Where direct is the better call

Checklist to verify in Google's documentation

  1. Current authentication options and which suits your runtime.
  2. Default quotas and the process to raise them, per region.
  3. Which Veo variants and resolutions your region exposes.
  4. How and for how long generated files are stored, and how you retrieve them.
  5. Pricing units, including what is charged when a job fails.

Compare the answers with the live host table, read the no-Google-Cloud walkthrough for the setup you would skip, or create a key and send one test job.

Frequently asked questions

Does using a gateway change which Veo model I get?

The model ids map to the same Veo variants, served by Google or by other hosts. Check the model page for which hosts serve a variant and any input differences between them.

How do spend controls differ from direct access?

A gateway key can carry its own monthly cap and model allow-list enforced per request. On Google's platform you configure budgets and alerts on the billing account; confirm the details in Google's docs.

What happens if one Veo host is unavailable?

Unpinned requests can fall back to another host serving the same checkpoint. A host suffix is only a soft preference; a hard pin needs provider.only with allow_fallbacks: false.

When should I call Google directly instead?

When your organisation is standardised on Google's cloud, needs contractual or residency terms with Google, or generates volume where a percentage platform fee outweighs the convenience.

Keep reading

Using Veo is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →