LOADING0
Same models. Same output.Same models. Same output.
One endpoint. A fractionOne endpoint. A fraction
of the cost.of the cost.
Powered by Google models
Nano BananaNano Banana
VeoVeo
OmniOmni
GeminiGemini

The APIThe API
CostCost
ProblemProblem

01/03
01

Teams pay full retail API rates on every request to Google Cloud and Vertex AI. Costs scale linearly as you grow — with no discount in sight.

02

Self-hosting your own inference layer means huge upfront cost, a DevOps team to maintain it, and no real access to Google's models.

03

An all-or-nothing decision: overpay for the API, or build infrastructure you never wanted. Puzzle is the third option — same models, fraction of the cost.

THE CORE PROBLEM
SAME
MODELS,
SAME
OUTPUT,
JUST
A
SMALLER
BILL
SAVINGS
SAVINGS
MECHANICS

OneOne
EndpointEndpoint

Keep the same Google APIs, SDKs, prompts and application logic. Update a single endpoint and you're done — no migration, no rewrites.

PIPELINEFOUR STEPS / ONE POST01PURCHASECLIENT BUYS A GENERATION INSIDEYOUR PRODUCT — UI, BOT, MARKETPLACE.YOU SET THE PRICE.011010011010011010011010011010011010011010011010011010011010011010011010011010011010011002APIONE POST /V1/GENERATIONS.IDEMPOTENT. ASYNCHRONOUS.SAME GOOGLE SDKS AND PROMPTS.011010011010011010011010011010011010011010011010011010011010011010011010011010011010100103GEMINILOAD-BALANCED ACROSS GEMINIMODELS. RETRIES. VALIDATION.UP TO 30,000 RPM.011010011010011010011010011010011010011010011010011010011010011010011010011010011010011004WEBHOOKHMAC-SIGNED WEBHOOK.PRESIGNED S3 LINK TO THE RESULT.DELIVERED TO YOUR CALLBACK.NANO BANANA PRO - 1024x1024READY — DELIVERED TO YOUR CALLBACK

OurOur
NetworkNetwork

We negotiate volume rates with Google and pass the savings to you. Post-paid billing. Built for production traffic — up to 30,000 RPM and 99.99% uptime.

NETWORKVOLUME RATESGEMINIFLASHNANOBANANAVEO 3RAW CAPACITY / 1-5S DISPATCHROUTERETRYQUEUEVALIDATETHROUGHPUTUP TO 30,000 RPMUPTIME99.99%POST-PAID BILLINGBUILT FOR PRODUCTION TRAFFIC. PAY AFTER USAGE —NO PREPAY, NO LOCK-IN.LIVE TRAFFIC
WHY PUZZLE

WhyWhy
PuzzlePuzzle
WinsWins

Lower costs
1
Lower costs
SIGNIFICANTLY BELOW GOOGLE RETAIL RATES. CONTACT US TO LEARN MORE.
Zero migration
2
Zero migration
SAME APIS, SDKS, PROMPTS. JUST CHANGE BASE URL
Post-payment first
3
Post-payment first
NO PRE-FUNDING. OVERDRAFT LIMIT UP TO $500K
Production scale
4
Production scale
30,000 RPM / 99.99% UPTIME / AUTO-RETRIES
Pay your way
5
Pay your way
CRYPTO PREFERRED, BANK TRANSFERS ACCEPTED
COVERAGE

EVERYEVERY
GOOGLEGOOGLE
MODEL, ONE APIMODEL, ONE API

token

NANO BANANA

token

VEO

token

OMNI

token

GEMINI

FAQ

FrequentlyFrequently
AskedAsked
QuestionsQuestions

Register in under 2 minutes, get your API key, and swap the base URL in your existing SDK. You are ready to go with zero configuration.

Unlike traditional prepaid services, Puzzle operates on a post-payment model. Every verified developer account starts with an overdraft limit of up to $500K, ensuring your production code never runs out of credits or drops requests during usage spikes.

No. Puzzle is 100% compatible with existing Google Gemini APIs and SDKs. You simply replace the base URL/endpoint in your configuration and leave the rest of your application logic, prompts, and parameters untouched.

Puzzle is built for high-throughput production workloads. If Google or Vertex AI encounters rate limits or brief outages, our infrastructure automatically retries the request up to 10 times with exponential backoff. We maintain a 99.99% uptime guarantee.

By default, verified accounts support up to 30,000 RPM (Requests Per Minute) and 3 million requests per day. For custom enterprise pools, this limit can be scaled to 100,000+ RPM.