Gemini-compatible generateContent
When to use this page
Use this page when your client expects the Gemini generateContent API shape.
Base URL
https://api.quotaflow.ai
Endpoint
POST https://api.quotaflow.ai/v1beta/models/{model}:generateContent
Supported models
Text, multimodal, and reasoning-capable Gemini models:
gemini-3.5-flashgemini-3.1-pro-previewgemini-3-flash-previewgemini-2.5-progemini-2.5-flashgemini-2.5-flash-litegemini-3.1-flash-lite
Reasoning-capable models support Gemini-style thinking controls where the model family exposes them. Use generationConfig.maxOutputTokens when you need a hard output cap. When it is omitted, Quotaflow sends a normal 1024-token default so the request has a bounded execution and billing contract.
Use Gemini-native generationConfig.thinkingConfig for Gemini thinking controls. Do not send a top-level thinking field to Gemini-compatible endpoints; Quotaflow returns a clear 400 invalid_request_error for unsupported parameter shapes instead of retrying the request on another execution path.
Protocol surfaces
Quotaflow supports Gemini text models on three customer-facing surfaces when your key/package enables the model:
- Gemini-compatible:
:generateContent,:streamGenerateContent?alt=sse, and:countTokens. Use this surface for Gemini SDKs, cached-content workflows, and token counting. - OpenAI-compatible chat:
/openai/v1/chat/completionswith the same Gemini model ids andstream: truefor SSE. Use this surface when your app already speaks Chat Completions. Quotaflow preserves the Chat Completions request and response shape on this path instead of requiring Gemini-native fields. - OpenAI-compatible responses:
/openai/v1/responseswith the same Gemini model ids andstream: truefor SSE. Use this surface when your app already speaks the Responses API. Quotaflow serves it by projecting the request onto the same Chat Completions cells (no upstream speaks the Responses protocol for Gemini), and rewrites the answer back into the Responses shape; see the Responses page for the field mapping. This surface is stateless:storemust befalseif you send it at all, andprevious_response_id,conversation,promptandbackgroundare refused with a400, because Quotaflow keeps no response store to read them back from. Send the whole turn ininputeach time.
Cached-content and countTokens requests require a Gemini-compatible route. The two OpenAI-compatible Gemini surfaces are for generation, not Gemini native cache management.
Capability availability and retries
Gemini availability is evaluated per model, protocol surface, and capability. Seeing a model in authenticated /models confirms key access, not that every combination of plain generation, streaming, function tools, hosted tools, image, document, audio, and video input is currently enabled.
The status depends on whether a candidate could be planned at all. A valid request naming a model with no published cell on the surface you called plans nothing, and that is a 404 — measured, not inferred: every one of the Gemini requests recorded on /responses over the seven days to 2026-09-21 was answered 404, before the projection below existed. A request that does plan candidates and then finds none of them serving returns 503. An unknown model returns 404 too, and malformed or protocol-incompatible parameters return 400. Retry only temporary 429, 503, and network failures; do not retry 400 request-shape errors, or a 404, on another model.
Which surfaces have supply today. All three surfaces above are served from 2026-09-22. The native :generateContent one and /openai/v1/chat/completions are served by cells of their own (three cells on Google's own OpenAI-compatible endpoint, published after their bootstrap captures landed). /openai/v1/responses is served by projection onto those chat cells: Google's OpenAI-compatible endpoint has no /responses route and no relay Quotaflow buys from serves Gemini on that wire, so Quotaflow rewrites a Responses request into a Chat Completions request at the edge and rewrites the answer back. The per-capability tables in the model dossiers describe cells, so they list gemini.chat.*; there is no gemini.responses.* capability because there is no cell on that wire, and that is not a gap in measurement.
The google/ prefix is accepted for compatible Gemini model ids when your key/package enables that alias. Call your model discovery endpoint to confirm what your key can use.
Gemini image models such as gemini-3.1-flash-image-preview return images on :generateContent in Google's native shape. Until more sizes are measured they are served at aspectRatio 1:1 with imageSize 1K (1024x1024), and :streamGenerateContent is not available for them. The accepted fields are listed on the Images API page.
Curl
curl "https://api.quotaflow.ai/v1beta/models/gemini-3.8-flash:generateContent" \
-H "x-goog-api-key: $QUOTAFLOW_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents": [
{
"role": "user",
"parts": [
{ "text": "Summarize the benefits of an API gateway in three bullet points." }
]
}
]
}'
Quotaflow returns a Gemini-compatible response for Gemini-style clients.