Quotaflow
llms.txtOpenAPIDashboard
Protocol Surfaces

Gemini-compatible generateContent

When to use this page

Use this page when your client expects the Gemini generateContent API shape.

Base URL

https://api.quotaflow.ai

Endpoint

POST https://api.quotaflow.ai/v1beta/models/{model}:generateContent

Supported models

Text, multimodal, and reasoning-capable Gemini models:

Reasoning-capable models support Gemini-style thinking controls where the model family exposes them. Use generationConfig.maxOutputTokens when you need a hard output cap. When it is omitted, Quotaflow sends a normal 1024-token default so the request has a bounded execution and billing contract.

Use Gemini-native generationConfig.thinkingConfig for Gemini thinking controls. Do not send a top-level thinking field to Gemini-compatible endpoints; Quotaflow returns a clear 400 invalid_request_error for unsupported parameter shapes instead of retrying the request on another execution path.

Protocol surfaces

Quotaflow supports Gemini text models on three customer-facing surfaces when your key/package enables the model:

Cached-content and countTokens requests require a Gemini-compatible route. The two OpenAI-compatible Gemini surfaces are for generation, not Gemini native cache management.

Capability availability and retries

Gemini availability is evaluated per model, protocol surface, and capability. Seeing a model in authenticated /models confirms key access, not that every combination of plain generation, streaming, function tools, hosted tools, image, document, audio, and video input is currently enabled.

The status depends on whether a candidate could be planned at all. A valid request naming a model with no published cell on the surface you called plans nothing, and that is a 404 — measured, not inferred: every one of the Gemini requests recorded on /responses over the seven days to 2026-09-21 was answered 404, before the projection below existed. A request that does plan candidates and then finds none of them serving returns 503. An unknown model returns 404 too, and malformed or protocol-incompatible parameters return 400. Retry only temporary 429, 503, and network failures; do not retry 400 request-shape errors, or a 404, on another model.

Which surfaces have supply today. All three surfaces above are served from 2026-09-22. The native :generateContent one and /openai/v1/chat/completions are served by cells of their own (three cells on Google's own OpenAI-compatible endpoint, published after their bootstrap captures landed). /openai/v1/responses is served by projection onto those chat cells: Google's OpenAI-compatible endpoint has no /responses route and no relay Quotaflow buys from serves Gemini on that wire, so Quotaflow rewrites a Responses request into a Chat Completions request at the edge and rewrites the answer back. The per-capability tables in the model dossiers describe cells, so they list gemini.chat.*; there is no gemini.responses.* capability because there is no cell on that wire, and that is not a gap in measurement.

The google/ prefix is accepted for compatible Gemini model ids when your key/package enables that alias. Call your model discovery endpoint to confirm what your key can use.

Gemini image models such as gemini-3.1-flash-image-preview return images on :generateContent in Google's native shape. Until more sizes are measured they are served at aspectRatio 1:1 with imageSize 1K (1024x1024), and :streamGenerateContent is not available for them. The accepted fields are listed on the Images API page.

Curl

curl "https://api.quotaflow.ai/v1beta/models/gemini-3.8-flash:generateContent" \
  -H "x-goog-api-key: $QUOTAFLOW_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [
      {
        "role": "user",
        "parts": [
          { "text": "Summarize the benefits of an API gateway in three bullet points." }
        ]
      }
    ]
  }'

Quotaflow returns a Gemini-compatible response for Gemini-style clients.