Quotaflow
llms.txtOpenAPIDashboard
Quotaflow API

Models

Use this page when you need to verify model ids, understand endpoint support, and choose the right OpenAI-compatible model before configuring a client. For a product-style visual catalog, start with the Quotaflow Model Library.

Quotaflow Model Library
Quotaflow Model Library

What /models means

GET /models is scoped to the authenticated key and base URL you call. Quotaflow lists only public model ids enabled for that key. A default customer key can include multiple model families; a product-scoped key may list only GPT, embeddings, images, video, or another enabled subset.

Unsupported model ids fail fast instead of being silently changed to another public model id.

Quotaflow model discovery uses OpenRouter-style author/model names. The author namespace describes the model identity, not the upstream vendor selected for a request: openai/* covers GPT, Codex, GPT Image, and embeddings; anthropic/* covers Claude; google/* covers Gemini, Gemini image, and Veo; z-ai/*, x-ai/*, moonshotai/*, and bytedance/* cover GLM, Grok, Kimi, and Seedance. Routing, pricing, supply snapshots, and receipts use one internal canonical model id, so changing the public author prefix never creates or pins a vendor route.

Legacy unprefixed ids remain accepted as inbound compatibility aliases. New integrations should send the prefixed id returned by /models. Unknown or mismatched author namespaces fail closed. A Claude key exposes exactly one assigned identity namespace—anthropic/claude-*, anthropic-enterprise/claude-*, or aws-bedrock/claude-*—across both /v1/models and /openai/v1/models; an unprefixed Claude alias follows that assignment. The aws-bedrock/claude-* namespace is a Quotaflow Claude response projection, not an OpenRouter author namespace and not a supplier selector. The anthropic-enterprise/claude-* prefix selects an isolated enterprise line. Enterprise ids appear only for a key assigned to the enterprise line. Google Veo names remain hidden until an exact executable route, price, and capability receipt are available.

Endpoint

GET https://api.quotaflow.ai/openai/v1/models

Curl

curl https://api.quotaflow.ai/openai/v1/models \
  -H "Authorization: Bearer $QUOTAFLOW_API_KEY"

Curated model families

Model familyModel idsEndpointsAvailabilityCommon casesReference
OpenAI author aliasesopenai/gpt-5, openai/gpt-5-mini, openai/gpt-5-nano, openai/gpt-5.1, openai/gpt-5.4, openai/gpt-5.4-mini, openai/gpt-5.4-nano, openai/gpt-5.4-pro, openai/gpt-5.5, openai/gpt-5.6, openai/gpt-5.6-sol, openai/gpt-5.6-terra, openai/gpt-5.6-luna, openai/gpt-6-astra, openai/gpt-6-sol, openai/gpt-6-luna, openai/gpt-6.1-sol/responses, /chat/completionsKey-scoped; verify with authenticated /modelsOpenRouter-style OpenAI model ids over the same canonical GPT and Codex routesResponses API
Claude protocol contractsanthropic/claude-fable-5, anthropic/claude-fable-5.1, anthropic/claude-sonnet-5.5, anthropic/claude-sonnet-5, anthropic/claude-opus-5.5, anthropic/claude-opus-5, anthropic/claude-opus-4.8, anthropic/claude-opus-4.7, anthropic/claude-opus-4.6, anthropic/claude-opus-4.5, anthropic/claude-sonnet-4.6, anthropic/claude-sonnet-4.5, anthropic/claude-haiku-5.5, anthropic/claude-haiku-4.5, anthropic-enterprise/claude-fable-5, anthropic-enterprise/claude-fable-5.1, anthropic-enterprise/claude-sonnet-5.5, anthropic-enterprise/claude-sonnet-5, anthropic-enterprise/claude-opus-5.5, anthropic-enterprise/claude-opus-5, anthropic-enterprise/claude-opus-4.8, anthropic-enterprise/claude-opus-4.7, anthropic-enterprise/claude-opus-4.6, anthropic-enterprise/claude-opus-4.5, anthropic-enterprise/claude-sonnet-4.6, anthropic-enterprise/claude-sonnet-4.5, anthropic-enterprise/claude-haiku-5.5, anthropic-enterprise/claude-haiku-4.5, aws-bedrock/claude-fable-5, aws-bedrock/claude-fable-5.1, aws-bedrock/claude-sonnet-5.5, aws-bedrock/claude-sonnet-5, aws-bedrock/claude-opus-5.5, aws-bedrock/claude-opus-5, aws-bedrock/claude-opus-4.8, aws-bedrock/claude-opus-4.7, aws-bedrock/claude-opus-4.6, aws-bedrock/claude-opus-4.5, aws-bedrock/claude-sonnet-4.6, aws-bedrock/claude-sonnet-4.5, aws-bedrock/claude-haiku-5.5, aws-bedrock/claude-haiku-4.5/v1/messages, /responses, /chat/completionsKey-scoped; verify with authenticated /modelsAnthropic-style and AWS Bedrock-style Claude clients over one managed capability poolMessages
Embeddingsopenai/text-embedding-3-small, openai/text-embedding-3-large/embeddingsKey-scoped; verify with authenticated /modelsSearch, RAG, clustering, semantic matchingEmbeddings
GPT imageopenai/gpt-image-2, openai/gpt-image-2.5-flare, openai/gpt-image-2.5-sunburst/images/generationsKey-scoped; verify with authenticated /modelsProduct images and marketing creativeGPT Image 2
Gemini text and reasoninggoogle/gemini-3.8-flash, google/gemini-3.7-flash, google/gemini-3.1-pro-preview/responses, /chat/completions, Gemini :generateContentKey-scoped; verify with authenticated /modelsGemini text, multimodal, and reasoning-capable appsgenerateContent
GLM textz-ai/glm-5.2, z-ai/glm-5.3, z-ai/glm-5.3-flash, z-ai/glm-5.3-flashx/chat/completionsKey-scoped; verify with authenticated /modelsOpenAI-compatible GLM chat for entitled keysChat Completions
Grok Responsesx-ai/grok-4.5, x-ai/grok-4.6/responsesKey-scoped; verify with authenticated /modelsResponses API for coding, agentic tasks, and knowledge workResponses API
Kimi code chatmoonshotai/kimi-k2.7-code, moonshotai/kimi-k3/chat/completionsKey-scoped; verify with authenticated /modelsCoding, agent, and long-context chatChat Completions
DeepSeek textdeepseek/deepseek-v4-flash, deepseek/deepseek-v4-pro/chat/completionsKey-scoped; verify with authenticated /modelsLong-context chat, coding, and extractionChat Completions
Gemini imagegoogle/gemini-3-pro-image, google/gemini-3.1-flash-image, google/gemini-3.1-flash-lite-image, google/gemini-3.1-flash-image-preview, google/gemini-2.5-flash-image/images/generationsKey-scoped; verify with authenticated /modelsFast image generation and creative variantsImages API

gpt-5.6, gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna are public ids for the same canonical GPT 5.6 model contract. They do not select different upstream routes. Use the exact id returned for your key. GPT 5.6 supports function and custom tools on /responses, including streaming agent requests, and function tools on /chat/completions. Capability-aware routing selects only supply with matching tool evidence and fails closed when no verified candidate is available. For Chat Completions requests that use reasoning with function tools, set reasoning_effort to none; use /responses when reasoning and tool use must be combined. gpt-6-astra is its own canonical id and does not share that contract. It serves /responses and /chat/completions, but function tools are available on /responses only: OpenAI refuses a function tool on /chat/completions for this model at every reasoning effort it accepts, and the reasoning_effort: none workaround above does not apply because gpt-6-astra accepts only low, medium, high and xhigh. gpt-6-sol and gpt-6-luna, published 2026-09-14, are each their own canonical id as well. Both serve /responses and /chat/completions. Unlike gpt-6-astra, the reasoning_effort: none workaround DOES apply to them: they accept none, low, medium, high and xhigh, and a function tool on /chat/completions is served when reasoning_effort is none and refused at every other value. Use /responses when reasoning and tool use must be combined. gpt-6.1-sol, published 2026-09-27, is its own canonical id and behaves like gpt-6-astra on tools, not like gpt-6-sol: it serves /responses and /chat/completions, function tools are available on /responses only, and the reasoning_effort: none workaround does not apply because gpt-6.1-sol accepts only low, medium, high and xhigh. Image generation uses the dedicated synchronous Images API.

For any Claude-compatible model id currently returned by your authenticated /models response, send that exact anthropic/claude-*, anthropic-enterprise/claude-*, or aws-bedrock/claude-* id. Do not send non-default temperature, top_p, or top_k; use adaptive thinking and endpoint-specific effort hints instead. Model availability is evaluated per key and may be withdrawn until a verified route is available; never substitute a different Claude model id automatically. A supplier fallback, when one is used inside Quotaflow, does not change the returned canonical model or the opaque message-id profile selected by your key.

The currently reviewed Claude family includes Fable 5, Sonnet 5, Opus 4.8, Opus 4.7, Opus 4.6, Opus 4.5, Sonnet 4.6, Sonnet 4.5, and Haiku 4.5 variants. See Anthropic-compatible Messages for the copyable namespace mapping and request rules.

If you are buildingStart withWhy
A tool-capable OpenAI-compatible app or agentopenai/gpt-5.4 with /responsesStrong general default.
A low-latency app chat flowopenai/gpt-5.4-miniGood speed and capability balance.
A cost-sensitive lightweight flowopenai/gpt-5.4-nanoSmall-model option for simple chat, extraction, and routing.
A coding-agent integrationopenai/gpt-5.3-codex with /responsesIntended for Codex-style usage.
A vector search pipelineopenai/text-embedding-3-smallGood default embedding model.
A high-quality image generation appopenai/gpt-image-2Current V3 generation-only image path.
A Gemini reasoning appgoogle/gemini-3.1-pro or google/gemini-3.5-flashUse /chat/completions or /responses for OpenAI-style clients — both take the same Gemini model ids, with or without the google/ prefix — or generateContent for Gemini-native features such as countTokens and cached content. /responses is stateless here: send the whole turn in input, and store: true, previous_response_id, conversation, prompt and background return 400.
A GLM chat appz-ai/glm-5.2 with /chat/completionsChat-only OpenAI-compatible text path for keys with GLM access. Use non-streaming messages or non-streaming tool calls. stream: true currently fails closed with 503 and will resume only after verified availability evidence exists. Strict JSON mode returns 400 invalid_request_error.
A Kimi coding chat appmoonshotai/kimi-k2.7-code with /chat/completionsKimi text path for coding, agent, and OpenAI-compatible chat clients that need streaming, tools, or strict JSON mode.
A fast image-iteration workflowgoogle/gemini-3.1-flash-image-previewFast creative iteration through the OpenAI-compatible Images endpoint.
An async video workflowbytedance/seedance-2.0-fastAvailable Seedance video task generation.

Validation checklist

  1. Call /models with the same key and base URL your app will use.
  2. Confirm the desired model id is returned.
  3. Confirm the endpoint in the matrix matches your request shape.
  4. Send one small request before enabling production traffic.
  5. Treat model_not_found or unsupported-model errors as configuration errors.
  6. Retry only temporary errors such as 429, 503, and network timeouts.

GPT Image 2.5

gpt-image-2.5-flare and gpt-image-2.5-sunburst are distinct models, exposed as openai/gpt-image-2.5-flare and openai/gpt-image-2.5-sunburst. There is no gpt-image-2.5 alias. Both use the synchronous Images API; check your authenticated /models response for availability. See GPT Image 2.5 parameters, sizes, and prices.