Models
Use this page when you need to verify model ids, understand endpoint support, and choose the right OpenAI-compatible model before configuring a client. For a product-style visual catalog, start with the Quotaflow Model Library.
What /models means
GET /models is scoped to the authenticated key and base URL you call. Quotaflow lists only public model ids enabled for that key. A default customer key can include multiple model families; a product-scoped key may list only GPT, embeddings, images, video, or another enabled subset.
Unsupported model ids fail fast instead of being silently changed to another public model id.
Quotaflow model discovery uses OpenRouter-style author/model names. The author namespace describes the model identity, not the upstream vendor selected for a request: openai/* covers GPT, Codex, GPT Image, and embeddings; anthropic/* covers Claude; google/* covers Gemini, Gemini image, and Veo; z-ai/*, x-ai/*, moonshotai/*, and bytedance/* cover GLM, Grok, Kimi, and Seedance. Routing, pricing, supply snapshots, and receipts use one internal canonical model id, so changing the public author prefix never creates or pins a vendor route.
Legacy unprefixed ids remain accepted as inbound compatibility aliases. New integrations should send the prefixed id returned by /models. Unknown or mismatched author namespaces fail closed. A Claude key exposes exactly one assigned identity namespace—anthropic/claude-*, anthropic-enterprise/claude-*, or aws-bedrock/claude-*—across both /v1/models and /openai/v1/models; an unprefixed Claude alias follows that assignment. The aws-bedrock/claude-* namespace is a Quotaflow Claude response projection, not an OpenRouter author namespace and not a supplier selector. The anthropic-enterprise/claude-* prefix selects an isolated enterprise line. Enterprise ids appear only for a key assigned to the enterprise line. Google Veo names remain hidden until an exact executable route, price, and capability receipt are available.
Endpoint
GET https://api.quotaflow.ai/openai/v1/models
Curl
curl https://api.quotaflow.ai/openai/v1/models \
-H "Authorization: Bearer $QUOTAFLOW_API_KEY"
Curated model families
| Model family | Model ids | Endpoints | Availability | Common cases | Reference |
|---|---|---|---|---|---|
| OpenAI author aliases | openai/gpt-5, openai/gpt-5-mini, openai/gpt-5-nano, openai/gpt-5.1, openai/gpt-5.4, openai/gpt-5.4-mini, openai/gpt-5.4-nano, openai/gpt-5.4-pro, openai/gpt-5.5, openai/gpt-5.6, openai/gpt-5.6-sol, openai/gpt-5.6-terra, openai/gpt-5.6-luna, openai/gpt-6-astra, openai/gpt-6-sol, openai/gpt-6-luna, openai/gpt-6.1-sol | /responses, /chat/completions | Key-scoped; verify with authenticated /models | OpenRouter-style OpenAI model ids over the same canonical GPT and Codex routes | Responses API |
| Claude protocol contracts | anthropic/claude-fable-5, anthropic/claude-fable-5.1, anthropic/claude-sonnet-5.5, anthropic/claude-sonnet-5, anthropic/claude-opus-5.5, anthropic/claude-opus-5, anthropic/claude-opus-4.8, anthropic/claude-opus-4.7, anthropic/claude-opus-4.6, anthropic/claude-opus-4.5, anthropic/claude-sonnet-4.6, anthropic/claude-sonnet-4.5, anthropic/claude-haiku-5.5, anthropic/claude-haiku-4.5, anthropic-enterprise/claude-fable-5, anthropic-enterprise/claude-fable-5.1, anthropic-enterprise/claude-sonnet-5.5, anthropic-enterprise/claude-sonnet-5, anthropic-enterprise/claude-opus-5.5, anthropic-enterprise/claude-opus-5, anthropic-enterprise/claude-opus-4.8, anthropic-enterprise/claude-opus-4.7, anthropic-enterprise/claude-opus-4.6, anthropic-enterprise/claude-opus-4.5, anthropic-enterprise/claude-sonnet-4.6, anthropic-enterprise/claude-sonnet-4.5, anthropic-enterprise/claude-haiku-5.5, anthropic-enterprise/claude-haiku-4.5, aws-bedrock/claude-fable-5, aws-bedrock/claude-fable-5.1, aws-bedrock/claude-sonnet-5.5, aws-bedrock/claude-sonnet-5, aws-bedrock/claude-opus-5.5, aws-bedrock/claude-opus-5, aws-bedrock/claude-opus-4.8, aws-bedrock/claude-opus-4.7, aws-bedrock/claude-opus-4.6, aws-bedrock/claude-opus-4.5, aws-bedrock/claude-sonnet-4.6, aws-bedrock/claude-sonnet-4.5, aws-bedrock/claude-haiku-5.5, aws-bedrock/claude-haiku-4.5 | /v1/messages, /responses, /chat/completions | Key-scoped; verify with authenticated /models | Anthropic-style and AWS Bedrock-style Claude clients over one managed capability pool | Messages |
| Embeddings | openai/text-embedding-3-small, openai/text-embedding-3-large | /embeddings | Key-scoped; verify with authenticated /models | Search, RAG, clustering, semantic matching | Embeddings |
| GPT image | openai/gpt-image-2, openai/gpt-image-2.5-flare, openai/gpt-image-2.5-sunburst | /images/generations | Key-scoped; verify with authenticated /models | Product images and marketing creative | GPT Image 2 |
| Gemini text and reasoning | google/gemini-3.8-flash, google/gemini-3.7-flash, google/gemini-3.1-pro-preview | /responses, /chat/completions, Gemini :generateContent | Key-scoped; verify with authenticated /models | Gemini text, multimodal, and reasoning-capable apps | generateContent |
| GLM text | z-ai/glm-5.2, z-ai/glm-5.3, z-ai/glm-5.3-flash, z-ai/glm-5.3-flashx | /chat/completions | Key-scoped; verify with authenticated /models | OpenAI-compatible GLM chat for entitled keys | Chat Completions |
| Grok Responses | x-ai/grok-4.5, x-ai/grok-4.6 | /responses | Key-scoped; verify with authenticated /models | Responses API for coding, agentic tasks, and knowledge work | Responses API |
| Kimi code chat | moonshotai/kimi-k2.7-code, moonshotai/kimi-k3 | /chat/completions | Key-scoped; verify with authenticated /models | Coding, agent, and long-context chat | Chat Completions |
| DeepSeek text | deepseek/deepseek-v4-flash, deepseek/deepseek-v4-pro | /chat/completions | Key-scoped; verify with authenticated /models | Long-context chat, coding, and extraction | Chat Completions |
| Gemini image | google/gemini-3-pro-image, google/gemini-3.1-flash-image, google/gemini-3.1-flash-lite-image, google/gemini-3.1-flash-image-preview, google/gemini-2.5-flash-image | /images/generations | Key-scoped; verify with authenticated /models | Fast image generation and creative variants | Images API |
gpt-5.6, gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna are public ids for the same canonical GPT 5.6 model contract. They do not select different upstream routes. Use the exact id returned for your key. GPT 5.6 supports function and custom tools on /responses, including streaming agent requests, and function tools on /chat/completions. Capability-aware routing selects only supply with matching tool evidence and fails closed when no verified candidate is available. For Chat Completions requests that use reasoning with function tools, set reasoning_effort to none; use /responses when reasoning and tool use must be combined. gpt-6-astra is its own canonical id and does not share that contract. It serves /responses and /chat/completions, but function tools are available on /responses only: OpenAI refuses a function tool on /chat/completions for this model at every reasoning effort it accepts, and the reasoning_effort: none workaround above does not apply because gpt-6-astra accepts only low, medium, high and xhigh. gpt-6-sol and gpt-6-luna, published 2026-09-14, are each their own canonical id as well. Both serve /responses and /chat/completions. Unlike gpt-6-astra, the reasoning_effort: none workaround DOES apply to them: they accept none, low, medium, high and xhigh, and a function tool on /chat/completions is served when reasoning_effort is none and refused at every other value. Use /responses when reasoning and tool use must be combined. gpt-6.1-sol, published 2026-09-27, is its own canonical id and behaves like gpt-6-astra on tools, not like gpt-6-sol: it serves /responses and /chat/completions, function tools are available on /responses only, and the reasoning_effort: none workaround does not apply because gpt-6.1-sol accepts only low, medium, high and xhigh. Image generation uses the dedicated synchronous Images API.
For any Claude-compatible model id currently returned by your authenticated /models response, send that exact anthropic/claude-*, anthropic-enterprise/claude-*, or aws-bedrock/claude-* id. Do not send non-default temperature, top_p, or top_k; use adaptive thinking and endpoint-specific effort hints instead. Model availability is evaluated per key and may be withdrawn until a verified route is available; never substitute a different Claude model id automatically. A supplier fallback, when one is used inside Quotaflow, does not change the returned canonical model or the opaque message-id profile selected by your key.
The currently reviewed Claude family includes Fable 5, Sonnet 5, Opus 4.8, Opus 4.7, Opus 4.6, Opus 4.5, Sonnet 4.6, Sonnet 4.5, and Haiku 4.5 variants. See Anthropic-compatible Messages for the copyable namespace mapping and request rules.
Recommended defaults
| If you are building | Start with | Why |
|---|---|---|
| A tool-capable OpenAI-compatible app or agent | openai/gpt-5.4 with /responses | Strong general default. |
| A low-latency app chat flow | openai/gpt-5.4-mini | Good speed and capability balance. |
| A cost-sensitive lightweight flow | openai/gpt-5.4-nano | Small-model option for simple chat, extraction, and routing. |
| A coding-agent integration | openai/gpt-5.3-codex with /responses | Intended for Codex-style usage. |
| A vector search pipeline | openai/text-embedding-3-small | Good default embedding model. |
| A high-quality image generation app | openai/gpt-image-2 | Current V3 generation-only image path. |
| A Gemini reasoning app | google/gemini-3.1-pro or google/gemini-3.5-flash | Use /chat/completions or /responses for OpenAI-style clients — both take the same Gemini model ids, with or without the google/ prefix — or generateContent for Gemini-native features such as countTokens and cached content. /responses is stateless here: send the whole turn in input, and store: true, previous_response_id, conversation, prompt and background return 400. |
| A GLM chat app | z-ai/glm-5.2 with /chat/completions | Chat-only OpenAI-compatible text path for keys with GLM access. Use non-streaming messages or non-streaming tool calls. stream: true currently fails closed with 503 and will resume only after verified availability evidence exists. Strict JSON mode returns 400 invalid_request_error. |
| A Kimi coding chat app | moonshotai/kimi-k2.7-code with /chat/completions | Kimi text path for coding, agent, and OpenAI-compatible chat clients that need streaming, tools, or strict JSON mode. |
| A fast image-iteration workflow | google/gemini-3.1-flash-image-preview | Fast creative iteration through the OpenAI-compatible Images endpoint. |
| An async video workflow | bytedance/seedance-2.0-fast | Available Seedance video task generation. |
Validation checklist
- Call
/modelswith the same key and base URL your app will use. - Confirm the desired model id is returned.
- Confirm the endpoint in the matrix matches your request shape.
- Send one small request before enabling production traffic.
- Treat
model_not_foundor unsupported-model errors as configuration errors. - Retry only temporary errors such as
429,503, and network timeouts.
GPT Image 2.5
gpt-image-2.5-flare and gpt-image-2.5-sunburst are distinct models, exposed as openai/gpt-image-2.5-flare and openai/gpt-image-2.5-sunburst. There is no gpt-image-2.5 alias. Both use the synchronous Images API; check your authenticated /models response for availability. See GPT Image 2.5 parameters, sizes, and prices.