Quotaflow
llms.txtOpenAPIDashboard
Model Library

Quotaflow Model Library — supported AI models by use case

Choose the right Quotaflow model for chat, agents, code, embeddings, image generation, and video generation. The catalog below is organized like a product model library: start with your use case, pick a curated model family, then open the endpoint-specific reference.

Quotaflow Model Library
Quotaflow Model Library

What is the Quotaflow Model Library?

The Model Library is the customer-facing list of model ids and endpoint families that Quotaflow documents for production integrations. It explains what each model is good for, which protocol surface it uses, and which reference page has copy-pasteable API examples.

Your authenticated GET /models response is still the source of truth for your key. A package, beta flag, or product-scoped key can make your returned list smaller than the public catalog.

Why use Quotaflow model routing?

Quotaflow lets teams keep official-style request shapes while centralizing API keys, team usage, package limits, and model access in one control plane. That means you can point OpenAI-compatible, Anthropic-compatible, Gemini-compatible, and creative-media clients at Quotaflow without teaching every app a new billing or credential system.

Use this page when you need to answer three practical questions before shipping:

Curated models

CategoryStart withEndpoint familyBest forReference
GPT 5.6 familygpt-5.6, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna/responses, /chat/completionsKey-scoped GPT 5.6 aliases over one canonical model contractOpenAI-compatible models
GPT agentsgpt-5.4/openai/v1/responsesProduction agents, tool use, app reasoningResponses API
Fast GPT chatgpt-5.4-mini/responses, /chat/completionsUI chat, extraction, summaries, routingOpenAI-compatible models
Lightweight GPT routinggpt-5.4-nano/responses, /chat/completionsCost-sensitive simple chat, extraction, and routingOpenAI-compatible models
Coding agentsgpt-5.3-codex/responsesCodex-style repository work and agent tasksConnect Coding Agents
Embeddingstext-embedding-3-small, text-embedding-3-large, text-embedding-ada-002/embeddingsSearch, RAG, clustering, semantic matchingEmbeddings
OpenAI-style imagesgpt-image-2/images/generationsProduct images and marketing creativeGPT Image 2
Fast image iterationgemini-3.1-flash-image-preview/images/generations, Gemini :generateContentDrafts and many generated variantsGemini 3.1 Flash Image
Seedance video generationseedance-2.0-fast/videosSeedance 2.0 text, image, portrait, and reference-media video tasksSeedance 2.0 Fast
Claude-compatible clientsclaude-fable-5, claude-sonnet-5, claude-opus-4-8, claude-opus-4-7, and the other exact ids returned by /models/v1/messagesClaude-style coding and reasoning through the key's assigned anthropic/* or aws-bedrock/* response identityMessages
GLM chatglm-5.2/openai/v1/chat/completionsOpenAI-compatible GLM text with non-streaming messages and non-streaming tool calls; stream: true currently fails closed with 503 pending verified availability evidence; strict JSON mode returns 400 invalid_request_errorChat Completions
Kimi coding chatkimi-k2.7-code/openai/v1/chat/completionsKimi text for coding, agent, long-context chat, streaming, tools, and strict JSON modeChat Completions
Gemini-compatible clientsgemini-3.5-flash, gemini-3.1-pro-preview, gemini-2.5-pro/v1beta/models/{model}:generateContent, /openai/v1/chat/completionsGemini text, multimodal, and reasoning-capable apps; Gemini-native features stay on generateContent/countTokensgenerateContent

Complete public model-id catalog

The table below is generated from the public model catalog. It lists currently documented public model ids; an authenticated /models response remains the final entitlement check for a specific key.

FamilyPublic model idsPrimary endpointsAvailability
OpenAI author aliasesopenai/gpt-4, openai/gpt-4-1106-preview, openai/gpt-4-turbo, openai/gpt-4.1, openai/gpt-4.1-mini, openai/gpt-4.1-nano, openai/gpt-4o, openai/gpt-4o-mini, openai/gpt-5, openai/gpt-5-mini, openai/gpt-5-nano, openai/gpt-5.1, openai/gpt-5.2, openai/gpt-5.4, openai/gpt-5.4-mini, openai/gpt-5.4-nano, openai/gpt-5.4-pro, openai/gpt-5.5, openai/gpt-5.5-pro, openai/gpt-5.6, openai/gpt-5.6-sol, openai/gpt-5.6-terra, openai/gpt-5.6-luna, openai/gpt-5-codex, openai/gpt-5.1-codex, openai/gpt-5.1-codex-mini, openai/gpt-5.1-codex-max, openai/gpt-5.2-codex, openai/gpt-5.3-codex/responses, /chat/completionsKey-scoped; verify with authenticated /models
Claude protocol contractsanthropic/claude-fable-5, anthropic/claude-sonnet-5, anthropic/claude-opus-4.8, anthropic/claude-opus-4.7, anthropic/claude-opus-4.6, anthropic/claude-opus-4.5, anthropic/claude-sonnet-4.6, anthropic/claude-sonnet-4.5, anthropic/claude-sonnet-4.5-20250929, anthropic/claude-opus-4.5-20251101, anthropic/claude-haiku-4.5, aws-bedrock/claude-fable-5, aws-bedrock/claude-sonnet-5, aws-bedrock/claude-opus-4.8, aws-bedrock/claude-opus-4.7, aws-bedrock/claude-opus-4.6, aws-bedrock/claude-opus-4.5, aws-bedrock/claude-sonnet-4.6, aws-bedrock/claude-sonnet-4.5, aws-bedrock/claude-sonnet-4.5-20250929, aws-bedrock/claude-opus-4.5-20251101, aws-bedrock/claude-haiku-4.5/v1/messages, /responses, /chat/completionsKey-scoped; verify with authenticated /models
Embeddingsopenai/text-embedding-3-small, openai/text-embedding-3-large, openai/text-embedding-ada-002/embeddingsKey-scoped; verify with authenticated /models
GPT imageopenai/gpt-image-2/images/generationsKey-scoped; verify with authenticated /models
Gemini text and reasoninggoogle/gemini-3.7-flash, google/gemini-3.5-flash, google/gemini-3.1-pro, google/gemini-3.1-pro-preview, google/gemini-3.1-flash-lite, google/gemini-3-flash-preview, google/gemini-2.5-pro, google/gemini-2.5-flash, google/gemini-2.5-flash-lite/chat/completions, Gemini :generateContentKey-scoped; verify with authenticated /models
GLM textz-ai/glm-5.2/chat/completionsKey-scoped; verify with authenticated /models
Grok Responsesx-ai/grok-4.5, x-ai/grok-4.6/responsesKey-scoped; verify with authenticated /models
Kimi code chatmoonshotai/kimi-k2.7-code, moonshotai/kimi-k3/chat/completionsKey-scoped; verify with authenticated /models
Gemini imagegoogle/gemini-3.1-flash-image-preview, google/gemini-2.5-flash-image/images/generations, Gemini :generateContentKey-scoped; verify with authenticated /models
Videobytedance/seedance-2.0, bytedance/seedance-2.0-fast, bytedance/seedance-2.0-mini, bytedance/seedance-2.5/videosKey-scoped; verify with authenticated /models

For endpoint-specific support and model restrictions, use the OpenAI-compatible model matrix. Do not send model ids absent from your key's /models response.

Key features

For the exact copyable list, open Supported Models. That page is generated from the same reviewed public catalog used by the runtime model-discovery contract; your key's authenticated /models response can be a smaller subset.

How to choose a model

  1. Pick the protocol your client already speaks: OpenAI-compatible, Anthropic-compatible, or Gemini-compatible.
  2. Call the matching /models endpoint with the exact key your app will use.
  3. Choose the first curated model in the table for your category.
  4. Send a small non-streaming test request.
  5. Add a streaming, multipart, or async polling test if your production flow uses that capability.

Pricing and access

Model availability depends on your Quotaflow package and enabled product surfaces. If a model appears in this catalog but not in your authenticated /models response, ask your Quotaflow admin to enable that model family for the key or create a key scoped to the right product.

For creative media workloads, budget for longer runtimes than text calls. Image and video jobs should use client timeouts, retry policies, and UI pending states that match media-generation latency.

Use cases

Use caseRecommended path
Build a coding-agent gatewaygpt-5.3-codex with /responses, then follow Codex setup.
Add generated product images to an ecommerce toolgpt-image-2 with /images/generations.
Create many visual concepts quicklygemini-3.1-flash-image-preview for fast variants, then optionally promote final candidates to a higher-quality image model.
Generate short social clipsseedance-2.0-fast with async /videos tasks and a polling UI.
Keep an existing Claude clientUse the Anthropic-compatible /v1/messages surface and a Claude-compatible model returned by your key.
Add GLM text chatUse glm-5.2 through /openai/v1/chat/completions after confirming it appears in /models; use non-streaming messages or non-streaming tools. Streaming currently fails closed with 503 and will resume only after verified availability evidence exists. Strict JSON mode returns 400 invalid_request_error.
Add Kimi coding chatUse kimi-k2.7-code through /openai/v1/chat/completions after confirming it appears in /models; streaming, tools, and strict JSON mode are supported on this chat path.

FAQ

How do I know which models my key can use?

Call GET https://api.quotaflow.ai/openai/v1/models for OpenAI-compatible keys, or the protocol-specific discovery endpoint for Anthropic-compatible and Gemini-compatible clients. Use the same key and base URL that your production app will use.

Are image and video models OpenAI-compatible?

Quotaflow currently exposes the OpenAI-compatible /images/generations endpoint. Image edits fail closed until separately audited and activated. Video generation uses an OpenAI-style task endpoint where the create call returns a task id and the client polls for completion.

Do all keys include every model?

No. Keys can be scoped by product, package, customer, or beta access. Treat this page as the public catalog and /models as the exact enabled list for a specific key.

What happens if I send an unsupported model id?

The request fails with a model or configuration error. Update the configured model id or enable the model for the key; do not retry unsupported-model errors as transient failures.

Next steps