Anthropic-compatible Messages
When to use this page
Use this page when your client expects the Anthropic Messages API shape.
Base URL
https://api.quotaflow.ai
Endpoint
POST https://api.quotaflow.ai/v1/messages
Supported models
Use the exact prefixed id returned by your authenticated /models response:
anthropic/claude-*selects the Anthropic identity projection, includingmsg_01...message ids andtoolu_01...client-tool ids.aws-bedrock/claude-*selects the AWS Bedrock identity projection, includingmsg_bdrk_...message ids andtoolu_bdrk_...client-tool ids.
Each Claude key is assigned one of these identity namespaces. Authenticated model discovery, including both /v1/models and /openai/v1/models, returns only the namespace assigned to that key. A request with the other explicit prefix fails as an unknown model (404); use the exact id returned by model discovery.
All three Claude lines use the same key-scoped, capability-verified scheduling rules. The anthropic/ and aws-bedrock/ lines share one managed supply pool, where the assigned prefix controls the customer-visible response contract and not the upstream vendor selected for a request. The anthropic-enterprise/ line is served only by the supply assigned to it, and that supply serves no other line; the one exception is Quotaflow's own first-party supply, which can serve the enterprise line while still serving the others. The enterprise line has platform fallback only when that first-party supply is assigned to it. The response model is the bare canonical Claude id without the routing prefix. For every completed message, Quotaflow generates the public message id before the first response byte from that selected customer projection; it is opaque and does not reveal or preserve an upstream message id. A managed supplier may make an internal fallback decision, but it cannot change the requested public model, message-id profile, or add fallback/provider narration to the response. Tool definitions and tool_use_id result references are projected together so their relationship remains intact. Documented cache and thinking-token details are retained when supplied with valid public values; Anthropic-only service_tier and inference_geo fields are omitted from the AWS Bedrock projection. Provider headers, hosts, account metadata, route details, and upstream request ids are not exposed on any projection. The request and event schema at /v1/messages remains Anthropic-compatible for all three lines.
Enterprise line
The enterprise customer prefix selects a separate supply line while keeping the documented response identity:
anthropic-enterprise/claude-*uses the Anthropic identity projection (msg_01...).
The anthropic-enterprise/ line is independent. The response contract, message-id shape, and provider-metadata sanitization are identical to the matching base namespace, and the request must still match your key's assigned identity namespace: an anthropic-enterprise-bound key may send anthropic-enterprise/claude-*, while an anthropic-bound key stays on anthropic/claude-*.
Currently reviewed base ids include claude-fable-5, claude-sonnet-5, claude-opus-4-8, claude-opus-4-7, claude-opus-4-6, claude-opus-4-5, claude-sonnet-4-6, claude-sonnet-4-5, and claude-haiku-4-5. Availability is evaluated per key and request shape, so /models is authoritative.
Claude namespace model list
The table shows all three public response identities for each reviewed base model. A single key returns only the prefixed column assigned to that key.
| Base model | Anthropic identity | Anthropic Enterprise identity | AWS Bedrock identity |
|---|---|---|---|
| Fable 5 | anthropic/claude-fable-5 | anthropic-enterprise/claude-fable-5 | aws-bedrock/claude-fable-5 |
| Sonnet 5 | anthropic/claude-sonnet-5 | anthropic-enterprise/claude-sonnet-5 | aws-bedrock/claude-sonnet-5 |
| Opus 4.8 | anthropic/claude-opus-4.8 | anthropic-enterprise/claude-opus-4.8 | aws-bedrock/claude-opus-4.8 |
| Opus 4.7 | anthropic/claude-opus-4.7 | anthropic-enterprise/claude-opus-4.7 | aws-bedrock/claude-opus-4.7 |
| Opus 4.6 | anthropic/claude-opus-4.6 | anthropic-enterprise/claude-opus-4.6 | aws-bedrock/claude-opus-4.6 |
| Opus 4.5 | anthropic/claude-opus-4.5 | anthropic-enterprise/claude-opus-4.5 | aws-bedrock/claude-opus-4.5 |
| Sonnet 4.6 | anthropic/claude-sonnet-4.6 | anthropic-enterprise/claude-sonnet-4.6 | aws-bedrock/claude-sonnet-4.6 |
| Sonnet 4.5 | anthropic/claude-sonnet-4.5 | anthropic-enterprise/claude-sonnet-4.5 | aws-bedrock/claude-sonnet-4.5 |
| Sonnet 4.5 dated | anthropic/claude-sonnet-4.5-20250929 | anthropic-enterprise/claude-sonnet-4.5-20250929 | aws-bedrock/claude-sonnet-4.5-20250929 |
| Opus 4.5 dated | anthropic/claude-opus-4.5-20251101 | anthropic-enterprise/claude-opus-4.5-20251101 | aws-bedrock/claude-opus-4.5-20251101 |
| Haiku 4.5 | anthropic/claude-haiku-4.5 | anthropic-enterprise/claude-haiku-4.5 | aws-bedrock/claude-haiku-4.5 |
Legacy unprefixed Claude ids remain accepted during migration and follow the key's assigned identity namespace, but new integrations should use the prefixed ids returned by /models.
Extended thinking
Use the same Anthropic-compatible request fields your client already sends.
| Model | Recommended thinking shape |
|---|---|
claude-sonnet-5 | thinking: { "type": "adaptive" }; omit manual budget_tokens |
claude-opus-4-8 | thinking: { "type": "adaptive" } |
claude-opus-4-7 | thinking: { "type": "adaptive" } |
claude-opus-4-6 | thinking: { "type": "adaptive" } or legacy manual thinking |
claude-sonnet-4-6 | thinking: { "type": "adaptive" } or legacy manual thinking |
claude-sonnet-4-5 | Standard Anthropic-compatible thinking fields when enabled for your key |
claude-opus-4-5-20251101 | Standard Anthropic-compatible thinking fields when enabled for your key |
For claude-sonnet-5, claude-opus-4-8, and claude-opus-4-7, adaptive thinking is the supported mode. Send thinking: { "type": "adaptive" } or omit the field to use the model default. Do not send legacy thinking: { "type": "enabled", "budget_tokens": ... } or thinking: { "type": "disabled" } for adaptive-only Claude models; Quotaflow rejects those request shapes with a customer-terminal 400. Streaming responses preserve the Anthropic-compatible SSE event shape.
Route-aware request transport
Quotaflow owns the provider transport selected by your key and model namespace. You do not need to add anthropic-version to a Quotaflow request. Anthropic SDKs may still send anthropic-version: 2023-06-01; it is accepted as compatibility metadata, while Quotaflow constructs the actual provider request itself.
For anthropic/*, send beta opt-ins in the native anthropic-beta header. For aws-bedrock/*, Quotaflow also accepts the Bedrock InvokeModel body forms "anthropic_version": "bedrock-2023-05-31" and "anthropic_beta": ["..."]. Those body fields are optional at the Quotaflow gateway. If present, the version must use the Bedrock value; body and header beta tokens are merged and deduplicated before dispatch. Bedrock body transport fields are not reinterpreted on an anthropic/* route. The selected customer namespace controls this adaptation even when the managed supplier pool contains different Claude resource origins.
The retired context-1m-* token is accepted for client compatibility but is not forwarded upstream and does not select or purchase a 1M context service. New integrations should remove that retired token. The effective context limit is determined by the requested model, request size, and availability for your key; Quotaflow never silently substitutes a smaller context tier because of this header.
Sampling controls on adaptive-only Claude models
For claude-sonnet-5, claude-opus-4-8, and claude-opus-4-7, do not send non-default temperature, top_p, or top_k. Those models use adaptive behavior instead of explicit sampling controls. Quotaflow strips those fields where possible before forwarding, but clients should omit them to match the model contract. Use prompting plus thinking: { "type": "adaptive" } for depth control.
Output configuration
Top-level output_config is an accepted request field. Quotaflow recognizes output_config.format.type=json_schema as a structured-output request and routes it only to supply with positive capability evidence. Other output_config members, including effort, remain provider-defined passthrough fields. An admitted field is forwarded; support is still model- and customer-line-specific, so an unsupported combination returns a customer-terminal 400 in the selected anthropic/* or aws-bedrock/* error vocabulary. In particular, AWS documents structured output for supported Claude models through Bedrock Runtime InvokeModel/Converse, not the Bedrock Anthropic Messages endpoint.
Errors and request IDs
Non-streaming errors use the Anthropic-compatible envelope. The request-id response header and body request_id contain the same neutral Quotaflow gateway identifier on both projections. Quotaflow does not synthesize an AWS infrastructure request id:
{
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "Invalid request"
},
"request_id": "req_..."
}
Quotaflow returns the standard error types for their matching HTTP status: invalid_request_error (400 and other unmapped 4xx), authentication_error (401), billing_error (402), permission_error (403), not_found_error (404), conflict_error (409), request_too_large (413), rate_limit_error (429), api_error (500-class service failures), timeout_error (504), and overloaded_error (529). Internal route, account, provider, credential, and upstream transport classifications are never returned as error types.
After a streaming response has started, an error is emitted as an Anthropic-compatible SSE event. Its data contains only the documented type and error fields; use the neutral request-id header for gateway correlation.
event: error
data: {"type":"error","error":{"type":"api_error","message":"The service is temporarily unavailable. Please try again later."}}
Curl
curl https://api.quotaflow.ai/v1/messages \
-H "x-api-key: $QUOTAFLOW_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"max_tokens": 128,
"messages": [
{ "role": "user", "content": "Return only: connected" }
]
}'
Adaptive thinking curl
curl https://api.quotaflow.ai/v1/messages \
-H "x-api-key: $QUOTAFLOW_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-opus-4.8",
"max_tokens": 1024,
"thinking": { "type": "adaptive" },
"messages": [
{ "role": "user", "content": "Solve this carefully, then give the final answer." }
]
}'
SDK environment pattern
export ANTHROPIC_AUTH_TOKEN="$QUOTAFLOW_API_KEY"
export ANTHROPIC_BASE_URL="https://api.quotaflow.ai"
Quotaflow preserves the Anthropic-compatible request and response shape for downstream clients.