How billing works
Quotaflow charges per token against your organization balance, which is the only spendable money account for customer-organization keys. Project budgets are reporting targets, not separate wallets. Per-model unit prices are returned by the authenticated /models response and listed on each model's reference page — treat those as the source of truth for exact rates.
What you are charged for
- Tokens actually processed, on requests that complete. Input and output tokens are metered per request at the model's unit price.
- Interrupted streams. If a streaming response is cut on the serving side before the model reported its final usage, the request is not charged, including any output tokens that were already delivered; it appears in your console request log with an "Interrupted" status and a "Not billed" amount. If your own client disconnects after the model has started processing the request, it is settled the way the official API settles a disconnect: the prompt that was processed (input, cache read, cache write) at your contracted price, plus the output produced before the disconnect — the model's reported count when it arrived, otherwise the delivered output counted at four bytes per token. A stream your client closes before the model processed anything is not charged. These requests appear in the request log as "Client disconnected (499)" or, when partial output had been delivered, "Interrupted".
- Long-context requests, at a surcharge. When a single request's input plus cache-read tokens exceed the model's long-context threshold, that request's input is billed at 2× and its output at 1.5× the normal rate. The threshold varies by model — for example, GPT-5.4 / 5.5 / 5.6 models use 272,000 tokens and several other models (such as some Gemini Pro models) use 200,000. Once the threshold is crossed, the multiplier applies to the whole request.
Caching
A prompt cache hit (cache-read tokens) is billed at a lower rate than normal input tokens, and cache writes are priced separately. Exact cache rates differ by model and are part of each model's pricing — see /models and the model reference pages.
Post-settlement corrections
Occasionally a request is served and its settlement does not complete at the time. If we charge for it later, the charge is attributed to the period and the model of the original request — it raises that period's usage totals instead of appearing as a separate entry dated the day the correction was applied. The "Charged to wallet" figure on your wallet page for a period therefore includes any corrections that belong to requests served in that period.
Downloading a statement
The console's Usage page offers a Monthly statement: pick a month, download a CSV.
A statement is one whole calendar month, and the month is cut at UTC midnight — not at your local midnight, and not at ours. Two people in different time zones downloading the same month therefore receive the same rows, and the same month downloaded again next year still means the same period. The period is stated beside the picker as the half-open interval it is: 2026-01-01T00:00:00Z — 2026-02-01T00:00:00Z, where the end bound is exclusive.
The file has one row per model plus a final TOTAL row, and every row repeats the period, so a statement stays readable after it is renamed or pasted beside another one.
| Column | What it is |
|---|---|
period_start_utc, period_end_utc | The statement period. The end bound is exclusive. |
model | The model the row accounts for. TOTAL on the last row. |
requests | Requests for that model whose charge has settled into the period. A request served but not yet settled is absent until its charge is applied. |
input_tokens, output_tokens, cache_creation_tokens, cache_read_tokens, total_tokens | Tokens metered in the period. The TOTAL row leaves the two cache columns empty: the account totals carry one combined cache figure rather than the split. |
cost_usd | The list price of that traffic. Amounts are written to ten decimal places, the precision charges are stored at, so a fraction-of-a-cent line is never rounded away to zero. |
actual_cost_usd | The charge for that traffic after any discount that applies to your account. This is the usage book's figure; the wallet ledger is the record of money that moved, and the two differ only in the rare case where a charge exceeded the admission hold taken for the request. |
A statement covers your whole account — every API key and every project in the scope the page is showing. It is not narrowed by the API-key, model, or window filters set on the Usage page above it; those filter the request log, not the statement.
A statement is a read of a period, not a snapshot frozen when you downloaded it. Because a post-settlement correction (see above) is attributed to the period of the original request, re-downloading a past month can produce a slightly larger total than an earlier download of the same month did. The period has not moved; a charge that belongs inside it landed late.
<Note>
The TOTAL row is the account total for the period as our billing read reports it, not a sum of the rows above it. The two are separate reads taken moments apart, so for a period still accruing they will not always agree to the cent — a request settling between them lands in one and not the other. Both figures go into the file unchanged. For a closed period they should agree; if they do not, send us the file.
</Note>
Reconciling a statement against your request log
The statement and the console's request log take the same period bounds, so a month's statement and that month's log can be made to describe the same requests. Two things have to match, not one:
- The window. Set the Usage window to the two instants the statement names — the control accepts an instant, and a value carrying an offset such as
2026-01-01T00:00:00Zis used exactly as written. - The filters. Clear the API-key and model filters. The statement always covers the whole account, so a log narrowed to one key is a smaller population than the file no matter how well the periods line up.
<Warning>
This does not extend to GET /v1/usage, the usage summary an API key can read. That endpoint accepts start_date and end_date as whole dates only, resolved in our server's time zone, and an RFC3339 value there is ignored in favour of its default 30-day window rather than rejected. Do not reconcile a UTC statement against it.
</Warning>
Exporting every request (CSV)
The statement above has one row per model. When you need to reconcile request by request, the Usage page also offers Export all requests (CSV): pick a month and request the export. It is built in the background; the Exports list beside the button shows its progress and offers Download once it is ready.
- Period. One whole UTC calendar month, cut at UTC midnight exactly like a statement. For the current month the export stops 15 minutes before the moment you requested it, and requests whose charge is still pending settlement are not included.
- Scope. The organization or project the Usage page is showing, every API key in it. Like a statement, it is not narrowed by the API-key, model, or window filters.
- File. A ZIP of one or more CSV parts, each holding at most 1,000,000 rows, one row per request.
- Prices. Every cost column is the official list price of the request. Your invoice reflects your contracted pricing, so the export's totals are not what you were charged when your account has a discount.
- Availability. A finished export can be downloaded for 7 days; after that it shows as Expired and you request a new one. A download link is short-lived, so start each download from the console.
- Limits. At most one export can be in progress per scope at a time, and at most 20 exports can be requested per 24 hours. When a limit is reached the console says so and, for the daily limit, when to try again.
- How long it takes. When you request an export the console shows an estimate of how long it will take, based on how many requests the month holds (for very large months it says "at least"); when no estimate can be made in time it says so rather than guessing. You do not need to stay on the page: the Exports list updates on its own, and when email notifications are set up for the service you also receive an email at your account address once the export is ready (or could not be built) with a link back to the Usage page, where you download it.
| Column | What it is |
|---|---|
time_utc | When the request was made, in UTC. |
request_id | The request's identifier, as shown in the request log. |
api_key_name | The name of the API key that made the request. |
model | The model the request was billed under. |
input_tokens, output_tokens, cache_read_tokens, cache_write_tokens | Tokens metered for the request. |
input_cost_usd, output_cost_usd, cache_read_cost_usd, cache_write_cost_usd, other_cost_usd | The list-price cost of each component of the request. |
official_cost_usd | The request's total at official list prices. |
Where to find exact prices
GET /openai/v1/balancereturns your current organization balance and billing context.GET /modelsreturns the models your key can use.- Per-model unit prices (input, output, and cache) are documented on each model's reference page.
<Note>
This page describes billing *behavior*. It intentionally does not hard-code per-model unit prices, which live in /models and the model reference pages so they stay authoritative.
</Note>