API reference
This is the request and response contract of MUSCLE Flex by Adhibita during the beta. It describes how the current Flex build is designed to behave. The compatibility matrix records what has been verified on the deployed beta: a row that reads Pending has not been verified yet.
Base URL and endpoints
Your invitation gives you the Flex base URL: the scheme and host only, written here as <FLEX_BASE_URL>. Every path below is relative to it. The examples read the base URL and your key from the shell:
bash
export FLEX_BASE_URL="<FLEX_BASE_URL>"
export MUSCLE_FLEX_API_KEY="<FLEX_API_KEY>"| Method | Path | Request shape | Same handler at | Charged |
|---|---|---|---|---|
POST | /v1/chat/completions | OpenAI-compatible Chat | /chat/completions | Yes |
POST | /v1/responses | OpenAI-compatible Responses | — | Yes |
POST | /v1/messages | Anthropic-compatible Messages | /messages | Yes |
POST | /v1/messages/count_tokens | Anthropic-compatible token count | /messages/count_tokens | No |
GET | /v1/models | Model list | /models | No |
GET | /v1/models/{model} | One model | /models/{model} | No |
GET | /v1/account | Account summary | — | No |
The paths without /v1 serve clients that strip /v1 from their base URL. Responses has no path without /v1.
Requests and responses are JSON, apart from streams. A request body that is not valid JSON returns 400 invalid_request_error, and one above the size limit returns 413 invalid_request_error.
POST /v1/messages/count_tokens takes a Messages request body and returns {"input_tokens": <n>}. The count is an estimate made by Flex; the usage on a reply is what counts for billing.
Authentication
Every endpoint, the model list included, needs a Flex API key. Create keys on the API keys page of the portal.
| Endpoint | Send the key as |
|---|---|
| Chat Completions, Responses, models, account | Authorization: Bearer <FLEX_API_KEY> |
| Messages and token count | Authorization: Bearer <FLEX_API_KEY> or x-api-key: <FLEX_API_KEY> |
- A missing, mistyped or revoked key returns
401 authentication_error. So does anx-api-keyheader alone on an OpenAI-shaped endpoint. - A suspended or inactive account returns
403 permission_denied. - Keep keys in headers. A request body with a field named like a credential or session token (for example
api_key,authorization,cookieorsession_id), at the top level or insidemetadata, returns400 invalid_request_error.
Check a key without spending credits:
bash
curl -i "$FLEX_BASE_URL/v1/account" \
-H "Authorization: Bearer $MUSCLE_FLEX_API_KEY"The model id
The model field is required. Send muscle/auto, the id GET /v1/models lists; auto is accepted as a short form of it. Flex chooses the route for each request, so the value does not select or pin an underlying model. Replies echo the model value you sent.
To choose the cost/quality profile for a request, send muscle/budget, muscle/balanced or muscle/max instead of muscle/auto. They route like muscle/auto and echo the id you sent. See Routing profile.
For client compatibility, Flex also accepts the model names that the supported coding agents and the official OpenAI and Anthropic SDKs send on their own, for example the ids Claude Code uses for its background requests and subagents. Flex routes them exactly like muscle/auto and echoes them; they never pin a model. Configure clients with muscle/auto.
Any other value, such as a misspelled id or another router's vendor/model id, returns 404 model_not_found before any credits are held, and nothing is charged. A missing or empty model returns 400 invalid_request_error. See Errors.
GET /v1/models lists muscle/auto with a context_length: the largest context window, in tokens, that any Flex route offers. A request fits a route when its input plus its output allowance fits that route's window. The output allowance is your max_tokens (or max_completion_tokens, or max_output_tokens). A request that fits no route returns 503 feature_unavailable; shorten the input or lower the output allowance.
GET /v1/models/muscle/auto returns the same entry on its own, for clients that look a model up by id. Any other id returns 404 model_not_found.
Without an output allowance, Flex uses 32,768 tokens, or the route's own output limit if that is lower, and less when a long conversation leaves less room than that in the route's context window. That default is a real cap: it counts toward the fit, it is sent to the route, and a reply that reaches it is cut off (see Streaming). An allowance you set is sent to the route up to the route's own output limit, with one exception: the output budget a model uses for reasoning is shared with the visible reply, so a small allowance can be used up by reasoning before any answer text appears. For that reason, when Flex serves a request on a long-context route (a very long conversation) or a graphics route (requests such as SVG or 3D scenes) and your allowance is below 4,096 tokens, it raises the allowance Flex sends to the route to 4,096. A larger allowance is sent as set. The raise does not apply when you set no allowance (the default above applies). It changes the budget the route can spend, not what you are billed: you are charged for tokens actually generated. The credit hold is sized from the allowance Flex sends, so 4,096 where it was raised (see Credits and metering). A smaller allowance elsewhere needs less available credit: on a low balance, a request can be refused with 402 budget_exceeded when a smaller one would be admitted. Set an output allowance that fits the reply you expect, and leave room for reasoning.
Requests
Each endpoint takes its own familiar request shape. Familiar does not mean identical: a field not listed here is ignored, unless this page or the compatibility matrix says it is rejected.
Chat Completions
bash
curl -i "$FLEX_BASE_URL/v1/chat/completions" \
-H "Authorization: Bearer $MUSCLE_FLEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "muscle/auto",
"messages": [{"role": "user", "content": "Name three prime numbers."}],
"max_tokens": 64
}'| Field | Behavior |
|---|---|
messages | system, developer (read as system), user, assistant and tool roles. |
max_tokens, max_completion_tokens | The output allowance. |
temperature | Passed to the route. |
stop | See Stop sequences. |
stream, stream_options.include_usage | See Streaming. |
tools, tool_choice | Function tools only. Tool choice auto, none, required or a named function. |
parallel_tool_calls | Accepted and treated as true. |
reasoning_effort | low, medium or high. See Reasoning effort. |
top_p, user | Accepted and ignored. |
n above 1, seed, logprobs, frequency_penalty, presence_penalty | 400 unsupported_field for any value other than null, even 0 or false. n: 1 is accepted. |
response_format | See Structured outputs. |
Responses
bash
curl -i "$FLEX_BASE_URL/v1/responses" \
-H "Authorization: Bearer $MUSCLE_FLEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "muscle/auto",
"input": "Write a one-line release note.",
"max_output_tokens": 80
}'| Field | Behavior |
|---|---|
input, instructions | A string or input items. Send the whole conversation every turn, including function_call_output items. |
max_output_tokens | The output allowance. |
temperature | Passed to the route. |
stop | A Flex addition to this shape. See Stop sequences. |
stream | See Streaming. |
tools, tool_choice | Function tools only. Hosted tools (web search, file search, code interpreter and others) return 400 unsupported_field. |
reasoning_effort | Top-level low, medium or high. reasoning.effort is not read. |
store: true, background: true, previous_response_id | 400 unsupported_field. Flex keeps no response state; store: false is accepted. |
text.format | See Structured outputs. |
Messages
bash
curl -i "$FLEX_BASE_URL/v1/messages" \
-H "x-api-key: $MUSCLE_FLEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "muscle/auto",
"max_tokens": 64,
"messages": [{"role": "user", "content": "Name three prime numbers."}]
}'| Field | Behavior |
|---|---|
system, messages | Text and tool_use / tool_result blocks. Image blocks return 400; see Images and audio. |
max_tokens | The output allowance. |
temperature | Passed to the route. |
stop_sequences | See Stop sequences. |
stream | See Streaming. |
tools, tool_choice | Client tools with an input_schema. Server-side tools are not run: Flex treats them as client tools, so remove them from requests. |
thinking | Not read. |
output_format, output_config.format | See Structured outputs. |
Images and audio
Image and audio input are not supported yet. Send text only.
An image or audio part anywhere in a request fails it before it is routed, on all three shapes: in any message role (system, developer, user, assistant or tool), in a Messages system array, inside a Messages tool_result, and in a Responses function_call_output. Flex never drops the media to run the rest as text. The response is 400 with the message "Image and audio input are not supported yet.": on Chat Completions and Responses, code is unsupported_content and type is invalid_request_error; on Messages, error.type is invalid_request_error. Nothing is charged. Retrying it unchanged fails again, and so does every later request that still carries the media, so remove it from the conversation. The token count (POST /v1/messages/count_tokens) refuses media the same way, with the same 400.
Reasoning effort
On Chat Completions and Responses, reasoning_effort tells Flex how to weigh cost against quality when it routes the request: low favours cost, high favours quality and medium sits between. Any other value returns 400 invalid_request_error.
Routing profile
A profile tells Flex how to weigh cost against quality: budget favours cost, max favours quality, balanced sits between. With none set, your workspace default applies, which is balanced unless your workspace was configured otherwise. Choose it on any of the three request shapes, in any of these ways:
| How | Example |
|---|---|
The x-flex-profile request header: budget, balanced or max | x-flex-profile: max |
| A profile model id | "model": "muscle/budget", muscle/balanced, muscle/max |
reasoning_effort (Chat Completions and Responses) | low is budget, medium is balanced, high is max |
When more than one is present, the most specific wins: the x-flex-profile header first, then a profile model id, then reasoning_effort. muscle/auto sets no profile of its own, so a header or reasoning_effort still applies to it. A header value is case-insensitive; any value other than budget, balanced or max, an empty value included, returns 400 invalid_request_error before anything is charged. The profile is part of the response-cache key, so a budget answer is never replayed to a max request. Replies echo the model value you sent, not the profile.
Clients that only let you set a model name, such as Claude Code, use the model ids. See Claude Code.
Streaming
Set "stream": true. Replies arrive as server-sent events in the shape of the endpoint you called.
| Endpoint | Events | End of stream |
|---|---|---|
| Chat Completions | data: lines, each a chat.completion.chunk. The first delta carries role: "assistant"; the last carries finish_reason. With stream_options: {"include_usage": true}, one more chunk with empty choices carries usage. | data: [DONE] |
| Responses | data: lines with typed events: response.created, response.in_progress, response.output_item.added, response.content_part.added, response.output_text.delta, response.function_call_arguments.delta, their .done events, then response.completed, or response.incomplete when the reply ran out of room (the output allowance, 32,768 tokens by default) or content policy stopped the reply. The final event carries usage. | data: [DONE] |
| Messages | event: and data: pairs, from message_start through content_block_start, content_block_delta and content_block_stop to message_delta (with stop_reason, stop_sequence and usage) and message_stop. | message_stop |
Keep-alives. Some models think for a minute or more before the first word of a reply, and that reasoning is not streamed. Some requests are also worked on in several steps before the reply starts. Every stream opens with its first event as soon as the request is admitted, before the reply exists: the
role: "assistant"chunk on Chat Completions,response.createdandresponse.in_progresson Responses,message_starton Messages. While a reply is quiet, Chat Completions and Responses streams send SSE comment lines (: keepalive) and Messages streams sendpingevents, about every 15 seconds, until the reply starts or the request's time limit ends it. Standard SSE clients ignore comments. Give your client's read timeout room for at least 15 seconds between events.Unknown events. Ignore event types your client does not recognise. A stream may carry more event types than those listed here.
Errors after the stream starts. The HTTP status is already
200and cannot change. A failure after that point ends the stream with an error frame, and nothing follows the frame, not even the end marker:- Chat Completions: a
data:line with{"error": {"message", "type", "code"}}. Nodata: [DONE]follows. - Responses: an
errorevent (event: error, withcode,messageandparam), then aresponse.failedevent whoseresponse.errorhascodeandmessage. Nodata: [DONE]follows. - Messages: an
errorevent with{"type": "error", "error": {"type", "message"}}.
The official OpenAI and Anthropic SDKs raise each of these as an API error. A frame carries the code and message the same failure returns as a status on a non-streamed request, for example
upstream_error(a failed route, or a timeout),budget_exceeded,rate_limit_exceeded,validator_failedorstream_error.A reply is complete only when its stream reaches the end marker in the table above with no error frame before it, and on Responses only after
response.completedorresponse.incomplete. A stream that closes without its end marker has failed, whether or not an error frame arrived: a dropped connection can end a stream with no frame at all. Do not treat the text received so far as a finished reply; retry the request. An SDK may end its stream loop quietly when the connection closes. If yours hides the end marker, at least check that the last event arrived: a chunk withfinish_reasonon Chat Completions,response.completedorresponse.incompleteon Responses,message_stopon Messages.- Chat Completions: a
A reply cut off before it starts. A model can spend the whole output allowance reasoning and reach the cap before any visible text. The stream then ends with an error frame whose code is
invalid_request_error, the error a non-streamed request returns as422, instead of an empty reply. If you set the output allowance, the request is charged for the work the model did. If you did not, it is free within the fair-use allowance described in the Terms. Raise the output allowance and retry. When a request worked on in several steps ends its last step with no visible text, the stream ends with anupstream_errorframe instead, the error a non-streamed request returns as502. That request is charged for the work the model did; retry with backoff.Errors while a streamed request is routed. Flex holds a streamed request's
200until the request is admitted to its first route, so an error before that, such as402 budget_exceededwhen the request's largest possible cost exceeds your available credits, or a502,503or504, returns its real status with a JSON error body, as it would on a non-streamed request, and its credit hold is already released. If admission takes longer than the keep-alive interval, or a request worked on in several steps is refused at a later step, the200was already sent, and the error ends the stream early with an error frame instead.Streamed requests are metered like non-streamed ones (see Credits and metering). A stream that ends with an error frame is a failed request. You are charged for verified upstream work, including work done before a failure or disconnect, except failures caused by MUSCLE, which are free (see the Terms). The final usage of a stream that reached its end marker carries
billed_microcentsandbilling_pending: false; an error frame carriesusage.billing_pending: true, because the amount for a failed request is settled after the frame is sent and is on your usage record afterwards.
Stop sequences
| Endpoint | Field |
|---|---|
| Chat Completions | stop: a string or an array of strings |
| Responses | stop: a string or an array of strings (a Flex addition) |
| Messages | stop_sequences: an array of strings |
Flex ends the reply at your stop sequences itself, on every route, streamed or not, whether or not the route supports them. A route can accept stop sequences and still generate past them, so Flex never relies on the route to stop.
- The reply ends just before the first stop sequence that its text completes. The stop sequence itself is not returned. If several complete at the same character, the longest one counts.
- Matching runs over the reply text as a whole, so a match across stream chunks is found, and a streamed and a non-streamed reply end at the same place.
- Text before and after a tool call is matched separately: a stop sequence never spans a tool call. After a match, the rest of the reply is dropped, later tool calls included.
- Empty strings and repeated entries are ignored.
- A request may carry at most 16 stop sequences, counted as sent, and each may be at most 256 characters (UTF-16 code units). A longer list or a longer stop sequence returns
400 invalid_request_errorbefore the request is routed. - Flex also sends the first 4 stop sequences, after dropping empty and repeated ones, to the route, so a route that honours them can stop early. The reply ends in the same place whether or not the route honours them.
How the reply reports a match:
| Endpoint | Reported as |
|---|---|
| Chat Completions | finish_reason: "stop" |
| Responses | status: "completed" |
| Messages | stop_reason: "stop_sequence" and the matched stop_sequence when Flex finds the match; stop_reason: "end_turn" and stop_sequence: null when the route stops first |
A route that honours stop sequences stops before it generates one of those it was sent, so Flex never sees the match. On Messages the reply then reports stop_reason: "end_turn" and stop_sequence: null, like a reply that ended on its own. Which report you get depends on the route, so it can differ between identical requests. The reply text ends before the first stop sequence either way; do not rely on stop_reason or stop_sequence to tell whether one matched.
When the route's reply ends on the content policy or fails, that ending is reported instead, even after a match.
Streaming. Flex holds back only a trailing fragment that could still become a stop sequence (never longer than your longest one) and sends it as soon as it cannot. After a match, your reply is complete, but Flex keeps reading the route until it finishes so that usage is final. A stream may send only keep-alives during that wait, the request's total time limit still applies, and an error in the unseen remainder still fails the request: the stream then ends with an error frame, without its end marker (see Streaming).
Billing. Usage and charges are what the route generated, which can include text after the stop sequence that you do not receive. A route that honours the stop sequences it is sent stops generating sooner; one that ignores them keeps generating, and that text is billed too.
The compatibility matrix row for stop sequences reads Pending until a verification run on the deployed beta passes.
Structured outputs
Not yet turned on
Structured output is not yet turned on for the beta. Until the compatibility matrix says otherwise, requests behave as described under While it is off.
While it is off
- Chat Completions and Responses: a
response_formatfield with any value other thannull, even{"type": "text"}, returns400 unsupported_field. Omit the field. - Responses
text.formatand Messagesoutput_formatare ignored. The reply is free text, and no error is returned.
Once it is on
| Endpoint | Field | Types |
|---|---|---|
| Chat Completions | response_format | text, json_object, or json_schema with a json_schema object: name, description, schema, strict |
| Responses | text.format, or a Chat-style response_format (not both) | text, json_object, or json_schema with name, description, schema, strict on the format |
| Messages | output_format or output_config.format (not both) | json_schema with a schema |
- On Chat Completions and Responses,
{"type": "text"}asks for an ordinary reply. json_objectasks the route for a JSON object. When no message mentions JSON, Flex adds a one-line instruction to reply with a single JSON object.json_schema: Flex does not enforce the schema. A route that a probe showed enforcing strict schemas receives yourjson_schemaas sent. Every other route is asked for JSON, and Flex adds the schema to the system instructions, asking the model to follow it. There the reply can still differ from the schema, andstrict: truedoes not change that. You do not choose the route, so validate every reply against your schema.schemamust be a JSON object,namemust be 1 to 64 letters, digits, underscores or dashes (it defaults toresponse),descriptionmust be a string andstricta boolean. Anything else returns400 invalid_request_error. Any othertypereturns400 unsupported_field.- Only routes that offer JSON output serve these requests, and a single route answers each one. When no such route can serve the request, Flex returns
400 unsupported_fieldwith the message "Structured output is not available for this request. Remove the JSON response format and retry." - The JSON arrives as ordinary text:
message.contenton Chat Completions, anoutput_textpart on Responses, atextblock on Messages. Flex does not validate it. A reply cut short by the output allowance, at most 32,768 tokens when you set none (finish_reason: "length",status: "incomplete"orstop_reason: "max_tokens"), or by a stop sequence, may not be valid JSON. - Any instruction or schema Flex adds is charged as input tokens.
Tool calls
Tools use the native shape of each endpoint: {"type": "function", "function": {...}} on Chat Completions, {"type": "function", "name", "parameters"} on Responses, and {"name", "description", "input_schema"} on Messages. Send each tool result back the same way: a tool message with the matching tool_call_id, a function_call_output input item with the matching call_id, or a tool_result block with the matching tool_use_id. Flex runs no tools itself.
Errors
Every response, errors included, carries an X-Request-Id header. Flex assigns it and does not reuse a value you send. Keep it when you contact support.
OpenAI-shaped endpoints (Chat Completions, Responses, models and account) return:
json
{
"error": {
"code": "rate_limit_exceeded",
"message": "API key rate limit exceeded",
"type": "rate_limit_exceeded"
}
}Messages and token count return:
json
{
"type": "error",
"error": {
"type": "rate_limit_exceeded",
"message": "API key rate limit exceeded"
}
}code and type both hold the Flex error code from the table below, except for unsupported_content, whose type is invalid_request_error, and model_not_found, whose type is invalid_request_error and which adds "param": "model". Messages bodies carry the type only; for model_not_found it is not_found_error. Branch on the status and the code; the message text may change. On a streamed request, an error can instead arrive after the 200 and end the stream early; see Streaming.
| Status | Code | Meaning | What to do |
|---|---|---|---|
400 | invalid_request_error | The request is malformed or breaks a rule on this page: invalid JSON, a missing field, an invalid reasoning_effort, too many stop sequences, or a request declined by the content policy. | Fix the request. Retrying it unchanged fails again. |
400 | unsupported_field | The request uses a field, value or tool that Flex does not offer. | Find it in the compatibility matrix and remove it. |
400 | unsupported_content | The request carries an image or audio part. See Images and audio. | Send text only. Retrying it unchanged fails again. |
401 | authentication_error | The key is missing, mistyped or revoked, or was sent the wrong way for the endpoint. | See Authentication. |
402 | budget_exceeded | Your available credits, token budget or spend cap do not cover the request. | Add credits or raise the limit, then retry. See Credits and metering. |
403 | permission_denied | The account is suspended or not active. | Contact support with the X-Request-Id. |
404 | model_not_found | The model value is not one Flex accepts. See The model id. | Send muscle/auto. Retrying it unchanged fails again. |
413 | invalid_request_error | The request body is larger than the size limit. | Send less, for example a shorter conversation. |
422 | invalid_request_error | The reply reached the output allowance before any visible text: the model spent it all reasoning. | Raise the output allowance and retry. |
429 | rate_limit_exceeded | A rate limit was reached. See Rate limits. | Wait for Retry-After seconds when present, otherwise back off, then retry. |
500 | internal_server_error | Flex failed. | Retry with backoff. If it persists, contact support with the X-Request-Id. |
502 | upstream_error | The route failed while serving the request. | Retry with backoff. |
502 | no_visible_answer | The route finished without a visible answer. If the message says the max_tokens budget was used up, your own limit ended the answer. | Retry with backoff; if the message names max_tokens, raise it first. |
503 | upstream_error | No route could take the request at that moment. | Retry later. If it persists, contact support with the X-Request-Id. |
503 | feature_unavailable | No route supports what the request needs, or it is too large for every route's context window. | Change the request: shorten the input, lower the output allowance, or remove the capability. Retrying it unchanged fails again. |
504 | upstream_error | The request ran past its time limit. | Retry, with a smaller output allowance if you can. |
A response with x-should-retry: false cannot succeed if it is retried unchanged. The official OpenAI and Anthropic SDKs read this header and skip their automatic retries.
Content policy
A request the content policy declines ends in one of these ways.
- Declined before routing:
400 invalid_request_errorwith the message "Request blocked by content filter". - Declined by the route:
400 invalid_request_errorwith other message text. Branch on the status and code, not the message. - Declined during the reply: the reply ends with
finish_reason: "content_filter"on Chat Completions,status: "incomplete"withincomplete_details.reason: "content_filter"on Responses, andstop_reason: "refusal"on Messages. When no text was produced, the reply text is: "This request was blocked by the content policy. If you believe this is a mistake, contact support."
An account that keeps sending declined requests may be suspended; its requests then return 403 permission_denied.
Rate limits
Flex limits requests per minute for each API key, for your whole account, and for each client IP address. The limits depend on your plan; contact support to review them.
A request over a limit returns 429 rate_limit_exceeded with two headers:
| Header | Value |
|---|---|
Retry-After | Seconds to wait before retrying. |
X-RateLimit-Reason | Which limit was reached: key_requests_per_minute (this key), tenant_requests_per_minute (your account), ip_requests_per_minute or pre_auth_ip_requests_per_minute (this IP address). Other values name a limit set for your account. |
A 429 without these headers means routing capacity was briefly exhausted. Back off for a few seconds and retry.
To stay under the limits, cap concurrency, spread bursts, and honour Retry-After instead of retrying at once.
Credits and metering
Flex runs on a prepaid balance of credits. One credit is US$0.01.
- Metering. Each charged request is priced exactly, in microcents: one credit is 1,000,000 microcents. The price follows the usage the route reports.
- Debits in whole credits. A request's charge is added to your accrued usage. Flex debits whole credits as the accrued usage fills them and keeps the remainder, always below one credit, for the next request. A request that costs a fraction of a credit is not rounded up. For example, three requests metered at 0.4, 0.4 and 0.3 of a credit debit nothing, nothing, then one credit, and leave 0.1 of a credit accrued.
- Holds. Before a charged request is routed, Flex holds credits against an estimate of what it could cost. The hold is released when the request settles, and only the metered charge is kept.
- Admission. A charged request needs available credits: your balance, minus credits on hold, minus one credit whenever usage has accrued (that credit is owed). With none available, the request returns
402 budget_exceededbefore it is routed. A request whose largest possible cost exceeds what is available also returns402, streamed or not. On a streamed request, that402ends the stream early instead only when the check takes longer than one keep-alive interval, or on a request worked on in several steps (see Streaming). The largest possible cost counts the longest reply the request allows: your output allowance (the raised one where The model id says it is raised) or, without one, at most 32,768 output tokens. A few routes can bill past the allowance; for those it counts the longest reply the route can produce. Set a smaller output allowance to keep it small. If your balance is zero or negative, a charged request returns402 budget_exceededuntil a top-up covers what you owe. - Account limits. A token budget or spend cap on your account also returns
402 budget_exceededwhen the request would cross it. The spend cap is checked before each request and before each upstream call it makes. Requests already in flight when the cap is reached can finish, so under concurrency your spend can slightly exceed the cap. - Failed requests. You are charged for the upstream work a request actually used, including a request that fails after the model did work, a failed fallback attempt and the seats of a panel. Failures caused by MUSCLE are free (see the Terms). Requests that end at a limit MUSCLE set are free within a fair-use allowance for your account; beyond it they are charged. A charge that is not yet proven stays unbilled, and a later verified receipt adds its charge once. Safety screening in enforcing mode costs $0.0000135 per request, including refusals. On a successful JSON reply,
usage.billed_microcentsis the amount charged for the request in microcents andusage.billing_pendingistruewhile that amount is not yet final. An error reply's body is not changed: read what a failed request was charged from your portal usage. A charge settled after the reply was sent appears in your portal usage. - Not charged.
GET /v1/models,GET /v1/models/{model},GET /v1/accountand the token count are never charged, and they keep working when your balance is used up.
GET /v1/account reports the balance:
json
{
"tenant": { "id": "<ACCOUNT_ID>", "name": "<ACCOUNT_NAME>", "key_mode": "managed" },
"credits": {
"balance_credits": 2500,
"reserved_credits": 12,
"available_credits": 2487,
"monthly_allotment_credits": 0,
"current_period_end": null,
"accrued_usage_microcents": 431200
},
"key": { "name": "<KEY_NAME>", "last4": "<LAST4>" }
}| Field | Meaning |
|---|---|
balance_credits | Credits in your balance. |
reserved_credits | Credits on hold for requests in progress. |
available_credits | What a new request can use: balance, minus holds, minus one credit while usage is accrued. |
accrued_usage_microcents | Metered usage not yet debited, always below one credit. |
monthly_allotment_credits | Credits your plan adds each period. |
current_period_end | When the current period ends, or null. |
The portal shows metered usage next to the whole credits debited.
Support
When a request misbehaves, send support:
- the
X-Request-Id; - the time, with its time zone;
- the endpoint, the HTTP status and the error code;
- whether the request streamed;
- your client or SDK and its version.
Never send API keys, full Authorization or x-api-key headers, or private prompts and code.