Prompt caching
The public contract
Same JSON for every alias. Put stable content first. Optional prompt_cache_key. Every response uses OpenAI usage shape.
{
"usage": {
"prompt_tokens": 1847,
"completion_tokens": 98,
"total_tokens": 1945,
"prompt_tokens_details": { "cached_tokens": 1792 }
}
}On POST /v1/responses the same count is usage.input_tokens_details.cached_tokens. In both cases cached_tokens is a subset of input tokens, not an extra line. Zero is expected on the first request with a new prefix, or after eviction.
Structure the prefix
Caches match from the start of the tokenized prompt. Tokens after the first difference are computed fresh.
- System prompt, instructions, and few-shot examples first.
- Tool definitions next — they should not change every turn.
- Conversation history grows as a prefix; only the new user turn is hot.
- Volatile bits (timestamps, the current question) last.
This ordering is the main lever on every alias.
prompt_cache_key
Optional. A stable string that names the shared prefix — an app or use case, not a per-user or per-request id. We forward it on Chat Completions and Responses where the upstream uses it. A hit still happens on the model host.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://gateway.codegraff.com/v1",
apiKey: process.env.CODEGRAFF_API_KEY, // cg_sk_...
});
const res = await client.chat.completions.create({
model: "muse-spark-1.3", // or gpt-5.6, grok-4.6, glm-5.3-flash, kimi-k3
prompt_cache_key: "my-app-system-prompt",
messages: [
{
role: "system",
content:
"You are a senior technical support engineer. Always give step-by-step instructions.",
},
{ role: "user", content: "How do I rotate API keys without downtime?" },
],
});
console.log(res.usage?.prompt_tokens_details?.cached_tokens);Responses-only retention
POST /v1/responses, prompt_cache_retention may be in_memory (default) or 24h. It is a hint, not a guarantee. Do not send it on Chat Completions.See what was cached
- Most of prompt_tokens cached — the prefix is working. Only the new turn needed prefill.
- Zero — first request, eviction, or the prefix changed early (edited system prompt).
- Lower than expected — something in the middle of the prefix moved. Check order.
Streaming still reports usage: we force stream_options.include_usage so the final chunk carries cached_tokens for billing.
Billing
Uncached input tokens bill at Input. Cached input tokens bill at Cached in. Output is unchanged. We do not add a separate cache-write line item even when an upstream charges one internally — we meter the cached_tokens we receive.
Example: Muse Spark 1.3 list is $1.25 / $0.15 cached / $4.25 per million. Contributor is $0.10 / $0.002 / $0.20. A 1,792-token cached prefix of 1,847 input is mostly the $0.15 rate.
Long-context price breaks