Docs/Gateway/Prompt caching

Prompt caching

You send one OpenAI-shaped request. Put a stable prefix first, optionally pass prompt_cache_key, and read cached_tokens. Hits bill at the cached-input rate on the models page.

The public contract

Same JSON for every alias. Put stable content first. Optional prompt_cache_key. Every response uses OpenAI usage shape.

json
{
  "usage": {
    "prompt_tokens": 1847,
    "completion_tokens": 98,
    "total_tokens": 1945,
    "prompt_tokens_details": { "cached_tokens": 1792 }
  }
}

On POST /v1/responses the same count is usage.input_tokens_details.cached_tokens. In both cases cached_tokens is a subset of input tokens, not an extra line. Zero is expected on the first request with a new prefix, or after eviction.

Structure the prefix

Caches match from the start of the tokenized prompt. Tokens after the first difference are computed fresh.

  • System prompt, instructions, and few-shot examples first.
  • Tool definitions next — they should not change every turn.
  • Conversation history grows as a prefix; only the new user turn is hot.
  • Volatile bits (timestamps, the current question) last.

This ordering is the main lever on every alias.

prompt_cache_key

Optional. A stable string that names the shared prefix — an app or use case, not a per-user or per-request id. We forward it on Chat Completions and Responses where the upstream uses it. A hit still happens on the model host.

ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://gateway.codegraff.com/v1",
  apiKey: process.env.CODEGRAFF_API_KEY, // cg_sk_...
});

const res = await client.chat.completions.create({
  model: "muse-spark-1.3", // or gpt-5.6, grok-4.6, glm-5.3-flash, kimi-k3
  prompt_cache_key: "my-app-system-prompt",
  messages: [
    {
      role: "system",
      content:
        "You are a senior technical support engineer. Always give step-by-step instructions.",
    },
    { role: "user", content: "How do I rotate API keys without downtime?" },
  ],
});
console.log(res.usage?.prompt_tokens_details?.cached_tokens);

Responses-only retention

On POST /v1/responses, prompt_cache_retention may be in_memory (default) or 24h. It is a hint, not a guarantee. Do not send it on Chat Completions.

See what was cached

  • Most of prompt_tokens cached — the prefix is working. Only the new turn needed prefill.
  • Zero — first request, eviction, or the prefix changed early (edited system prompt).
  • Lower than expected — something in the middle of the prefix moved. Check order.

Streaming still reports usage: we force stream_options.include_usage so the final chunk carries cached_tokens for billing.

Billing

Uncached input tokens bill at Input. Cached input tokens bill at Cached in. Output is unchanged. We do not add a separate cache-write line item even when an upstream charges one internally — we meter the cached_tokens we receive.

Example: Muse Spark 1.3 list is $1.25 / $0.15 cached / $4.25 per million. Contributor is $0.10 / $0.002 / $0.20. A 1,792-token cached prefix of 1,847 input is mostly the $0.15 rate.

Long-context price breaks

GPT-5.6 flips the whole request to long rates at 272k prompt tokens. grok-4.6 flips at 200k. Those are price breaks, not context windows — see the models page.