Gateway contract
What you send
The public contract is OpenAI-shaped JSON. POST /v1/chat/completions works for every alias on the models page, including Claude and Gemini. POST /v1/responses (and the same path as a WebSocket) is for OpenAI models and grok-4.6 — those two providers already speak Responses natively. There is no public /v1/messages or /v1/interactions on the gateway. Claude and Gemini are translated at the edge so billing still reads OpenAI cached_tokens.
Chat Completions
Point any OpenAI client at the gateway. Same shape for GPT-5.6, grok-4.6, Claude, and Gemini.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://gateway.codegraff.com/v1",
apiKey: process.env.CODEGRAFF_API_KEY, // cg_sk_...
});
const res = await client.chat.completions.create({
model: "grok-4.6", // or gpt-5.6, claude-opus-4.8, gemini-3.7-flash
messages: [{ role: "user", content: "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}" }],
reasoning_effort: "high",
});
console.log(res.choices[0].message.content);curl https://gateway.codegraff.com/v1/chat/completions \
-H "Authorization: Bearer cg_sk_your_key" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-5.6","messages":[{"role":"user","content":"Hello"}]}'Responses (OpenAI + Grok)
Official Grok and GPT examples use client.responses.create. That maps 1:1 on the gateway for gpt-5.6 / gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna and grok-4.6. We rewrite the alias to the upstream model id and POST to api.openai.com/v1/responses or api.x.ai/v1/responses. Claude and Gemini return 400 here — use Chat Completions for those.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://gateway.codegraff.com/v1",
apiKey: process.env.CODEGRAFF_API_KEY,
});
const response = await client.responses.create({
model: "grok-4.6",
prompt_cache_key: "session:bugfix-v1",
input: [
{
role: "user",
content:
"Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}",
},
],
});
console.log(response.output_text);curl https://gateway.codegraff.com/v1/responses \
-H "Authorization: Bearer cg_sk_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6",
"prompt_cache_key": "session:bugfix-v1",
"input": "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}"
}'WebSocket
GET /v1/responses with Upgrade: websocket keeps one client socket open. We accept that immediately, then open an upstream WebSocket to OpenAI or xAI on the first response.create — after we know which host to use. If the upstream WS fails, that turn falls back to HTTP/SSE; your socket stays up. Switching from gpt-5.6 to grok-4.6 on the same client socket reconnects upstream.
Auth: Authorization: Bearer cg_sk_... from Node, or pass the key as a WebSocket subprotocol from a browser (browsers cannot set that header). Each block below has a copy button.
// gpt-5.6 — Node 22+
const ws = new WebSocket("wss://gateway.codegraff.com/v1/responses", {
headers: { Authorization: `Bearer ${process.env.CODEGRAFF_API_KEY}` },
});
ws.addEventListener("open", () => {
ws.send(JSON.stringify({
type: "response.create",
model: "gpt-5.6",
prompt_cache_key: "session:ws-gpt56",
input: "Reply with the single word pong",
}));
});
ws.addEventListener("message", (ev) => {
const event = JSON.parse(String(ev.data));
if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
if (event.type === "response.completed") ws.close();
});// grok-4.6 — Node 22+
const ws = new WebSocket("wss://gateway.codegraff.com/v1/responses", {
headers: { Authorization: `Bearer ${process.env.CODEGRAFF_API_KEY}` },
});
ws.addEventListener("open", () => {
ws.send(JSON.stringify({
type: "response.create",
model: "grok-4.6",
prompt_cache_key: "session:ws-grok46",
input: "Reply with the single word pong",
}));
});
ws.addEventListener("message", (ev) => {
const event = JSON.parse(String(ev.data));
if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
if (event.type === "response.completed") ws.close();
});// Browser — key must be a subprotocol, not an Authorization header
const KEY = "cg_sk_your_key";
const ws = new WebSocket("wss://gateway.codegraff.com/v1/responses", [KEY]);
ws.onopen = () => {
ws.send(JSON.stringify({
type: "response.create",
model: "gpt-5.6", // or "grok-4.6"
input: "Reply with the single word pong",
}));
};
ws.onmessage = (ev) => {
const event = JSON.parse(String(ev.data));
if (event.type === "response.output_text.delta") {
document.body.append(event.delta);
}
};Anthropic Messages
Do not point the Anthropic SDK at the gateway. Send the same OpenAI Chat Completions body; we translate it to Anthropic Messages so cache reads show up as cached_tokens. You send:
{
"model": "claude-opus-4.8",
"messages": [{ "role": "user", "content": "Summarize this file" }],
"reasoning_effort": "high",
"tools": [{ "type": "function", "function": { "name": "lookup", "parameters": {} } }]
}We POST that to https://api.anthropic.com/v1/messages as Messages JSON (adaptive thinking, ephemeral cache_control on system / tools / last turn) and rewrite the reply back to a chat.completion. The Anthropic SDK's /v1/messages path is not a public gateway route.
Gemini Interactions
Same idea for Gemini. You send OpenAI Chat Completions; we translate to the Interactions API (not generateContent) so implicit cache hits land in total_cached_tokens.
const res = await client.chat.completions.create({
model: "gemini-3.7-flash",
messages: [{ role: "user", content: "Continue from last turn" }],
// Optional: continue a stored Interactions thread for cache hits.
previous_interaction_id: process.env.GEMINI_THREAD,
});Upstream that becomes a POST to Google's Interactions API with store on by default and previous_interaction_id when you pass it. Thought tokens bill once as output. There is no public gateway /v1/interactions route.
Prompt caching
Stable prefix first, read usage.prompt_tokens_details.cached_tokens, billed at the cached-input rate. Optional prompt_cache_key. Details: prompt caching.
Long-context price breaks
Tools we reject
Caller-defined function tools pass through. Hosted tools that bill off the token meter are rejected: OpenAI web_search / file_search / code_interpreter, xAI web_search / x_search / code_execution, Gemini google_search / google_maps / url_context and the rest of that set. Use our POST /v1/search if you need live web results on the token meter.