Docs/Gateway/Gateway contract

Gateway contract

You talk to Codegraff. We hold the upstream contracts with OpenAI, Anthropic, Gemini, and xAI. Every public call uses https://gateway.codegraff.com/v1 and a cg_sk_ key. You do not need an OpenAI, Anthropic, Google, or xAI key.

What you send

The public contract is OpenAI-shaped JSON. POST /v1/chat/completions works for every alias on the models page, including Claude and Gemini. POST /v1/responses (and the same path as a WebSocket) is for OpenAI models and grok-4.6 — those two providers already speak Responses natively. There is no public /v1/messages or /v1/interactions on the gateway. Claude and Gemini are translated at the edge so billing still reads OpenAI cached_tokens.

Chat Completions

Point any OpenAI client at the gateway. Same shape for GPT-5.6, grok-4.6, Claude, and Gemini.

ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://gateway.codegraff.com/v1",
  apiKey: process.env.CODEGRAFF_API_KEY, // cg_sk_...
});

const res = await client.chat.completions.create({
  model: "grok-4.6", // or gpt-5.6, claude-opus-4.8, gemini-3.7-flash
  messages: [{ role: "user", content: "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}" }],
  reasoning_effort: "high",
});
console.log(res.choices[0].message.content);
bash
curl https://gateway.codegraff.com/v1/chat/completions \
  -H "Authorization: Bearer cg_sk_your_key" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-5.6","messages":[{"role":"user","content":"Hello"}]}'

Responses (OpenAI + Grok)

Official Grok and GPT examples use client.responses.create. That maps 1:1 on the gateway for gpt-5.6 / gpt-5.6-sol / gpt-5.6-terra / gpt-5.6-luna and grok-4.6. We rewrite the alias to the upstream model id and POST to api.openai.com/v1/responses or api.x.ai/v1/responses. Claude and Gemini return 400 here — use Chat Completions for those.

ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://gateway.codegraff.com/v1",
  apiKey: process.env.CODEGRAFF_API_KEY,
});

const response = await client.responses.create({
  model: "grok-4.6",
  prompt_cache_key: "session:bugfix-v1",
  input: [
    {
      role: "user",
      content:
        "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}",
    },
  ],
});
console.log(response.output_text);
bash
curl https://gateway.codegraff.com/v1/responses \
  -H "Authorization: Bearer cg_sk_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6",
    "prompt_cache_key": "session:bugfix-v1",
    "input": "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}"
  }'

WebSocket

GET /v1/responses with Upgrade: websocket keeps one client socket open. We accept that immediately, then open an upstream WebSocket to OpenAI or xAI on the first response.create — after we know which host to use. If the upstream WS fails, that turn falls back to HTTP/SSE; your socket stays up. Switching from gpt-5.6 to grok-4.6 on the same client socket reconnects upstream.

Auth: Authorization: Bearer cg_sk_... from Node, or pass the key as a WebSocket subprotocol from a browser (browsers cannot set that header). Each block below has a copy button.

ts
// gpt-5.6 — Node 22+
const ws = new WebSocket("wss://gateway.codegraff.com/v1/responses", {
  headers: { Authorization: `Bearer ${process.env.CODEGRAFF_API_KEY}` },
});

ws.addEventListener("open", () => {
  ws.send(JSON.stringify({
    type: "response.create",
    model: "gpt-5.6",
    prompt_cache_key: "session:ws-gpt56",
    input: "Reply with the single word pong",
  }));
});

ws.addEventListener("message", (ev) => {
  const event = JSON.parse(String(ev.data));
  if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
  if (event.type === "response.completed") ws.close();
});
ts
// grok-4.6 — Node 22+
const ws = new WebSocket("wss://gateway.codegraff.com/v1/responses", {
  headers: { Authorization: `Bearer ${process.env.CODEGRAFF_API_KEY}` },
});

ws.addEventListener("open", () => {
  ws.send(JSON.stringify({
    type: "response.create",
    model: "grok-4.6",
    prompt_cache_key: "session:ws-grok46",
    input: "Reply with the single word pong",
  }));
});

ws.addEventListener("message", (ev) => {
  const event = JSON.parse(String(ev.data));
  if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
  if (event.type === "response.completed") ws.close();
});
ts
// Browser — key must be a subprotocol, not an Authorization header
const KEY = "cg_sk_your_key";
const ws = new WebSocket("wss://gateway.codegraff.com/v1/responses", [KEY]);

ws.onopen = () => {
  ws.send(JSON.stringify({
    type: "response.create",
    model: "gpt-5.6", // or "grok-4.6"
    input: "Reply with the single word pong",
  }));
};

ws.onmessage = (ev) => {
  const event = JSON.parse(String(ev.data));
  if (event.type === "response.output_text.delta") {
    document.body.append(event.delta);
  }
};

Anthropic Messages

Do not point the Anthropic SDK at the gateway. Send the same OpenAI Chat Completions body; we translate it to Anthropic Messages so cache reads show up as cached_tokens. You send:

json
{
  "model": "claude-opus-4.8",
  "messages": [{ "role": "user", "content": "Summarize this file" }],
  "reasoning_effort": "high",
  "tools": [{ "type": "function", "function": { "name": "lookup", "parameters": {} } }]
}

We POST that to https://api.anthropic.com/v1/messages as Messages JSON (adaptive thinking, ephemeral cache_control on system / tools / last turn) and rewrite the reply back to a chat.completion. The Anthropic SDK's /v1/messages path is not a public gateway route.

Gemini Interactions

Same idea for Gemini. You send OpenAI Chat Completions; we translate to the Interactions API (not generateContent) so implicit cache hits land in total_cached_tokens.

ts
const res = await client.chat.completions.create({
  model: "gemini-3.7-flash",
  messages: [{ role: "user", content: "Continue from last turn" }],
  // Optional: continue a stored Interactions thread for cache hits.
  previous_interaction_id: process.env.GEMINI_THREAD,
});

Upstream that becomes a POST to Google's Interactions API with store on by default and previous_interaction_id when you pass it. Thought tokens bill once as output. There is no public gateway /v1/interactions route.

Prompt caching

Stable prefix first, read usage.prompt_tokens_details.cached_tokens, billed at the cached-input rate. Optional prompt_cache_key. Details: prompt caching.

Long-context price breaks

GPT-5.6 flips the whole request to long rates at 272k prompt tokens (2× input, 1.5× output). grok-4.6 flips at 200k (2× input, 2× output, 2× cache). Those numbers are price breaks, not context windows.

Tools we reject

Caller-defined function tools pass through. Hosted tools that bill off the token meter are rejected: OpenAI web_search / file_search / code_interpreter, xAI web_search / x_search / code_execution, Gemini google_search / google_maps / url_context and the rest of that set. Use our POST /v1/search if you need live web results on the token meter.