Docs/Getting Started/Models

Models

Requests run through the codegraff gateway, which proxies these model aliases to their upstream providers. Prices are per million tokens (USD), read live from the gateway catalog.

Up or down: status.

Looking for Graff subscription login or direct API connections? See Model providers for setup and route inspection.

Available models

ModelProviderContextInputOutputCached in
claude-opus-5Anthropic1M$5.00$25.00$0.5
claude-sonnet-5Anthropic1M$2.00$10.00$0.2
deepseek-flashDeepSeek1M$0.3$1.20$0.006
deepseek-v4-flashDeepSeek1M$0.3$1.20$0.006
deepseek-v4-flash-fastDeepSeek1M$0.3234$0.6468$0.0809
deepseek-v4-proDeepSeek1M$0.435$0.87$0.0036
gemini-3.7-flashGemini1M$0.75$3.75$0.075
gemini-3.8-flashGemini—$0.75$3.75$0.075
glm-5.2Zhipu128K$1.68$5.28$0.312
glm-5.3Zhipu—$1.68$5.28$0.312
glm-5.3-flashZhipu1M$0.075$0.25$0.015
gpt-5.5OpenAI1.05M$4.00$20.00$0.4
gpt-5.6OpenAI1.05M$4.00$20.00$0.4
gpt-5.6-lunaOpenAI1.05M$0.2$1.20$0.02
gpt-5.6-solOpenAI1.05M$4.00$20.00$0.4
gpt-5.6-terraOpenAI1.05M$2.00$12.00$0.2
gpt-6-astraOpenAI1.05M$10.00$50.00$1.00
gpt-6-lunaOpenAI1.05M$0.1$0.5$0.01
gpt-6-solOpenAI1.05M$2.00$10.00$0.2
grok-4.6xAI500K$2.00$6.00$0.5
grok-4.7xAI500K$2.00$6.00$0.5
grok-buildxAI256K$1.00$2.00$0.2
hy4-previewTokenHub—$1.00$3.00$0.0504
jevTypeSafe64K$0.042$0$0
jev-1.13.0TypeSafe64K$0.042$0$0
jev-latestTypeSafe64K$0.042$0$0
jev-previewTypeSafe64K$0.042$0$0
kimi-k2.6Moonshot256K$1.14$4.80$0.228
kimi-k2.7-codeMoonshot256K$1.14$4.80$0.228
kimi-k2.7-code-highspeedMoonshot1M$2.28$9.60$0.456
kimi-k3Moonshot1M$3.60$18.00$0.36
ling-3.0-flash-finInclusionAI256K$0$0$0
mimo-v2.6-flashXiaomi1M$0.147$0.294$0.0029
mimo-v2.6-proXiaomi1M$0.4567$0.9135$0.0038
mimo-v2.6-pro-ultraspeedXiaomi1M$4.57$9.13$0.0378
minimax-m3MiniMax1M$0.72$2.88$0.144
muse-spark-1.2Meta1M$1.25$4.25$0.15
muse-spark-1.2-contributorMeta (Contributor)1M$0.1$0.2$0.002
muse-spark-1.3Meta1M$1.25$4.25$0.15
muse-spark-1.3-contributorMeta (Contributor)1M$0.1$0.2$0.002
qwen-3.8-27bCerebras128K$0.99$1.49$0.99

muse-spark-1.3 reasoning

muse-spark-1.3 accepts reasoning_effort: "max" as the top tier (Chat Completions top-level; Responses as reasoning.effort). none still 400s — Spark always reasons. Older 1.2 aliases that still hit a 1.2 upstream, and the contributor slug, do not have max (Meta 400s); the gateway sends xhigh instead.

Muse Spark contributor data use

Requests sent to muse-spark-1.3-contributor may be used by Meta to train future models. Use muse-spark-1.3 when prompts and completions must not be used for Meta model training. The older muse-spark-1.2 and muse-spark-1.2-contributor aliases still work and route to 1.3.

MiMo V2.6

mimo-v2.6-pro, mimo-v2.6-flash, and mimo-v2.6-pro-ultraspeed speak Chat Completions and Responses. Cache is automatic prefix matching on both; prompt_cache_key is stripped on Responses because Xiaomi does not document it. Hits bill at the Cached in rate. Keep prior reasoning_content on Chat Completions tool loops. Reasoning is currently an on/off switch: Chat Completions uses thinking.type, and Responses uses reasoning.effort (none means off). Low, medium, and high all enable the same reasoning mode; they do not set distinct intensity or speed. For OpenAI-shaped Chat clients, the gateway maps reasoning_effort to the on/off thinking.type switch.

grok-4.7

grok-4.7 is xAI's flagship: 500K context at $2 / $6 / $0.50. Prompts of 200k tokens or more bill the whole request at $4 / $12 / $1.00 — 200k is a price break, not the window. Use prompt_cache_key on Responses so cache hits stick. Hosted xAI search and code execution are blocked; function tools still work. grok-4.6 stays at the same rates.

GPT-6 family

gpt-6-astra is the flagship, gpt-6-sol the everyday tier, and gpt-6-luna the fastest and cheapest. All three have a 1.05M window, take image input, and speak Chat Completions and Responses. Prompts over 272k tokens bill the whole request at the long-context rate (2× input and cache, 1.5× output). Writing to the prompt cache costs 1.25× the input rate; reading it back bills at the Cached in rate. All three can search the web.

GPT-5.6 family

gpt-5.6 and gpt-5.6-sol are Sol. gpt-5.6-terra and gpt-5.6-luna are cheaper tiers of the same 1.05M window. gpt-5.5 still works as a Sol compatibility alias. Prompts of 272k tokens or more bill the whole request at the long-context rate (2× input, 1.5× output, 2× cache) — 272k is a price break, not the context window.

Prompt caching

Cached input tokens bill at the Cached in rate. One OpenAI-shaped contract for every provider — see prompt caching. Gemini built-in Search, Maps, computer-use, file-search, code-execution, and URL-context tools are blocked; your function tools still work. Gemini 3.8 Flash intro rates run through 2026-12-31.

Using a model

Pass any alias as --model on the CLI, as the model option in the SDKs, or as the model field of an OpenAI-compatible request. Chat Completions works for every row; Responses (HTTP or WebSocket) is for OpenAI, grok-4.7 / grok-4.6, Muse Spark, and mimo-v2.6-pro / mimo-v2.6-flash / mimo-v2.6-pro-ultraspeed. See the gateway contract for copy-paste examples, Anthropic/Gemini translation, and prompt caching.

OpenAI models (gpt-6-astra, gpt-6-sol, gpt-6-luna and the gpt-5.6 family) can search the live web while they answer. Add {"type": "web_search"} to tools on a Responses request, over HTTP or WebSocket. The model decides when to search and cites its sources:

bash
curl https://gateway.codegraff.com/v1/responses \
  -H "Authorization: Bearer cg_sk_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6-luna",
    "tools": [{"type": "web_search"}],
    "input": "What changed in the latest Node.js release?"
  }'

With the OpenAI SDK pointed at the gateway, it is the same call you would make to OpenAI:

python
from openai import OpenAI

client = OpenAI(base_url="https://gateway.codegraff.com/v1", api_key="cg_sk_your_key")
response = client.responses.create(
    model="gpt-6-sol",
    tools=[{"type": "web_search"}],
    input="Summarize this week's biggest AI announcements, with sources.",
)
print(response.output_text)

How web search is billed

Each search the model runs costs $0.011 and appears as its own web_searchline in your usage. Opening a page or searching within one is free. The page text the model reads is billed as ordinary input tokens at the model's rate. A request runs at most 10 tool calls (set max_tool_calls for fewer). Credits for the worst case are reserved up front and settled to what was used.

Web search needs an OpenAI model and the Responses API. On other models, or on Chat Completions, use the POST /v1/search endpoint instead.