GPT‑6.1 Sol Standard API rates are $2 input, $0.10 cached input and $10 output per million tokens for short-context requests. Regular input and output rates are 80% below GPT‑6 Astra and equal to GPT‑6 Sol. A migration decision still depends on task success, reasoning usage and retries.
How much does GPT‑6.1 Sol cost?
Official Standard API prices below are USD per million tokens for input lengths up to 272K. These are not ChatGPT subscription prices or a gateway's final rates.
| Model | Regular input | Cache read | Output |
|---|---|---|---|
| GPT‑6.1 Sol | $2 | $0.10 | $10 |
| GPT‑6 Sol | $2 | $0.20 | $10 |
| GPT‑6 Astra | $10 | $1 | $50 |
Sources: OpenAI pricing and the GPT‑6 Sol model page. Rates can change. Cache writes, tools and other billable features are additional.
What does one request cost? Three worked examples
For short-context requests, excluding cache writes and tools: cost = (uncached input × 2 + cache reads × 0.10 + output × 10) / 1,000,000. These are our calculations from published rates, not measured bills.
- 10,000 input + 2,000 output: Sol costs $0.04 per request, or $40 for 1,000 requests. Astra costs $0.20 per request with identical usage. That comparison assumes identical token consumption.
- 10,000 fresh input + 90,000 cache reads + 2,000 output: Sol costs $0.049, versus $0.22 if all 100,000 input tokens are uncached. Total savings are about 77.7%; a 95% reduction in the cached-input rate is not a 95% reduction in the whole bill.
- 300,000 input + 10,000 output: This crosses the 272K threshold. Applying the long-context $4 input / $15 output rates to the full request gives $1.35, rather than the $0.70 short-context calculation.
Reasoning consumes tokens too. usage.output_tokens includes billed output; reasoning_tokens in its details is a subset, so do not add it twice. max_output_tokens covers both visible output and reasoning. Too small a budget can leave a long task without a complete answer. See the official reasoning guide.
6.1 Sol vs Astra, 6 Sol and Opus 5.5: which should you choose?
OpenAI's launch report shows Sol approaching or exceeding more expensive models on some coding and office tasks. Those are vendor results under specific evaluation settings, not a guarantee for your workload. We have not run independent comparative benchmarks for this article.
| Your situation | Suggested next step | What to check |
|---|---|---|
| Astra is expensive for routine work | Sample routine tasks on Sol; retain Astra as a hard-task baseline | Success rate and retry costs |
| You already use GPT‑6 Sol | Test a replacement at the same regular rates | Consistent quality and cache hits |
| You are comparing Opus 5.5 | Run the same real tasks on both | Tool behavior, format compliance, completion time and bills |
| You search “6.1 Sol vs 5.6 Sol” | Keep old results as a baseline before switching | Prompt assumptions, long conversations and failure cases |
Our suggested starting point is 20 real tasks with explicit acceptance criteria. Keep inputs, tool permissions and output constraints comparable, and repeat each task several times. Record first-pass success, time to first token, median and 95th-percentile total duration, and total cost per accepted result. Include failed attempts. A single impressive answer is not a migration test. Region, service tier and tools affect latency; the model name alone does not promise greater speed.
How do you call the GPT‑6.1 Sol API?
The official model ID is gpt-6.1-sol. It accepts text and images and produces text, with a 1,050,000-token context window and up to 128,000 output tokens. Use Responses for tool workflows; Chat Completions supports calls without tools. See the model documentation.
This minimal example uses the official Responses API. Set OPENAI_API_KEY in your server or local terminal environment. Never put a secret key in a public webpage.
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6.1-sol",
"input": "Explain a safe rollout plan for a database migration.",
"reasoning": { "effort": "medium" },
"max_output_tokens": 4096
}'
For Token King or another gateway, confirm its base URL, model ID, Responses support and live rates before replacing the official endpoint and key. Gateway model names may differ from the official ID.
What is “6.1 Sol ultra”? Which reasoning effort works?
The public API checked for this guide supports low, medium (default), high, xhigh and max. Do not send ultra just because it appears in search suggestions. none and minimal are also outside this model's supported list. A client's option label does not establish API support for the same value.
Start with medium. Try low for simple classification or extraction and inspect mistakes. Try high on multistep tasks that repeatedly fail, then check whether its additional cost buys better acceptance. Change one setting at a time to distinguish model gains from prompt changes or larger reasoning budgets.
What should you check before migrating?
- Confirm that your account or gateway actually exposes the model.
- Update the model ID and check SDK support for Responses and reasoning parameters.
- Test your real long-context, image, tool and structured-output workflows on small traffic.
- Keep a fallback to the previous model and set output and spending limits.
- Reconcile usage with bills, then expand only after your quality and cost targets pass.
Frequently asked questions
Does 80% lower pricing than Astra mean 80% cheaper tasks?
No. It describes the regular input and output rates in the table. Reasoning volume, retries, tools and long-context pricing can change the task's final cost.
Is GPT‑6.1 Sol worth upgrading to?
From GPT‑6 Sol, regular input and output rates are unchanged, so test quality first. From Astra, routine-task routing offers larger potential savings, but preserve a quality baseline for difficult tasks.
Does API pricing include a ChatGPT subscription?
No. This guide uses API token billing. Subscription quotas, client allowances and gateway balances are separate pricing systems.