DeepThinking AI

Which GPT-6 model should you use for AI engineering workflows?

AI Architect

Key takeaways

  • OpenAI's GPT-6 model pages say GPT-6 Astra and GPT-6.1 Sol do not support `none` reasoning effort, while GPT-6 Luna does.
  • OpenAI's GPT-6.1 Sol and GPT-6 Astra docs say tool calling belongs on the Responses API, even though Chat Completions is otherwise supported.
  • OpenAI's pricing page lists Standard short-context output at $0.50 per million tokens for GPT-6 Luna, $10 for GPT-6.1 Sol, and $50 for GPT-6 Astra.
  • OpenAI's pricing page doubles input and cache rates, and raises output 1.5x, once a GPT-6 request crosses 272K input tokens.
  • OpenAI's model-selection guide places Luna, GPT-6.1 Sol, and Astra on a ladder of reasoning effort rather than a single default winner.

OpenAI’s current GPT-6 documentation hides the most important model-selection fact in the model pages rather than in the marketing summary. The family does not differ only by intelligence and price. It also splits on reasoning modes and API behavior. GPT-6 Astra and GPT-6.1 Sol do not support none reasoning effort, while GPT-6 Luna does. Astra and GPT-6.1 Sol also push tool calling onto the Responses API.

That changes how an engineering team should choose. My opinion is that most teams should not start with Astra, even though the models page recommends it as the place to begin. Start with the lightest GPT-6 model that fits the workflow’s API and reasoning constraints, then earn a move upward with a benchmark. The family is expensive enough, and operationally different enough, that “default to the strongest model” is a procurement habit rather than a systems decision.

What does the GPT-6 compatibility table tell you first?

The first read of the GPT-6 family should be about compatibility, because that decides whether your existing integration even survives the move. OpenAI’s GPT-6 guide says GPT-6 Astra and GPT-6.1 Sol do not support none reasoning effort, and both model pages say tool calling should use the Responses API. GPT-6 Luna is the exception. Its model page allows none, low, medium, high, xhigh, and max, and says Chat Completions supports function calling only when reasoning_effort is none.

That is a stronger boundary than most “best model” explainers mention. If your application still leans on Chat Completions for simple function calls, Luna is the only GPT-6 path that lets you defer a Responses migration. If you already run tool-rich backends through Responses, the family becomes easier to compare. This matters even more in managed-runtime setups such as hosted agent backends, because the API surface becomes part of the orchestration contract rather than a local implementation detail.

Compatibility decides the first GPT-6 choice

Compatibility decides the first GPT-6 choiceDiagram: 7 ordered layers. GPT-6 Luna with none reasoning, then GPT-6 Responses tool path, then GPT-6 Luna with higher effort, then Compatibility gate (breakpoint), then GPT-6.1 Sol forbids none, then GPT-6 Astra forbids none, then Region and mode limits narrow lanes.1GPT-6 Luna with none reasoningChat Completions can call functions here.2GPT-6 Responses tool pathShared default for reasoning plus tools.3GPT-6 Luna with higher effortSame cheap model, deeper thinking.Compatibility gate4GPT-6.1 Sol forbids noneLowest supported effort is low.5GPT-6 Astra forbids noneSame Responses-first rule.6Region and mode limits narrow lanesEU residency removes some fast paths.
Show as text
Compatibility decides the first GPT-6 choice. Diagram: 7 ordered layers. GPT-6 Luna with none reasoning, then GPT-6 Responses tool path, then GPT-6 Luna with higher effort, then Compatibility gate (breakpoint), then GPT-6.1 Sol forbids none, then GPT-6 Astra forbids none, then Region and mode limits narrow lanes.
#LayerNote
1GPT-6 Luna with none reasoningChat Completions can call functions here.
2GPT-6 Responses tool pathShared default for reasoning plus tools.
3GPT-6 Luna with higher effortSame cheap model, deeper thinking.
·Compatibility gate (breakpoint)
4GPT-6.1 Sol forbids noneLowest supported effort is low.
5GPT-6 Astra forbids noneSame Responses-first rule.
6Region and mode limits narrow lanesEU residency removes some fast paths.
The family split that changes architectures first is reasoning and API support. Cost matters after you know whether the workflow can stay in Chat Completions, or has to move to Responses for tool use.

When does GPT-6 Luna beat the rest of the GPT-6 family?

GPT-6 Luna wins when the workflow is frequent, scoped, and cheap enough that token price dominates prestige. OpenAI’s pricing page lists Standard short-context output at $0.50 per million tokens, compared with $10 for GPT-6.1 Sol and $50 for Astra. The same ratio holds on input pricing. Luna is also the only GPT-6 model whose page explicitly keeps a none reasoning lane, which matters for teams that need function calling in Chat Completions or want the narrowest possible reasoning budget on routine work.

That makes Luna stronger than “small model” language suggests. OpenAI’s own model-selection page places Luna at both low and extra-high reasoning settings, depending on workload shape. My read is that Luna should be the default for bounded extraction, triage, structured rewriting, and high-volume internal automations. The failure mode is pretending Luna is a drop-in answer for multi-step tool choreography. If a job depends on several revisions or conflicting evidence, Luna can become expensive in retries, exactly the pattern that shows up in prompt-cache economics once a workflow starts looping.

When should GPT-6.1 Sol be your default engineering model?

GPT-6.1 Sol is the practical default once a workflow is complex enough that Luna’s savings stop surviving contact with reality. OpenAI describes it as “near-Astra performance for complex work at a lower cost,” and the model page keeps the same 1,050,000 token context window, 922,000 token maximum input, and 128,000 token maximum output that Astra has. Standard pricing, however, is $10 per million output tokens instead of Astra’s $50. That is still 20 times Luna on output, but it is one fifth of Astra.

The operational catch is important. GPT-6.1 Sol does not support none reasoning effort, and OpenAI says tool calling should use Responses. That means Sol is a good default only if your system already accepts a reasoning-first runtime. For AI engineering teams building code agents, browser tasks, or multi-tool research flows, that is usually acceptable. My opinion is that Sol is the best first benchmark for serious agentic work, because it captures most of the family capability without turning every weak architecture decision into an Astra bill.

Standard short-context output prices across the GPT-6 family

Standard short-context output prices across the GPT-6 familyBar chart. GPT-6 Luna: 0.5 dollars per 1M output tokens. GPT-6.1 Sol: 10 dollars per 1M output tokens. GPT-6 Astra: 50 dollars per 1M output tokens.GPT-6 Luna0.5 dollars per 1M output tokenslow-cost default for volumeGPT-6.1 Sol10 dollars per 1M output tokens20x Luna on output priceGPT-6 Astra50 dollars per 1M output tokens5x Sol and 100x Luna
Show data
Standard short-context output prices across the GPT-6 family. Bar chart. GPT-6 Luna: 0.5 dollars per 1M output tokens. GPT-6.1 Sol: 10 dollars per 1M output tokens. GPT-6 Astra: 50 dollars per 1M output tokens.
ItemValue (dollars per 1M output tokens)Note
GPT-6 Luna0.5low-cost default for volume
GPT-6.1 Sol1020x Luna on output price
GPT-6 Astra505x Sol and 100x Luna
Output-token price expands far faster than most teams budget for. The ratio is the useful fact. Astra is 5 times GPT-6.1 Sol and 100 times GPT-6 Luna on Standard short-context output.

When is GPT-6 Astra worth the premium?

GPT-6 Astra is worth the premium when the expensive part of the workflow is human repair or failed execution, rather than raw tokens. OpenAI’s GPT-6 guide says Astra achieved stronger results in several evaluations while using substantially fewer output tokens, and that its estimated API cost per task was lower than earlier models despite higher per-token pricing. That is a vendor claim, so it deserves verification on your own workload. The model and pricing pages still make the basic tariff plain: Astra costs 5 times GPT-6.1 Sol and 100 times Luna on Standard short-context output.

That premium is rational only for hard tasks. Use Astra when a run has to coordinate code, documents, browser work, and careful reasoning in one pass, or when a wrong answer triggers expensive review. Astra also gets OpenAI’s fullest speed menu, including Fast and Ultrafast modes, although the GPT-6 guide says those lanes narrow under regional processing rules. I would keep Astra behind a strict quality gate rather than a product slogan. Teams that promote it to the whole fleet usually pay for architectural uncertainty they should have benchmarked away.

How should you benchmark a GPT-6 model before standardising it?

Benchmark the family at the workflow level, because the docs already tell you the token tables. What they cannot tell you is where your system burns money: extra reasoning, repeated tool turns, long prompts, or human cleanup. Run the same workload through Luna, GPT-6.1 Sol, and Astra with fixed tools, prompt, output cap, and evaluation rubric. Record total input, total output, failures, and time to a usable result. Then rerun above and below OpenAI’s 272K short-context boundary, because crossing it doubles input and cache rates and raises output pricing by 1.5 times for the full request.

The benchmark should also include deployment lanes. Batch and Flex cut rates by half on the pricing page, but they change operating assumptions the same way OpenAI’s background mode and Batch API do. My opinion is simple: standardise on the lightest GPT-6 model that clears your quality bar with stable retries, then isolate Astra to the workflows that still fail. That keeps model choice tied to evidence instead of taste.

Do this

Choose a GPT-6 model for an AI engineering workflow

Start with compatibility, then benchmark price and quality. That order removes migration mistakes before they become operating costs.

  1. Map the workflow to the API surface first

    Decide whether the job can live in Chat Completions or needs the Responses API for tool use. GPT-6 Astra and GPT-6.1 Sol push tool-driven work to Responses immediately, while GPT-6 Luna still leaves one narrow Chat Completions lane at `reasoning_effort: none`.

  2. Pick the lowest reasoning mode that matches the task

    Use GPT-6 Luna for bounded extraction, edits, and high-volume flows. Move upward only when the workload needs multi-step judgment, conflict resolution, or sustained revision inside one run.

  3. Compare GPT-6 Luna and GPT-6.1 Sol on identical traces

    Keep the prompt, tools, and output cap fixed. Measure total tokens, wall-clock time, and failure rate. A price table alone does not tell you whether retries or bad outputs erase Luna's savings.

  4. Add GPT-6 Astra only where GPT-6.1 Sol still misses quality

    Astra's premium makes sense when the task is costly to get wrong, or when one better answer replaces many repair loops. Keep that decision local to the hard workflow rather than promoting Astra to a blanket default.

  5. Rerun the benchmark above and below 272K input tokens

    OpenAI's pricing page moves the full request onto long-context pricing past that threshold. Also test overnight runs with Batch or Flex, because [background jobs and offline jobs diverge operationally](/openai-background-mode-vs-batch-api/).

Frequently asked questions

Can all GPT-6 models use Chat Completions with tools?
No. GPT-6 Astra and GPT-6.1 Sol support Chat Completions, but their model pages say tool calling should use the Responses API. GPT-6 Luna allows Chat Completions function calling only when `reasoning_effort` is `none`.
Which GPT-6 model is cheapest for high-volume automations?
GPT-6 Luna is the clear price leader in OpenAI's pricing table. Standard short-context output is $0.50 per million tokens, versus $10 for GPT-6.1 Sol and $50 for GPT-6 Astra.
When should a team skip GPT-6 Luna and start on GPT-6.1 Sol?
Skip Luna when the workflow needs more than bounded extraction or routine edits, especially when several tools, revisions, or conflicting inputs need to be coordinated inside one run. That is the lane OpenAI describes for GPT-6.1 Sol.
Does GPT-6 Astra always cost more per task than GPT-6.1 Sol?
Not always. OpenAI's GPT-6 guide says Astra can finish some evaluations with fewer output tokens and a lower estimated cost per task. Treat that as a vendor claim until your own benchmark confirms it on your workload.
What is the easiest GPT-6 pricing mistake?
Ignoring the 272K input threshold. OpenAI's pricing page says the full request moves to long-context pricing once the prompt crosses that boundary, which can change the winner even when the model choice stays the same.

Sources

  1. Using GPT-6OpenAI
  2. Model selectionOpenAI
  3. GPT-6 AstraOpenAI
  4. GPT-6.1 SolOpenAI
  5. GPT-6 LunaOpenAI
  6. PricingOpenAI

openaigpt-6model-selectionreasoningresponses-apipricing