OpenAI Releases GPT-6 Astra, Sol and Luna in the API: Prices, Limits and How to Choose
OpenAI moved its API to a new model generation in September 2026. According to the OpenAI API changelog, GPT-6 Astra was released on September 3, 2026, GPT-6 Sol and GPT-6 Luna followed on September 22, and GPT-6.1 Sol arrived on September 29. That is four models in under four weeks, with a hundredfold price gap between the cheapest and the most expensive.
For a business, the question is rarely "which model is smartest". It is which model belongs behind which feature, what it will cost at your volume, and whether switching is worth the testing effort. Those are engineering and budget decisions, and they sit alongside the rest of the technology stack your product runs on.
This article sets out what OpenAI has published about the four models, what is generally available versus limited, and how to decide what, if anything, to change.
What OpenAI released, and when
The changelog describes GPT-6 Astra, released September 3, 2026, as OpenAI's most capable model, intended for reasoning, coding, computer use, research and document creation. The September 22 entry introduces GPT-6 Sol and GPT-6 Luna as reasoning models that take text and image input and produce text through the Responses and Chat Completions APIs. On September 25, OpenAI logged a fix for an image encoding bug that had degraded image understanding in Sol and Luna during their first days.
The September 29 entry adds GPT-6.1 Sol, positioned for complex coding and professional work at a lower cost than Astra. OpenAI's latest model guide summarizes it as near-Astra performance at a lower price. All four are listed as released API models, not previews.
The four models at a glance
OpenAI's model pages give a short description of each one. Prices below are the standard API rates in USD per one million tokens as shown on those pages when we read them on October 4, 2026. A token is a fragment of a word, so a million tokens is very roughly several hundred thousand words of text.
- GPT-6 Astra: the top model for the most demanding work. USD 10 input and USD 50 output.
- GPT-6.1 Sol: near-Astra results for complex work at lower cost. USD 2 input and USD 10 output.
- GPT-6 Sol: built for complex coding and agentic workflows. USD 2 input and USD 10 output.
- GPT-6 Luna: the most efficient model, for focused, high-volume tasks. USD 0.10 input and USD 0.50 output.
On those listed rates, GPT-6.1 Sol costs one fifth of Astra, and Luna costs one hundredth. OpenAI has changed prices more than once this year (the changelog records cuts to the earlier GPT-5.6 models on July 30 and August 21, 2026), so treat these as a snapshot and confirm on the model pages before you budget.
Pricing details that change the bill
The headline rates are not the whole picture. Each model page lists a context window of 1,050,000 tokens and up to 128,000 output tokens, which means a single request can carry a very large set of documents. But the same pages state that a request with more than 272,000 input tokens is charged at twice the input rate and one and a half times the output rate. Feeding the model everything you have is possible; it is not cheap.
Cached input works the other way. Text the model has seen in a recent request, such as a long system prompt or a product catalog, is billed at a fraction of the normal input rate. The Astra page lists cached input at USD 1 against USD 10 for fresh input. For products that send the same instructions with every request, structuring prompts so they cache well can matter more than which model you choose.
The knowledge cutoffs listed are April or May 2026 depending on the model. Anything newer has to be supplied through search or your own data.
Speed modes and a data residency caveat
The changelog entry for September 29, 2026 also adds Ultrafast mode for GPT-6 Astra in the Responses API, aimed at reducing the time between output tokens. Ultrafast had been announced on August 13, 2026 as a limited preview for select customers.
UK and European teams should read the fine print. OpenAI's latest model guide states that Fast mode is not available with EU data residency for the GPT-6 models, and that Ultrafast mode supports US data residency only. If your contracts or your own data protection assessment depend on regional processing, the faster modes may not be an option for you. This article is not legal advice; confirm the position with whoever handles data protection for your organization.
What changes for developers
Moving to GPT-6 is not always a one-line model name change. The latest model guide lists several adjustments, among them removing sampling settings such as temperature and top_p when the model is reasoning, and a renamed prompt cache setting. The guide also flags behavior differences in Astra: it is more inclined to ask the user a clarifying question, it pays closer attention to instructions found in files and skills, and it tends to write longer, more formatted answers.
Those behavior notes matter outside engineering. A support assistant that starts asking follow-up questions, or a report generator that changes its formatting, is a visible product change for your customers even if no code was touched beyond the model name.
What this means for your business
The points below are our guidance, not OpenAI's.
The wide price spread rewards matching the model to the task. Sorting support tickets, tagging products, extracting fields from invoices and similar high-volume, narrow jobs are what the cheapest tier is described as being for. Multi-step work, such as an agent that plans, calls your systems and checks its own results, is where the middle tier fits. The top tier is hard to justify for anything a customer triggers thousands of times a day, and easier to justify for low-volume, high-stakes work such as a difficult code migration or a long analytical report.
Many products should use more than one model: a cheap one for routing and simple answers, and a stronger one only when the request needs it. If you are still deciding whether your use case needs an agent at all, our comparison of AI agents, chatbots and workflow automation is a better starting point than a model price list.
There is also no obligation to move today. A feature that works well on an earlier model keeps working until OpenAI announces a retirement date for that model, and those dates are published in advance.
What to do next
Start by listing every place your product calls an OpenAI model, which model it uses, and roughly how many requests it handles a month. Then pick the one or two features where quality complaints or cost are highest and test those first. Use real inputs from your own logs, compare answers side by side, and measure cost per request including long-context and cached-input effects.
Check the non-functional points before switching production traffic: rate limits for your account tier, data residency requirements, and whether your prompts rely on settings the new models do not accept. Finally, set a hard spend limit for the test project. OpenAI added organization and project spend limits on July 22, 2026, according to the changelog, and a testing phase is exactly when a runaway loop gets expensive.
Conclusion
Between September 3 and September 29, 2026, OpenAI released GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and GPT-6.1 Sol in its API, with listed prices that range from USD 0.10 to USD 10 per million input tokens. The models share a very large context window, charge more for very long requests, and come with behavior and configuration changes that need testing before a switch.
If you want a second opinion on which model fits a planned feature, or an estimate for upgrading an existing integration, you can request a quote from Entrant Technologies.