• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - GPT-5.6 Up to 80% Price Cut! Not Only Are APIs Getting Cheaper – ChatGPT and Codex Users Benefit Too

GPT-5.6 Up to 80% Price Cut! Not Only Are APIs Getting Cheaper – ChatGPT and Codex Users Benefit Too

Rocky by Rocky
July 31, 2026 - Updated on August 5, 2026
in AI Trends and Related News

It’s truly unexpected that OpenAI announced a price reduction less than a month after GPT-5.6 officially launched. Earlier today, OpenAI announced adjustments to the pricing and processing options for the GPT-5.6 series, with Luna seeing the steepest reduction at 80%, Terra dropping by 20%, and the flagship model Sol gaining a new Fast mode that delivers up to 2.5x speed improvements. Whether you’re using the API directly or executing tasks through Codex or ChatGPT Work, this update benefits you across the board—accomplishing more work while spending less budget or usage quota.

OpenAI Slashes GPT-5.6 API Pricing: Luna Down 80%, Sol Adds 2.5x High-Speed Mode

Starting July 30, OpenAI officially reduced the API pricing for GPT-5.6 Luna and Terra. Luna is the fastest and lowest-priced model in the series, primarily targeting high-frequency, high-volume processing tasks; Terra strikes a balance between performance and cost, suitable for general daily tasks.

Both models are priced per million tokens. Luna’s input price dropped from $1 to $0.2, and its output price fell from $6 to $1.2—a reduction of 80%. Terra’s input price was reduced from $2.5 to $2, with its output price dropping from $15 to $12, a 20% decrease.

model Positioning Original input price Enter new price Original export price New Output Price decrease
GPT-5.6 Luna high-frequency, high-volume processing $1 $0.20 6 dollars $1.2 80%
GPT-5.6 Terra Balance cost and performance $2.5 2 dollars $15 $12 20%

Luna’s latest adjustment is particularly noticeable. Assuming a job requires processing 100 million input tokens and 10 million output tokens, the original price would be $160, while the new price is only $32.

Many might wonder if only developers using the API benefit from this. To address that, OpenAI states that the subscription prices and quota budgets for ChatGPT and Codex remain unchanged. However, when using Terra or Luna in Codex or ChatGPT Work, the system deducts fewer credits. In other words, the monthly fee and the listed credit amount stay the same, but the same credits can accomplish more work.

We are committed to pushing the model frontier across cost efficiency, capability, and speed.

Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.

Luna and Terra’s lower prices are… pic.twitter.com/rFhK7XKedp

— OpenAI (@OpenAI) July 30, 2026

As for the most powerful GPT-5.6 Sol, while there was no reduction in standard API pricing this time—with input tokens still at $5 per million and output tokens at $30—OpenAI offers a brand new Fast mode that allows developers with real-time processing needs to trade higher rates for speed:

According to OpenAI, GPT-5.6 Sol in Fast mode can be up to 2.5x faster than standard processing, with unchanged intelligence and a rate of 2x standard mode. This translates to $10 per million input tokens and $60 per million output tokens. Therefore, it’s not necessarily ideal for background tasks that simply prioritize low cost, but rather more suitable for interactive coding, responsive agent services, or enterprise applications where latency directly impacts operations.

Fast mode will also replace the original Priority Processing, though developers don’t need to rush to modify their existing code for now. API requests previously marked as `priority` will automatically switch to Fast mode, aligning with the `/fast` mode in Codex. In other words, existing integration methods can largely continue to be used, with the main differences reflected in processing speed and billing.

In terms of inference systems, OpenAI has also improved request routing, scheduling, load balancing, caching, GPU kernel execution, and model program ordering. Even if the model itself is efficient, if requests are unevenly distributed, hardware is waiting for data, or memory is constantly being moved, GPUs may still sit idle. Therefore, OpenAI distributes global requests based on geographic location, available capacity, and accelerator type. Once requests enter a cluster, they are then routed to appropriate model instances based on current load, context length, and cache status.

Another key technology is speculative decoding. It works by having a smaller “draft model” predict multiple tokens that might come next, which the main model then verifies in parallel. If the predictions are accepted, the main model can generate multiple output tokens in a single computation, reducing the expensive operations required for generating content step by step.

Interestingly, GPT-5.6 Sol also participated in this optimization effort. OpenAI says Sol uses Codex to analyze production traffic, identify load imbalances that weren’t previously noticed, and can test new routing strategies. It also rewrote and optimized GPU kernels used in production on its own, and combined with improvements to other kernels, reduced end-to-end model serving costs by 20%.

GPT-5.6 Sol also designed and ran hundreds of experiments on its own draft model, testing different scales, architectures, and functionalities, even handling the launch and monitoring of the training pipeline, intervening when hardware failures or training instability occurred. These improvements boosted token generation efficiency by over 15%.

Source: KOCPC Chinese

Tags: aiArtificial IntelligenceChatGPTGPT-5.6OPENAI

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology