Since Google… Gemini switched to calculating based on compute usage.This quickly drew complaints from many paid users, especially Google AI Pro subscribers, who found their quotas depleting much faster than expected—sometimes hitting the cap after just a few intensive tasks, a noticeable change from before. So for anyone who frequently uses Gemini for coding, long-document analysis, Deep Research, or video generation, the experience must feel pretty frustrating.
Google already made adjustments last week. AntigravityNow finally starting to make changes to Gemini too! Earlier. Google Vice President Josh Woodward announced on social mediaDetails of Gemini usage limits will be changed, including fixes for issues such as incorrect quota deductions for error requests, abnormal video generation consumption, and heavy prompts consuming too much quota.

Google Adjusts Gemini Usage Limits: Failed Requests No Longer Deduct Quota, Flash-Lite Is Free, and Ultra Doubles Video Generation Output
The background of this controversy stems from Google adjusting Gemini Apps usage limits starting May 17, 2026. Gemini will now calculate usage limits based on “compute usage” instead of the previous “number of prompts.” Compute usage is affected by factors such as prompt complexity, features used, and conversation length; the allowance refreshes every 5 hours until the weekly cap is reached.
Although this approach is the same as Claude’s, since Google just changed it, no one knows what the right 5-hour usage limit should be. It was set too low at the start, and even instances where no result was returned are counted, causing many paid users to hit their usage quota after barely using it.
One of the most obvious issues is Omni video generation. Many users have run into it: just making one or two videos fills up the quota. For those who want to test and repeatedly adjust video prompts, it’s basically getting stuck before they even start. Google said this time that the bug related to Omni video generation has been fixed, and it has also increased the quota for heavy users. Google AI Ultra users will immediately see their Omni video generation quota doubled:

Another focus of the correction involves the more complex “Pro model” prompts—tasks like long prompts, large file uploads, and multi-step reasoning already consume more computing power than ordinary chat. However, the previous calculation method could cause a single request to use up too much, so Google will now impose a per-request cap on prompt consumption for these heavy tasks, preventing an especially complex task from eating up too much of the available quota:

More importantly, failed requests will no longer count against your quota. Josh Woodward noted that roughly 1 in 10 requests may fail due to system errors. Previously, if Gemini itself errored after a user submitted a prompt, it could still be counted toward the usage limit. After this fix, no quota will be deducted as long as the request fails:

Additionally, the Flash-Lite model will no longer count toward your quota going forward. Google officially positions Flash-Lite as a lightweight, fast model suited for everyday tasks like summarization and brainstorming. If you’re just organizing text, doing quick Q&A, or rewriting a short passage, you don’t necessarily have to use the Pro model—you can switch to Flash-Lite instead, saving your quota for tasks that genuinely require reasoning or long-context processing:

Deep Research will also get clearer usage breakdowns and notifications. Going forward, Google will let users see their Deep Research usage more clearly, so they can at least tell which type of task is consuming more, rather than just seeing their quota suddenly drop sharply:

There’s also a small but practical tweak: Gemini now remembers the model you’ve selected. So if you frequently use a particular model for writing, research, or organizing information, you won’t have to select it again the next time you open Gemini. However, if you hit a usage limit, the system will still automatically switch to a lighter model so the conversation can continue:

Here is a summary of the Gemini usage limit changes announced this time:
| Adjustment Items | Corrected content |
|---|---|
| failed request | Usage quota will no longer be deducted for requests that fail due to system errors. |
| Flash-Lite | Flash-Lite prompts no longer count toward usage quotas. |
| Advanced Pro Tips | Google will impose a per-use consumption cap on heavy requests such as complex prompts, long conversations, and large files. |
| Omni video generation | The related bug has been fixed, and AI Ultra users will also have their Omni video generation quota doubled. |
| Deep Research | Google will improve usage breakdowns and notifications for Deep Research. |
| Model selection | Gemini remembers the model the user selected, but it may still automatically switch to a lighter-weight model when limits are reached. |
Source: KOCPC Chinese