Sources: 

Google has announced adjustments to its Gemini AI usage limits in response to client feedback about quickly hitting quotas and soaring costs. As businesses adapt to new AI capabilities, CEO
Sundar Pichai noted,
"Companies are already blowing through their annual token budgets and it's only May."The tech giant's new
compute-used system introduces a refresh cycle every five hours until the weekly limit is reached, reflecting the complexity of user prompts and tool utilization.
In a bid to alleviate cost pressures, Google highlighted that
heavy tasks like
Deep Research require more resources and announced plans to deliver
better usage analytics to help users optimize their consumption.
Additionally, Pichai remarked that a strategic shift to using a blend of Flash and other advanced models could yield substantial savings.
Data indicated that customers could save over
$1 billion annually should they transition the majority of their workloads to a combination of Gemini 3.5 Flash and other models.
In a notable shift, Google has made 3.1 Flash-Lite prompts
free, stating that they will no longer count against users' quotas.
Furthermore, the company rectified a glitch that previously drained quotas for some users over minimal video usage, emphasizing its commitment to customer satisfaction and efficient resource management.
As AI complexity increases and processes lengthen,
Dan Morgan, analyst at Synovus Trust, noted,
"long-running processes have become the norm."Google, benefiting from its
TPU chips, boasts lower AI compute costs—up to
75% less than competitors—ensuring its alignment with market demands while addressing client concerns.
Sources: 
Google is adjusting usage limits for its Gemini AI following complaints from businesses about exceeding budgets. CEO Sundar Pichai highlighted that some companies are rapidly depleting their annual token budgets and emphasized potential savings with a more strategic usage of AI models.