Quick answer: Starting September 2, 2026, Google is changing Gemini Notebook from simple fixed-looking daily generation limits to a compute-based system. Your usage will depend on prompt complexity, chat length, number of sources, the model, and the feature you use. The allowance refreshes every five hours until you reach a weekly limit, and Gemini Notebook will show remaining usage plus an estimated cost bar before Studio generations.[1][4]
The practical change is simple: a short source-grounded question should consume less of your allowance than a long notebook with many sources or a resource-heavy Video Overview. Instead of guessing how many prompts remain, check the usage meter inside the notebook before starting an expensive artifact.[1][4]
Rollout note: Google says the new limits begin rolling out to consumer accounts on web and mobile on September 2. Your account may not change at the exact same moment as another user’s.[1]
What changes on September 2
Google’s announcement and support page confirm five important changes:[1][4]
- Usage becomes compute-based. There is no single universal “one prompt equals one credit” rule.
- Different tasks can consume different amounts. Prompt complexity, chat length, model choice, source count, and feature choice all matter.
- The short-term allowance refreshes every five hours. A weekly ceiling also applies.
- Gemini Notebook shows usage information in the product. Chat can show your remaining allowance and reset time, while Studio displays an expected-usage bar.
- You can queue some Studio artifacts for later. If you do not have enough allowance, the web version can generate the artifact after capacity becomes available and notify you when it is ready.
Google has not published a reliable conversion such as “one Video Overview equals ten chats.” Do not trust unofficial calculators that claim an exact prompt-to-credit formula unless Google adds one to its documentation.
How to check your remaining usage
On the web or mobile app:
- Open Gemini Notebook.
- Select Settings in the top-right corner.
- Select Usage.
- Review your current allowance and weekly limit.
The chat panel can also display messages such as an approaching-limit warning or the time when access refreshes. In Studio, check the expected-usage bar before generating an Audio Overview, Video Overview, Slide Deck, infographic, report, or another artifact.[4]
This in-product meter is more dependable than counting prompts yourself because two prompts can require very different amounts of compute.
How “Generate later” works
Google is adding a queue for users who reach their current allowance. The feature is currently documented as web-only.[4]
- Open the notebook on the web.
- Go to the Studio panel.
- Choose the artifact you want.
- Select Generate later when the option appears.
- Enable notifications if you want an alert when the output is complete.
Google says a delayed generation can take a couple of hours. The announcement specifically gives Video Overviews and Slide Decks as examples of outputs that can be deferred.[1][4]
Do not use the queue for a presentation or lesson you need in the next few minutes. Generate time-sensitive artifacts earlier, then review them for factual errors, missing citations, layout problems, and unusable audio before sharing.
Which tasks are likely to use more allowance?
Google does not provide an exact formula, but it explicitly names the factors used by the system.[1][4] Based on those documented factors, expect a task to consume more of the available budget when it has one or more of these characteristics:
- a long conversation with substantial prior context;
- many attached sources;
- a complex prompt requiring several steps;
- a more compute-intensive model or feature;
- a Studio artifact such as a Video Overview or Slide Deck;
- repeated revisions of a generated artifact.
This does not mean every long PDF automatically costs the same amount or that every Video Overview has one fixed cost. The product’s expected-usage bar is the live indicator to follow.
Best ways to make the allowance last
1. Start a focused notebook
Use one notebook for one project or closely related set of sources. Mixing unrelated documents increases noise and can force longer prompts to explain what the model should ignore.
2. Select only the sources needed for the question
If your notebook contains many sources, narrow the active set before asking a targeted question. Google specifically says source count factors into usage.[1]
3. Ask for a plan before an expensive artifact
Before generating a slide deck or video, use a concise chat prompt to confirm the audience, structure, key evidence, and required sections. Fixing the outline first can reduce wasteful artifact revisions.
Copy-paste prompt:
Using only the selected sources, propose a concise outline for a [slide deck/video overview/report] for [audience]. List the key claim and supporting source for each section. Do not generate the final artifact yet.
4. Split exploration from production
Use chat to explore the sources, then begin a fresh, well-scoped production prompt once you know the desired output. Very long chat history is one of the factors Google says can increase usage.[1][4]
5. Queue non-urgent outputs
Use Generate later for a Video Overview or Slide Deck that does not need immediate review. Keep urgent client, class, or meeting deliverables out of the delayed queue.
6. Check the meter before revising
A vague instruction such as “make it better” can trigger another expensive generation without fixing the actual problem. Name the required change: shorten section three, remove two slides, correct a date, simplify the vocabulary, or use only specified sources.
Free, Plus, Pro and Ultra limits
Google’s new compute-limit help page describes the relative AI allowances this way: Standard for users without a plan, 2x standard for AI Plus, 4x standard for AI Pro, and either 5x or 20x higher than AI Pro for AI Ultra depending on the subscription.[4]
Google’s public plans page uses different wording for artifact generations: Plus lists 2x standard generations, Pro lists 5x, and Ultra lists up to 50x.[3] These are not safe to treat as one interchangeable formula. One page describes the broader compute-based AI allowance, while the other describes generation access by plan.
The same plans page currently lists source capacity of up to 50 sources per notebook on Standard, 100 on Plus, 300 on Pro, and 600 on the highest Ultra option.[3] Availability of paid plans varies by region.[3]
Google’s upgrade help page still displays per-day feature caps and labels usage limits as subject to change.[2] Treat those figures as a plan reference, not as a durable conversion between prompts and the new compute budget.
Because the September 2 system is compute-based, do not buy a plan based on an assumed fixed number of prompts. Check the live plan screen in your country and compare it with your real workload after the rollout.
Practical workflow for students and researchers
Use this sequence to avoid consuming the larger part of your allowance before the analysis is ready:
- Create a project-specific notebook.
- Add only authoritative and relevant sources.
- Ask for a source inventory and identify missing evidence.
- Request an outline with citations.
- Correct the outline and source selection.
- Generate the final report, deck, infographic, or overview once.
- Verify every important claim against the original source.
- Export or save the finished work before making optional stylistic revisions.
Gemini Notebook can accelerate source review, but it can still misunderstand a document or produce an inaccurate synthesis. Human verification remains necessary for academic citations, legal or medical information, financial decisions, public claims, and quotations.
Frequently asked questions
When do the new Gemini Notebook limits start?
Google says the rollout to consumer accounts on web and mobile starts on September 2, 2026.[1][4]
When does the allowance reset?
The short-term allowance refreshes every five hours until the account reaches its weekly limit.[1][4] Check Settings → Usage for the live reset time shown for your account.
Is there still a fixed daily prompt limit?
The new system is compute-based, so Google does not describe it as one fixed prompt count that applies equally to every task. Complexity, chat length, source count, model, and feature choice can change usage.[1][4]
Can I generate a Video Overview after hitting the limit?
On the web, you may be able to select Generate later. Google says delayed outputs can take a couple of hours and can trigger a notification when complete.[1][4]
Does Generate later work on mobile?
Google’s support page currently labels the delayed-generation feature as web-only.[4]
Where can I see how much usage remains?
Open Settings → Usage. Chat can also show remaining usage and the next refresh time, while Studio displays an expected-usage bar for a generation.[4]
Are paid-plan limits the same in every country?
No. Google says AI Plus, Pro, and Ultra availability is region-dependent.[3] Check the plan screen attached to your own Google account for current local availability and terms.
Sources
[1] https://blog.google/innovation-and-ai/products/gemini-notebook/new-flexible-usage-limits — Google: More compute flexibility in Gemini Notebook
[2] https://support.google.com/gemininotebook/answer/16213268?hl=en — Gemini Notebook Help: Usage limits
[3] https://notebooklm.google/plans — Gemini Notebook plans
[4] https://support.google.com/gemininotebook?p=usagelimits — Gemini Notebook Help: Manage usage limits