← Back to whitepapers

White paper · August 2026

White paper — Usage limits: producing with AI without buying time

Usage limits: producing with AI without buying time. 43 pages on what usage caps change in the organisation of work — with no conflict of interest: Clarendis resells no subscription, no AI platform, and takes no referral commission.

  • The typology of limits: five families of caps, and how to read your own terms of use
  • The real cost of a lockout: interrupted task, lost context, and the wait-or-pay trade-off made in the cold rather than under pressure
  • Producing lean: seven practices that cut consumption without cutting output
  • The decision grid: when to buy credits, when to wait, when to switch tools, when to bring it in-house
Cover of the white paper
White paper · August 2026
White paper — Usage limits: producing with AI without buying time

Read the opening pages

The 43 pages of this white paper. The first 12 are readable in full; the rest is sent by email.

Page 1 of the white paper “Usage limits: producing with AI without buying time”
Page 1

White paper · August 2026

A message from your AI agent · 3 pm

"Limit reached. Come back in four hours."

Producing with AI without buying time. What usage limits change about how work is organized, and how to work around them without overpaying.

"We resell no subscription, no platform. We are users, like you, and we learned everything that follows at our own expense. Nobody documents this subject: vendors have no interest in it, consultancies sell adoption."

5 families

7 practices

1 grid

of limits: per message, per rolling window, per credit, per seat, per organization, and what each imposes on a team

that cut consumption without cutting output, from task slicing to model choice

of decision: when to buy, when to wait, when to switch tools, when to bring it in-house

For leaders, team managers and AI leads, with a typology of limits, a 90-day plan, the committee's five numbers and a self-assessment. No prices, no rankings: the tools (ChatGPT, Claude, Gemini, Copilot, Le Chat...) are named as landmarks, and the mechanisms will still hold when their offers have changed.

Page 2 of the white paper “Usage limits: producing with AI without buying time”
Page 2

Contents

Foreword 03 Eight findings your invoice will never show you 04 Where to start, depending on your situation 06 The subject in plain language, for everyone 07 1 Why the ceilings exist What a token really is · Why the cost is not linear · The long conversation versus ten short ones 09 2 The typology of limits Five families of ceilings · What each model implies for a team · Reading your own terms of use 13 3 The real cost of a lockout The interrupted task and the lost context · The restart · The wait-or-pay trade-off, made cold rather than hot 17 4 Producing lean: seven practices Slice the tasks · Manage the context · Pick the model for the job · What you delegate and what you keep 21

5 Organizing the team Pool or individualize · Who arbitrates when the ceiling falls · Sharing a limited resource without creating a black market 27 6 The decision grid When to buy credits · When to wait · When to switch tools · When to bring it in-house · API access, the other door 30 7 The blind spots The oversized subscription · Dormant seats · Clandestine AI on expense reports · Single-vendor dependency 34 8 The 90-day plan Measure, tune, anchor · The D1-D90 timeline · The five-number dashboard 38 Moving to execution 41 Appendix · The twenty-question self-assessment 42

The stance

The usage ceiling is neither a scandal nor a fate: it is a production constraint, like budget or deadline. A team that understands it produces more on the same subscription; a team that endures it buys time at a premium, every month, without ever really deciding to.

Page 3 of the white paper “Usage limits: producing with AI without buying time”
Page 3

Foreword

Three o'clock, a Tuesday. The consultant has spent the morning getting the assistant to produce the deliverable due tomorrow; the summary remains. The message drops: limit reached, come back in four hours. She has three options, all bad: wait and deliver late, start over on another account and re-explain the whole context, or pull out the credit card for top-up credits whose fair price and duration she does not know. She will pay. Like last week. Multiplied across every team working seriously with generative AI, this Tuesday at 3 pm has become a cost line, a source of delay and a source of friction, that nobody steers because nobody names it.

This subject has no literature. Vendors document their limits minimally, change them without notice, and have no interest in teaching you to consume less. Consultancies sell adoption: always more usage, never the art of doing as much with less. That leaves the forums, where everyone guesses the rules of the game by suffering them. This white paper fills that gap: it explains why the ceilings exist, how they work across vendors, what a lockout really costs, and how to organize a team to produce as much while consuming less.

Our position, so you know where we speak from: we resell no subscription, no AI platform, and take no referral commission. We are users, like you, intensively, every day, and this document says what we learned at our own expense: the credits bought in a panic at hours when nobody calculates, the river-conversations that devour entire quotas, the subscriptions stacked up to dodge a wall that a better practice would have removed.

A writing precaution. You will find here neither prices, nor rankings, nor offer comparisons: all of that would expire before autumn. The tools are named, ChatGPT, Claude, Gemini, Copilot, Le Chat, because that is the reality of your teams, but always as landmarks, never as recommendations. We describe the mechanisms, tokens, windows, queues, credits, because the mechanisms will outlive the price lists. Page 6 orients you by situation, and pages 7-8 translate all the vocabulary into plain language, so that technical and non-technical readers read the same book; in a hurry, read the eight findings (p. 4-5), the decision grid (ch. 6) and the 90-day plan (ch. 8).

Enjoy the read, and may your next Tuesday at 3 pm go differently.

Cédric Guittard

Anna Hoang

cedric.guittard@clarendis.com

anna.hoang@clarendis.com

for the Clarendis team · August 2026

Page 4 of the white paper “Usage limits: producing with AI without buying time”
Page 4

Eight findings your invoice will never show you

The document in eight statements. Each is developed and equipped in the chapter indicated.

01 The ceiling is not a punishment: it is the real price surfacing.

02 The long conversation is waste item number one.

Every answer costs compute, and the flat-rate subscription pools very unequal usages. The limit is where the economics of the plan stop subsidizing you. (ch. 1)

At every turn, the whole history is reread: the hundredth message costs a multiple of the first. Ten short, targeted conversations produce more than the single thread dragging on since Monday. (ch. 1, 4)

03 Two identical subscriptions do not grant the same right of use.

04 The lockout costs more than the waiting time.

Per message, per rolling window, per credit, per seat or per organization: five families of limits, five ways of breaking down. You do not organize the same way for each. (ch. 2)

The interrupted task loses its context, the restart is paid in re-explanation, and the purchase decided in the heat of the moment is always made at the worst price. The real cost counts in completed tasks, not hours. (ch. 3)

05 Leanness is practiced: it cannot be decreed.

06 A limited resource without a sharing rule creates a black market.

Slice the tasks, start fresh when the thread grows heavy, reserve the big model for what deserves it: seven practices that cut consumption by a third without touching output. (ch. 4)

Lent accounts, quotas siphoned by two heavy users, emergencies that jump every queue: without a named arbitrator and a published rule, the limit gets shared by nerve. (ch. 5)

Page 5 of the white paper “Usage limits: producing with AI without buying time”
Page 5

07 The emergency purchase is a confession of disorganization, not a solution.

08 What is not measured gets overpaid.

Five numbers suffice: consumption per user, lockout rate, cost per completed task, share of emergency credits, split between models. Without them, every subscription renewal is a bet. (ch. 8)

The credits bought at 3 pm to finish by 5 pm are the most expensive of the month, and they return every month as long as the cause (a badly sliced task, a badly shared quota, a badly chosen model) goes untreated. (ch. 3, 6)

What this document deliberately leaves out

No prices: the lists change several times a year, and a dated figure would do more harm than good. No vendor recommendations: the tools are named as landmarks (ChatGPT, Claude, Gemini, Copilot, Le Chat), but the offers converge, and your choice depends on your estate, not on a ranking. No model comparison: it would expire before printing. What remains is what lasts: the mechanisms, the practices, the team rules and the numbers to track, whatever the logo on the invoice.

The question that sums up the book:how many completed tasks does your team get out of a month of subscription, and what stops it from getting a third more without spending one euro more? If nobody can answer, you are not steering your usage: you are observing it on the invoice.

Page 6 of the white paper “Usage limits: producing with AI without buying time”
Page 6

Where to start, depending on your situation

Four situations cover most readers. Find yours: it gives the reading order and the move that can follow this very week.

"We already pay, and we still hit the wall"

Ch. 1 → ch. 4 → ch. 5. Understand what consumes, install the seven practices, then tune the team's sharing. In most cases the wall recedes before any additional purchase.

The most common case

This week: ask everyone how old their oldest still-active conversation is. Threads over a week old are your first suspects.

"We are choosing our subscriptions for the year"

Ch. 2 → ch. 6 → ch. 7. The typology of limits is your reading grid for the offers; the decision grid replaces the "let's take the tier above" reflex; the blind spots will keep you from paying for sleeping seats.

Purchase or renewal

This week: pull the list of paid seats and their last real use. The surprise is almost guaranteed.

"I steer AI for the company"

Full read; ch. 5 and ch. 8 first. The sharing rules and the five-number dashboard are your two instruments: the first prevents friction, the second turns the annual renewal into an informed decision.

AI program lead

This week: institute the reflex "who got locked out, on what, for how long" in your team meeting. That is the lockout rate starting to be measured.

"The bill doubled and nobody knows why"

Ch. 7 (the blind spots) → ch. 3 (the cost of a lockout) → ch. 8 (measure). Look first for recurring emergency credits, dormant seats and AI on expense reports: the three explain most doublings, and they correct within a quarter.

The emergency path

Page 7 of the white paper “Usage limits: producing with AI without buying time”
Page 7

The subject in plain language

The ten words of the subject, translated into plain language

This book speaks as much to the person who uses AI every day as to the leader who signs the subscription. These two pages put everyone on the same footing: ten words, one simple image each, no prerequisites. The rest of the book builds on them.

The token. The "pinch" of text the machine counts for billing: three or four letters. Everything you write and everything it answers is weighed in tokens, like vegetables by the gram.

The thread (or conversation). The continuous exchange you carry on in one window. Keep the parcel image: at every message, the whole parcel is reshipped, and it grows at every turn.

The context. Everything the machine must reread to answer you: your previous messages, its replies, the attached documents. It is the weight of the parcel, and it is what you pay for.

The model. The "engine" that produces the answers: GPT at OpenAI (ChatGPT), Claude at Anthropic, Gemini at Google, Le Chat at Mistral. Each tool offers several versions, from the most frugal to the most powerful, like a range of vehicles from one manufacturer.

Fast model, big model. The city car and the truck. The city car covers the vast majority of trips (summarize, rephrase, sort); the truck is reserved for heavy loads. Taking the truck to fetch the bread: the most common waste of all.

The reasoning mode. An option where the machine "thinks" at length before answering. Powerful on hard problems, and billed even when the thinking never appears on screen.

The ceiling (or quota, or limit). The volume of use your subscription allows before stopping you. The message "limit reached, come back later" is the ceiling falling.

The rolling window. A ceiling computed over the last few hours, continuously, like a tide: it does not reset at midnight, it recedes little by little. That is why you can be locked out at 3 pm for a morning's excess.

The credit. The top-up bought on top of the subscription, like an exceeded phone plan. Useful when planned, ruinous when bought in a panic one hour before the deadline.

The seat. One person's subscription inside a team contract. A sleeping seat (nobody uses it) costs the same as a producing one: hence the chapter 7 inventory.

If you keep only one image: a conversation with an AI is a parcel reshipped in full at every message, growing at every turn. Almost everything this book recommends, closing threads, lightening attachments, slicing tasks, comes down to keeping the parcel light.

Page 8 of the white paper “Usage limits: producing with AI without buying time”
Page 8

The subject in plain language

What happens when you press Enter

1

The whole thread ships, not just your message

The machine has no memory between two messages. For it to "remember", the tool resends it the entire conversation, your questions, its answers, the attached documents, every time. Your two-line question therefore travels with everything that precedes it.

2

Rare and expensive machines do the work

Every answer mobilizes, for a few seconds, specialized computers the whole world is fighting over. That is power, hardware and queuing: which is why nothing is truly free or unlimited, whatever the advertising says.

3

Your meter advances, in silence

The vendor deducts what your message cost, in tokens, from your right of use. A long thread with a heavy attachment can consume in one afternoon what a week of short threads would not have dented, and nothing on screen tells you.

4

At the end of the meter, the wall

Depending on the subscription: a waiting message, a switch to a slower engine, or an invitation to pay. That is the famous "Tuesday at 3 pm" of the foreword, and this whole book explains how to see it coming, push it back, and make sure it breaks nothing when it arrives.

Why does the plan have a ceiling, then? Because it works like a gym: most members come rarely, a minority comes every day, and the single price holds because the former finance the latter. The ceiling is the barrier that keeps the most assiduous from overflowing the gym. It is neither a breakdown nor a slight: it is the business model surfacing , and a team that knows it gets organized, where a team that ignores it gets irritated and overpays.

And now, two possible readings: chapter 1 retraces these mechanisms with one more notch of precision, for those who want to understand finely. Those who prefer to go straight to the moves can jump to chapter 4 (the seven practices) and come back: the book is built for that.

Page 9 of the white paper “Usage limits: producing with AI without buying time”
Page 9

Chapter 1

Why the ceilings exist

1

You only negotiate well with what you understand. This chapter explains what a token is, why the cost of a conversation is not linear, and why the truly unlimited plan will never exist.

1.1 The token, the unit nobody sees

Everything you write to an assistant, and everything it answers, is cut into tokens: word fragments, three to four characters on average in a European language. A page of text makes a few hundred; a long document, tens of thousands. The token is the industry's real billing unit: "per message" or "flat rate" plans are only commercial dressings laid on top. Every token processed mobilizes compute on rare and expensive machines, which has two consequences every user notices without naming them: the vendor caps what it loses on heavy consumers, and a voluminous task exhausts the right of use faster than a short one, even within a single "message".

Second notion, less known and heavier in consequences: the model remembers nothing. At every turn of the conversation, the entire exchange, your messages, its replies, the attached documents, is sent back to it so it can "remember". What you experience as memory is a full rereading, billed every time. This is the key to almost everything that follows: the conversation is not a channel, it is a parcel reshipped in full at every message, growing at every turn.

The mental conversion to install in the team: a "message" is not a unit of cost. The cost is the size of everything the model must reread to answer. Asking a short question in a three-day-old thread costs more than asking a long question in a fresh one.

Page 10 of the white paper “Usage limits: producing with AI without buying time”
Page 10

1

Chapter 1 · Why the ceilings exist

1.2 The thread that grows: why the cost is not linear

Since every turn rereads the whole history, a conversation's cost does not add up: it compounds on a slope. The chart below compares the same day's work run as a single thread or as short threads, with identical output:

Relative consumption of the same work, by slicing (index: turn 1 of the thread = 1)

Single thread · turn 5

≈ 5

Single thread · turn 20

≈ 15-25

Single thread · turn 50, with attachments

≈ 60-100

Same work · 10 short targeted threads

≈ 15-30

Orders of magnitude: the mechanics (full rereading at every turn) are universal; exact figures vary by vendor and their caching optimizations. The slope does not vary.

Three cost behaviours deserve to be known by all. The attachment is permanent: the fifty-page report dropped at turn 2 is reread at every following turn, even when the discussion has moved on. Generation costs more than reading: asking for ten variants of a text means paying for ten texts; asking for the best approach then a single draft means paying for two. Reasoning gets billed: the "deep thinking" modes consume invisible tokens, those of internal reasoning, which can exceed the displayed answer. Useful for hard problems, ruinous for rewording an email.

The reflex that changes everything: treat the conversation as a worksite you open and close, not a companion kept open all week. One thread = one task = one closure. That is practice #1 of chapter 4, and by itself it pushes back most mid-afternoon walls.

© Clarendis · August 2026 edition 10

Page 11 of the white paper “Usage limits: producing with AI without buying time”
Page 11

1

Chapter 1 · Why the ceilings exist

1.3 The economics of the plan: why unlimited will not exist

A fixed-price subscription sitting on a variable cost only holds through pooling: most subscribers consume little, a minority consumes enormously, and the plan lives off that gap, exactly like a gym. The usage ceiling is the tool that keeps the minority from sinking the model: it does not target your comfort, it protects the vendor's margin on its most profitable subscribers, that is, the least active ones. Understanding this dissolves two stubborn illusions: any advertised "unlimited" will eventually grow a "fair use policy", and the tier above does not remove the wall, it moves it back, for a user who will meanwhile have learned to consume more.

Add a constraint vendors never detail: their own scarcity. Available compute does not keep up with demand, and peak hours are managed like a power grid, by shedding load. Hence the behaviours observed but never announced: limits that seem to vary with the hour or the load, queues during office hours, reset windows that change without notice. It is not arbitrariness aimed at you: it is the scarcity management of a saturated infrastructure, of which you are one consumer among millions.

Practical consequence for an organization: the contract says what you pay, never exactly what you get. The real right of use depends on the hour, the vendor's load and the length of your threads. Disconcerting for a buyer used to classic licences, and precisely why internal observation (ch. 8) beats the brochure: your lockout history is the only reliable documentation of your real contract.

The thirty-second test: ask three heavy users at what time their limit usually falls, and when it resets. If they answer precisely, you have a team that learned the rules by suffering them: this book will save it time. If they do not know, chapter 8 starts there.

© Clarendis · August 2026 edition 11

Page 12 of the white paper “Usage limits: producing with AI without buying time”
Page 12

1

Chapter 1 · Why the ceilings exist

1.4 Seven received ideas, corrected in one sentence

These seven sentences circulate in every team. Correcting them upfront saves months of bad habits:

"A short message consumes little."

False in a long thread: the cost depends on everything that precedes, not on the question asked.

"Better keep everything in one thread, the AI has the context."

The useful context fits in ten summary lines; the whole thread, meanwhile, gets re-billed at every turn.

"The big model is always better."

For rephrasing, sorting, summarizing: the fast model does just as well, several times cheaper (ch. 4).

"The tier above will fix the problem."

It moves the wall back without changing the practices that lead to it; six months later, same meeting.

"The limits are written in the contract."

The most binding ones (peak hours, load shedding, moving windows) are documented nowhere.

"Buying credits means losing."

Buying cold, planned, can be the right call (ch. 6); it is the hot purchase, under pressure, that overpays.

"It's a tool problem, not an organization problem."

Two teams on the same subscription produce one to two times as much: the difference lies in the practices (ch. 4) and the sharing (ch. 5).

Next: understanding the mechanism is not enough; you still need to know which exact limit you are under. Chapter 2 draws the typology: five families of ceilings, and for each, what it imposes on how a team organizes.

© Clarendis · August 2026 edition 12

The remaining 31 pages are in the complete document.

Download the white paper
Page 13, blurred — available in the full document
Page 13 · in the full document
Page 14, blurred — available in the full document
Page 14 · in the full document
Page 15, blurred — available in the full document
Page 15 · in the full document
Page 16, blurred — available in the full document
Page 16 · in the full document
Page 17, blurred — available in the full document
Page 17 · in the full document
Page 18, blurred — available in the full document
Page 18 · in the full document
Page 19, blurred — available in the full document
Page 19 · in the full document
Page 20, blurred — available in the full document
Page 20 · in the full document
Page 21, blurred — available in the full document
Page 21 · in the full document
Page 22, blurred — available in the full document
Page 22 · in the full document
Page 23, blurred — available in the full document
Page 23 · in the full document
Page 24, blurred — available in the full document
Page 24 · in the full document
Page 25, blurred — available in the full document
Page 25 · in the full document
Page 26, blurred — available in the full document
Page 26 · in the full document
Page 27, blurred — available in the full document
Page 27 · in the full document
Page 28, blurred — available in the full document
Page 28 · in the full document
Page 29, blurred — available in the full document
Page 29 · in the full document
Page 30, blurred — available in the full document
Page 30 · in the full document
Page 31, blurred — available in the full document
Page 31 · in the full document
Page 32, blurred — available in the full document
Page 32 · in the full document
Page 33, blurred — available in the full document
Page 33 · in the full document
Page 34, blurred — available in the full document
Page 34 · in the full document
Page 35, blurred — available in the full document
Page 35 · in the full document
Page 36, blurred — available in the full document
Page 36 · in the full document
Page 37, blurred — available in the full document
Page 37 · in the full document
Page 38, blurred — available in the full document
Page 38 · in the full document
Page 39, blurred — available in the full document
Page 39 · in the full document
Page 40, blurred — available in the full document
Page 40 · in the full document
Page 41, blurred — available in the full document
Page 41 · in the full document
Page 42, blurred — available in the full document
Page 42 · in the full document
Page 43, blurred — available in the full document
Page 43 · in the full document
Share this white paper