Skip to content

Token Anxiety

Approaching daily Pro limit · resets in 6h

NOTE 📜 This is a post about AI. All views are my own and do not represent my employer. Please review my Disclosures.
range anxiety /reɪndʒ æŋˈzaɪ.ə.ti/ noun The fear that an electric vehicle will run out of charge before reaching a destination or charging station.

There's a similar phenomenon in the AI community: token anxiety.

token anxiety /ˈtoʊ.kən æŋˈzaɪ.ə.ti/ noun The fear that an LLM will exhaust its context or its credits before arriving at a solution.

You know the moment. You're chipping away at a hard problem, on the best model available, on the highest effort setting. You're almost there, and then the banner drops: Approaching daily Pro limit · resets in 6h. Will you make it to the terminating token?

That's one flavor; here's another. You hear AI is the future, you don't want to fall behind, so you buy one of the big-lab plans. You polish a shelved project here, spin up a new project there. A month in, you've used 12% of your quota, and you can't decide whether to feel relieved or guilty.

Or here's a third. You work in tech, your employer encourages AI use, and your usage caps may as well be unlimited. You watch the cool projects your colleagues are shipping and want one of your own. You know AI fluency is a big part of performance now, so you start hunting for ways to spend all those tokens.

Three scenarios, two failure modes. An empty tank, and a full one; not having enough tokens, and having too many.

Empty Tank

When tokens are scarce, the symptoms are both predictable and compounding.

Gemini CLI showing Pro and Flash both at 100% used, hours until reset

Rationing. You start asking the model less than you'd otherwise ask. The cap forces a triage that the work doesn't actually need. Each prompt becomes a small negotiation with yourself before it ever reaches the model.

Attachment avoidance. The first thing you cut is context. You stop pasting the file and describe it in prose. You hint at the error instead of showing it. Each attachment costs tokens, so you give the model less to ground on, and it generalizes from less. The output drifts toward generic, and you don't always notice why.

Model degradation. When the output disappoints, you assume the model was overkill and step down. Pro to Flash, Flash to Flash-Lite. Each step is a small concession: this question doesn't really need the smart model. Sometimes you're right. The trick is that you stop noticing when you're wrong.

Premature compression. The same anxiety that pushed the downgrade pushes you to trim the conversation before it's full. You /compress early to stretch the session, or auto-compression catches you mid-flow. The plan, the failed attempt, the thread you were following: gone. You restart on a thinner version of the problem and the model takes a wrong turn it had already corrected once.

Session splitting. Eventually you skip compression and start over. Work that should live in one continuous thread gets chopped into three. You spend tokens re-explaining where you were, paid in setup instead of progress.

Provider hopping. When Gemini finally caps out, you bounce to ChatGPT. ChatGPT is congested, you bounce to Mistral. Each switch loses the context you'd just rebuilt. You're chasing free tokens and paying in continuity.

Meter watching. By now /stats is a tic. You glance at the percentage between every prompt and weight each question by what it might cost. That's the wrong frame. The question is whether the answer is worth having, not whether you can afford to ask.

Settling. When every prompt feels taxed, you stop iterating. You accept the first draft because you can't afford another round. You stop asking for refactors. You stop asking for alternatives. The output is okay. Okay becomes the ceiling.

Burn out. Living inside these constraints is exhausting in a quiet, attritional way. You're working with a model that should make things easier, and instead you're rationing your access to it. Eventually you decide it isn't worth the friction, and you reach for the tools you trusted before.

Full Tank

You'd think the cure was more tokens. It isn't.

Claude Code session usage at 12%, weekly at 7% — plenty of room left

When tokens are plentiful, every incentive points toward burning them. You think you're falling behind. AI is showing up in your performance review. The dollars you spent on a plan demand to be amortized. So you use, and use, and use.

Trivial offloading. You start asking the model things you would have done in two seconds yourself. Renaming a variable. Looking up a flag. Reformatting a paragraph. The model's latency is higher than yours, but the tokens are "free," so you keep doing it. The muscle for the small things atrophies.

Drift. The bigger work catches the same habit. You let contexts grow long with stale code, dead chats, and abandoned plans. You retry from scratch instead of iterating. You let the model wander without guidance, because it has plenty of room to figure it out. The output gets worse, not better, and you respond by spending more tokens on it.

Project sprawl. With a great number of tokens comes a great number of side-projects. You start three demos in a week. Two are vibes, one has a real idea, and none of them ship. By Friday there's a fourth. The repos outpace the finished work. I'm still figuring out how to keep the portfolio from turning into a graveyard.

Burn out. This is the one I want to be honest about, because it's the reason I'm writing this.

When you're paying for the tooling, you want your money's worth. $20, $100, $200 a month is real money. So you use AI every chance you get. In the checkout line? Check out the long-running operation on your phone. Reading before bed? Perfect time to kick one off. Just woke up? Perfect time to check on it. The financial cost is fixed; the effort cost feels free.

It isn't. You're babysitting agents, holding the architecture in your head, deciding what they should work on next. You never stop thinking about the work. The model gets to forget. You don't.

We're not AI agents; we're human. We need rest, whether we choose it or not. The pause is where the work consolidates, where the bad ideas drop out and the good ones surface, where you remember what you were trying to build in the first place. None of that happens while you're checking on a build from the grocery store.

My current practice is to downgrade my plan once every few months for a month at a time. The lower cap forces me into intentional use, and the spare hours go into the things that aren't coding. It's the only thing that's worked; I'm curious if you've found others.

The Middle Lane

Range anxiety in EVs didn't get solved by 1,000-mile batteries. It got solved by chargers along the route, by trip planning, by drivers learning their cars. The fear faded as the fit got better.

OpenAI Codex showing 5h limit at 51% left, weekly at 61% left — the middle of the tank

Token anxiety won't be solved by infinite limits either. The cure isn't a bigger battery; it's knowing the route. Decide what the work is worth before you ask. Spend where the answer earns it. Hand some of the small tasks back to yourself, so the big ones get the version of you that actually shows up.

A low tank exhausts you quickly. A full tank exhausts you slowly. Only the middle is sustainable.

Comments

Related

Recent