How to Reduce Claude Code Token Usage (and Stop Hitting Limits)
Claude Code hitting limits? Concrete ways to reduce token usage: context hygiene, a lean CLAUDE.md, compaction, model and effort choice, subagents and caching.
Legacies is a software and web studio from Romania, founded by Horia Stan and Alexandru Talnaci, and Claude Code is open on our machines most of the day. When we hit a limit early in the week, the cause is almost always waste, not work. Here is how we cut it.
The core idea: Claude Code resends your whole conversation, plus every tool result, with each request. Anything that keeps that context small, keeps it cached, or sends routine work to a cheaper model reduces token usage. The biggest wins are clearing between tasks, a short CLAUDE.md, the right model and effort level, and pushing noisy work into subagents or hooks.
Quick answer
- Clear often. Run
/clearbetween unrelated tasks. It costs nothing, while stale context costs tokens on every message. - Keep CLAUDE.md under 200 lines. Move workflow-specific instructions into skills or path-scoped rules that load only when needed.
- Match model and effort to the task. Sonnet for routine work, Opus for hard problems, Haiku for simple subagents, and lower effort for easy edits.
- Keep noise out of the main context. Use subagents for test runs and log digging, hooks to filter output, and CLI tools instead of heavy MCP servers.
Why Claude Code burns through tokens
Before the techniques, the mechanics. The Claude Code cost docs explain it plainly: Claude Code sends your full conversation with every request, and each time Claude uses a tool it sends another request carrying that batch of tool results.
So a session is not a series of small messages. It is a growing pile that gets re-read many times per task. Prompt caching makes re-reading cheaper, but not free, and caches expire. On a subscription the cache lives for one hour. On usage credits, an API key or a cloud provider, it is five minutes by default.
That one fact explains almost every technique below. If you want the plan-level picture, our post on Claude usage limits explained covers the five-hour and weekly windows.
Measure first
You cannot fix what you do not see. Three commands tell you where tokens go:
/usageshows your plan bars and, on paid plans, what consumed them: skills, subagents, plugins and individual MCP servers as a share of the total. It flags behaviors like long context or cache misses when one passes 10% of recent usage./contextshows what is taking space in the current context window right now.- A status line can display context usage continuously, so you notice a bloated session before it hurts.
Newer versions also show a prompt cache line in /usage, with the share of input served from cache and the likely cause of the last miss. If that number is low, something in your workflow keeps breaking the cache.
Context management techniques
Clear between tasks
This is the single highest-impact habit. When you switch from a bug in the checkout to a copy change on the homepage, run /clear. Everything from the checkout work would otherwise ride along on every request. Use /rename first if you want to come back later with /resume.
The reasoning: a fresh session pays only for your system prompt, CLAUDE.md and the new task. A stale one pays for all of that plus hours of irrelevant history.
Compact with intent
/compact summarizes the conversation to free space. You can tell it what to keep:
/compact Focus on the API changes and the failing tests
You can also add standing instructions to CLAUDE.md:
## Compact instructions
When compacting, keep code changes, test output and open decisions.
One catch the docs point out: compacting reads the conversation it summarizes, so compacting a huge context is itself a large request. If you want continuity, compact. If you want a fresh start, /clear is free.
Know your auto-compact window
Opus 5.5 and Sonnet 5.5 have a 1 million token context window, and by default auto-compaction kicks in near the top of it, per the model configuration docs. A huge window is convenient, but every request in a 600,000 token session is expensive. You can set a lower threshold with /autocompact 300k so Claude summarizes earlier.
Mind the cache lifetime
Your first message after a break longer than the cache lifetime reprocesses the full context. If you step away for two hours, it is often cheaper to /clear and restate the task in three lines than to resume a giant session. On Pro and Max, Claude Code offers to resume a large session from a summary for exactly this reason.
Make CLAUDE.md lean
CLAUDE.md is loaded at the start of every session, so every line is paid for on every request. Anthropic's memory docs recommend targeting under 200 lines per file.
Keep in it only what applies to almost every task: build and test commands, conventions, project layout. Move the rest:
- Workflow instructions like release steps or migration procedures go into skills, which load only when invoked.
- Area-specific rules go into
.claude/rules/with apathsfield, so they load only when Claude touches matching files. - Imports do not save tokens. Files pulled in with
@pathload at launch too. They organize a long file, but do not shrink it.
A path-scoped rule looks like this:
---
paths:
- "app/api/**/*.ts"
---
All API routes validate input with zod and return typed errors.
That rule costs nothing when Claude is editing CSS.
Choose the model and effort deliberately
Pick the cheapest model that does the job
Claude Code defaults to Opus 5.5 on Pro, Max, Team and the API. It is the right tool for architecture and hard debugging, and overkill for renaming a prop. The docs say it directly: Sonnet handles most coding tasks well and costs less. On the API, Opus 5.5 is $4 input and $20 output per million tokens, Sonnet 5.5 is $2 and $10, and Haiku 4.5 is $1 and $5.
Two useful options: /model sonnet for the bulk of the day, and opusplan, which uses Opus in plan mode and switches to Sonnet to execute. For simple subagents, set model: haiku in the subagent configuration.
Lower effort on easy work
Thinking tokens are billed as output tokens. Opus 5.5 and Sonnet 5.5 always use adaptive thinking, so you cannot turn it off, but you can steer it with /effort. Both default to medium. Use low for quick exchanges where you review each result, and save high, xhigh or max for problems with real edge cases.
Keep noisy work out of the main context
Delegate to subagents
Running a test suite, reading documentation or digging through logs produces thousands of lines you do not need to keep. A subagent does that work in its own context and returns only a summary. The catch: subagent requests still count against your usage, so give simple ones a smaller model.
Be careful with agent teams. Anthropic says they use about 7x more tokens than a standard session when teammates run in plan mode.
Filter output with hooks
A hook can preprocess data before Claude sees it. The docs show a PreToolUse hook that rewrites test commands to return only failures, which cuts a long test log down to the lines that matter. The same idea works for build logs and linters.
Prefer CLI tools over MCP where you can
MCP tool definitions are deferred by default now, so only names enter context until a tool is used. Still, CLI tools like gh or aws add no per-tool listing at all. Run /mcp and disable servers you are not using.
Work in smaller, sharper tasks
Most of the savings hide here.
- Be specific"Add input validation to the login function in auth.ts" lets Claude read two files. "Improve the codebase" makes it scan everything.
- Plan first on anything bigPlan mode lets Claude explore and propose an approach before writing code, so you do not pay to build the wrong thing.
- Stop earlyIf Claude heads the wrong way, press Escape. Use /rewind to return to a checkpoint instead of arguing it back on course.
- Give it a way to verifyA failing test or an expected output lets Claude check its own work instead of you spending turns on corrections.
A quick checklist
| Technique | Why it saves tokens | Effort to adopt |
|---|---|---|
/clear between tasks | Stops stale history riding on every request | None |
| CLAUDE.md under 200 lines | It loads on every request | Low |
| Skills and path-scoped rules | Instructions load only when relevant | Medium |
| Sonnet by default, Opus on demand | Lower price per token | None |
Lower /effort on easy tasks | Fewer thinking tokens, billed as output | None |
| Subagents for verbose work | Noise stays out of the main context | Low |
| Hooks to filter output | Fewer tokens per tool result | Medium |
| Specific prompts and plan mode | Fewer file reads and less rework | None |
Where a team fits
If you apply all of this and still burn through Max every week, the question changes. Is the quota going into learning, or into building a product that has to be secure and maintained? The second case is often cheaper to hand off. We cover the usual gaps in taking a vibe-coded app to production, and if you are deciding between tools, see Claude Code vs Codex and Cursor vs Claude Code.
When to do it yourself and when to bring in a team
Do it yourself when you are prototyping, learning or building internal tools. With the habits above, a Pro or Max plan goes a long way, and our post on how AI coding agents change small teams covers the review side.
Bring in a team when the app has real users, payments or personal data, or when your hours are worth more on the business than on prompting. We use the same agents and review every line they write. Our services page has fixed starting prices, from 1,499 lei (about EUR 285) for a landing page to 3,499 lei (about EUR 665) for a web app, and the free website audit checks an existing site in seconds.
Frequently Asked Questions
Why is Claude Code using so many tokens?
Claude Code resends the full conversation and every tool result with each request, so long sessions grow expensive even when your messages are short. Big files, verbose test output, a long CLAUDE.md, Opus as the default model and cache misses after breaks are the usual causes. Run /usage and /context to see which one applies to you.
Does /compact save tokens in Claude Code?
It saves tokens on the requests that follow, because the conversation becomes a shorter summary. The compaction itself is a large request, though, since it reads the whole conversation it summarizes. If you do not need continuity, /clear is free and saves more.
Is Sonnet good enough for Claude Code?
Yes, for most coding work. Anthropic's own docs say Sonnet handles most coding tasks well and costs less than Opus, and recommend reserving Opus for complex architecture or multi-step reasoning. The opusplan setting gives you Opus for planning and Sonnet for execution.
How long should a CLAUDE.md file be?
Anthropic recommends keeping each CLAUDE.md under 200 lines. Longer files cost tokens on every request and reduce how well Claude follows them. Move workflow-specific instructions into skills and area-specific rules into path-scoped files in .claude/rules.
Do subagents use fewer tokens?
Subagents keep verbose work out of your main conversation, which keeps the main context small and cheaper per request. Their own requests still count against your usage, so the real saving comes from giving simple subagents a smaller model like Haiku and only returning a short summary.
How do I see what is using my Claude Code tokens?
Run /usage in Claude Code. On paid plans it shows your session and weekly bars and breaks recent usage down by skills, subagents, plugins and MCP servers, and it flags habits like long context or cache misses. The /context command shows what fills the current context window.