I burned through my Claude limits in two days the first time I used Claude Code seriously. No warning, no gradual slowdown, just a hard stop.
That sent me down a research rabbit hole that changed how I use Claude Code and Cowork for good.
So this guide does two things. First, I explain how Claude usage limits actually work, using only what Anthropic itself publishes.
Then I share the 20 habits that cut my Claude token usage by over 60% without hurting output quality.
Key Takeaways
How Claude Usage Limits Work
Here's the short version. Every fact in this section comes from Anthropic's own help center, docs or pricing page.
Where Anthropic doesn't publish a detail, I say so instead of guessing.
Claude and Claude Code share the same limits
This is the one that catches most people. On Pro and Max, all activity in Claude and Claude Code counts against the same usage limits.
Anthropic's usage limits article says the same about claude.ai, Claude Code and Claude Desktop. So a long research chat in the morning eats into the budget you have for coding in the afternoon.
The 5-hour session window
According to Anthropic's pricing page, every plan has usage limits that reset on a rolling five-hour session window. That includes Free.
In Settings > Usage you can see how much of your current session you've used and how much time is left in it.
Anthropic doesn't publish exactly what starts the clock or how many messages fit in a session. It says usage depends on things like conversation length, the features you use, the model and the effort level you pick (source).
The weekly cap
Paid plans add a weekly limit on top of the session limit, and it applies across all models.
The weekly limit resets at a fixed day and time assigned to your account. Per the Pro plan article, that time stays the same no matter when you start using Claude or when your subscription began.
Fable models have their own rule. On Max you can spend up to 50% of your weekly limits on Fable. On Pro, Fable sits outside your plan limits and only runs on usage credits (source).
Claude plans and limits compared
| Plan | Price (US) | Usage | Claude Code | Fable 5.1 |
|---|---|---|---|---|
| Free | $0 | Five-hour session limits (Anthropic doesn't publish exact numbers) | Not included | Not included |
| Pro | $20/month, or $17/month billed annually ($200 up front) | Session limit plus weekly limit | Included | Usage credits only |
| Max 5x | $100/month | 5x Pro's per-session usage, plus weekly limit | Included | Up to 50% of weekly limits |
| Max 20x | $200/month | 20x Pro's per-session usage, plus weekly limit | Included | Up to 50% of weekly limits |
Sources: Claude pricing, What is the Pro plan?, What is the Max plan? and Claude Fable models on your plan. Prices exclude tax. Team and Enterprise use seat-based limits, which I don't cover here.
One more thing from the Max plan article: Anthropic says it may also limit usage in other ways, like weekly and monthly caps or model and feature limits, to manage capacity.
What happens when you hit your limit
In Claude Code you'll see a message like "You've hit your session limit" or "You've hit your weekly limit", and it tells you when the window resets (Claude Code docs).
Those two limits cover all models, so switching models with /model won't get you going again. A model-specific message like "You've hit your Opus limit" is different: switching to a model outside that family does keep you working.
Your options at that point:
Claude Code asks for your consent before it switches you to API billing, so you won't get charged by surprise (source).
One trap to watch for: if ANTHROPIC_API_KEY is set in your environment, Claude Code uses that key instead of your subscription, and you pay API rates.
Limit resets: a free refill
Anthropic occasionally gives eligible plans a limit reset. It sets your five-hour session limit or your weekly limits back to full, right away.
You use it from Settings > Usage on the web or Claude Desktop by clicking "Reset for free". It isn't available in Claude Code or on mobile, and once you use it you can't undo it.
Don't sit on it forever, though. An unused reset expires on the day and time listed on the offer, and if you downgrade or cancel before using it, it's gone (source).
How to check your Claude usage
In the Claude app: go to Settings > Usage. On Pro, Max, Team and seat-based Enterprise plans you get progress bars for your five-hour session and weekly limits, plus your next reset time (source).
In Claude Code: run /usage. It shows your plan usage bars, activity stats and a breakdown of what's eating your limits: skills, subagents, plugins and individual MCP servers, each as a share of the total. /cost and /stats are aliases.
Press d or w inside /usage to switch between the last 24 hours and the last 7 days. The breakdown only counts sessions on that machine, so your claude.ai chats and other devices don't show up in it (docs).
I also run /context to see what's sitting in the context window right now. It's the fastest way to spot a bloated setup.
Why Your Claude Limits Drain Faster Than You Expect
Now for the fixes. First, here's the mechanic behind most of the waste: Claude Code sends your full conversation with every request. Every single turn.
Prompt caching re-reads that history at the cheaper cached rate. But a one-line question in a session that has been open all day still draws usage for the whole conversation (docs).
And if you step away longer than the cache lifetime (an hour on a subscription), your next message misses the cache and reprocesses your whole context.
On top of that, Claude Code loads your CLAUDE.md, MCP server info, skills and system context before your first message. On a bloated setup, that adds up before you ask Claude anything.
This isn't Claude being wasteful. It's how large language models work, and once you get it, the 20 fixes below make a lot more sense.
1. Master /clear and /compact

I use these two commands more than almost anything else in Claude Code. They're the fastest way to stop token waste.
/clear: start fresh
/clear wipes the conversation history. After you run it, only your CLAUDE.md and session configuration reload, and the context window is empty.
I use /clear every time I switch to a different task: between projects, between unrelated features, and after any long debugging session.
Not dragging 40 turns of old context into a new task saves a lot. Anthropic's docs also point out that /clear costs nothing, while compacting doesn't.
One tip from the official docs: run /rename before you clear, so you can find the session later with /resume.
/compact: summarize without losing the thread

/compact replaces your conversation history with a compressed summary. You keep a record of what was discussed, but the history gets much smaller.
Use /compact when you're mid-task and the session is getting long, but you still need continuity.
Keep in mind that /compact reads the whole conversation to summarize it, so compacting a huge context is itself a big request. My rule of thumb is to compact around 60% context usage instead of waiting for 90%.
You can also tell Claude what to preserve when it compacts:
/compact Focus on code samples, file paths, and the current task only
Or set it permanently in your CLAUDE.md:
# Compact instructions When compacting, prioritize: current task state, file changes, test results. Drop: exploratory conversations, resolved errors, tangential discussions.
I added this to my CLAUDE.md a while back, and it noticeably improved my compacted summaries.
2. Keep CLAUDE.md Under 200 Lines

This is one of the most specific recommendations in Anthropic's docs, and I see most people ignore it.
Your CLAUDE.md loads into context at the start of every session, before your first message. Every line in it rides along on every turn.
A 600-line CLAUDE.md doesn't just cost more at startup. It adds up across every message in the session.
The right content for CLAUDE.md is essentials only:
The wrong content is anything task-specific, workflow-specific or situational. That belongs in Skills, which I cover next.
My own CLAUDE.md sits at 87 lines.
When I'm tempted to add something, I ask: "Is this needed for every session, or just specific workflows?" If it's for specific workflows, it goes into a skill instead.
3. Move Specialized Instructions Into Skills

Skills are one of the most underused ways to save usage in both Claude Code and Cowork.
Unlike CLAUDE.md, a skill only loads when Claude invokes it. If you're not doing a PR review, your PR review instructions cost zero tokens.
Here's the workflow I follow:
The result: my base session context dropped by about 35% from this move alone.
The output quality is the same, often better, because the full skill instructions come in clean instead of buried in a long CLAUDE.md.
Skills work the same way in Cowork. Each installed skill only adds overhead when it's invoked.
4. Disable MCP Servers You Aren't Using

This one surprised me when I first measured it. Every MCP server you keep connected adds some overhead to your sessions.
Older versions of Claude Code loaded every tool definition up front. One developer measured 14,214 tokens per session from a single server with 20 tools.
Claude Code now defers MCP tool definitions by default, so only tool names and server instructions load until Claude actually uses a tool. That helps a lot, but idle servers still aren't free.
The fix is simple:
Run /mcp in Claude Code to see your configured servers. Disable anything you're not using for the current task, and turn it back on when you need it.
Then check /usage. Its breakdown shows each MCP server's share of your recent usage, so you can see which ones actually cost you.
Anthropic's docs also say to prefer CLI tools over MCP servers when both exist.
Tools like gh, aws, gcloud and sentry-cli are more context-efficient because they don't add any per-tool listing. Claude just runs the command.
The same idea applies to plugins in Cowork. Every connected plugin with MCP tools adds to your session context.
I review my plugin connections monthly and disconnect anything I haven't used in two weeks.
5. Choose the Right Claude Model for Each Task

Not every task needs the biggest model. Anthropic lists the model you pick as one of the things that decides how fast you use your limits (source).
It's a habit I had to build on purpose, because the pull to always use the most capable model is real.
Here's the current lineup as of October 2026, from Anthropic's models overview, and how I use each one:
Claude Haiku 4.5: the fastest model. I use it for simple rewrites, boilerplate, formatting data, quick questions and simple subagent tasks.
Claude Sonnet 5.5: Anthropic calls it the best mix of speed and intelligence, and it launched September 28, 2026. It handles most of my coding, writing, research and debugging.
The Claude Code docs say the same thing: Sonnet handles most coding tasks well and costs less than Opus.
Claude Opus 5.5: built for long-running agentic coding and knowledge work. At launch on September 22, 2026, Anthropic said it performs at the level of Fable 5.1 on most work and costs 40% less to run than Opus 5.
I reach for it on architecture decisions, big multi-file refactors and deep analysis across a large codebase.
Claude Fable 5.1: for the most demanding reasoning and long-horizon agentic work (launch post). On Max it can use up to 50% of your weekly limits, and on Pro it only runs on usage credits. I treat it as a special-occasion model.
API prices give you a feel for how heavy each one is, per million input/output tokens: Haiku 4.5 is $1/$5, Sonnet 5.5 is $2/$10, Opus 5.5 is $4/$20 and Fable 5.1 is $10/$50 (pricing).
Switch models mid-session with /model, or set a default in /config. Switching your session to Opus also applies to subagents that inherit your session's model.
For simple subagent tasks, the official docs recommend setting model: haiku in the subagent configuration.
6. Turn Down Effort and Thinking

Extended thinking is on by default in Claude Code. It's powerful, but thinking tokens bill as output tokens.
The default budget can reach tens of thousands of tokens per request, depending on the model. For simple, well-defined tasks, that's mostly waste.
Here's how I control it:
Lower the effort level for simple tasks:
/effort low
This tells the model to think less without turning thinking off. On the newest models, it's the main lever you have.
Cap the thinking budget on older models:
MAX_THINKING_TOKENS=8000
This only works on models with a fixed thinking budget, like Haiku 4.5. Opus 5.5, Sonnet 5.5 and Fable use adaptive reasoning and ignore it, so use /effort there (docs).
Turn thinking off in /config. Heads up: you can't turn it off on Opus 5.5, Sonnet 5.5 or the Fable models, which always think. On those, /effort is your tool.
Anthropic's guidance is simple: extended thinking helps a lot on complex planning, and for simpler tasks it's cost you don't need. Match the effort to the task.
One real comparison makes this concrete. The same bug-finding task at low effort took 7 turns, 1,028 output tokens, 18 seconds and $0.16.
At extra-high effort it took 9 turns, 1,363 output tokens, 21 seconds and about $0.19. That's about 1.3x the cost for the same bug found, and across repeated runs the gap tends to widen closer to 1.5x.
So I default to low and save high effort for tasks that need it.
7. Use Subagents for Verbose Operations
This is one of the most useful and least understood ways to save tokens.
When Claude processes a 10,000-line log file, runs a full test suite or reads big documentation pages, all that output lands in your main conversation.
It then rides along on every later message in the session, not just the one where it happened.
Subagents fix this. Hand a verbose job to a subagent and the heavy output stays in its own context window.
Only the summary comes back to your main conversation, so your main context stays clean. One catch: the subagent's own requests still count toward your usage, so give it a smaller model when the job is simple.
Where I use subagents:
The official Anthropic docs put it simply:
"Delegate verbose operations to subagents so the verbose output stays in the subagent's context while only a summary returns to your main conversation."
Agent teams take this pattern all the way, but the docs warn they use about 7x more tokens than standard sessions when teammates run in plan mode.
Keep teams small and tasks focused, and shut teammates down when the work is done. Each active teammate keeps using tokens until it exits.
8. Use Hooks to Preprocess Data Before Claude Sees It

Hooks run before and after Claude's tool calls. They can cut down a lot of what Claude actually has to read.
The example I use most: test output filtering.
Without a hook, npm test returns hundreds of lines: passing tests, coverage reports, timestamps, file paths. I only care about failures.
With a PreToolUse hook that filters test commands down to failures, I went from 3,000-token test outputs to 80-token summaries.
Here's the script from the official docs. Save it as ~/.claude/hooks/filter-test-output.sh, make it executable, and register it as a PreToolUse hook on Bash in your settings.json:
#!/bin/bash
input=$(cat)
cmd=$(echo "$input" | jq -r '.tool_input.command')
# If running tests, filter to show only failures
if [[ "$cmd" =~ ^(npm test|pytest|go test) ]]; then
filtered_cmd="$cmd 2>&1 | grep -A 5 -E '(FAIL|ERROR|error:)' | head -100"
echo "$input" | jq --arg filtered "$filtered_cmd" \
'{hookSpecificOutput: {hookEventName: "PreToolUse", permissionDecision: "allow", updatedInput: (.tool_input + {command: $filtered})}}'
else
echo "{}"
fiBeyond test filtering, I use hooks for:
Every hook that trims data before Claude reads it is a direct saving. The filtering runs locally for free, and Claude has less to process.
9. Block Junk Folders With Read Deny Rules
The old advice, including an earlier version of this guide, was to create a .claudeignore file. Don't bother.
Anthropic's permissions docs say a .claudeignore file has no effect, and that you should move its entries into Read deny rules instead.
For big projects this matters. Without it, Claude can spend tokens reading node_modules, build output, generated files and vendor folders while it explores.
Add a deny list to .claude/settings.json in your project. Here's my standard set:
{
"permissions": {
"deny": [
"Read(node_modules/**)",
"Read(dist/**)",
"Read(build/**)",
"Read(.next/**)",
"Read(vendor/**)",
"Read(coverage/**)",
"Read(__pycache__/**)",
"Read(*.log)",
"Read(*.lock)",
"Read(.env)",
"Read(.env.*)"
]
}
}It costs nothing to maintain and stops whole categories of pointless file reads.
10. Write Specific, Scoped Prompts
This habit separates efficient Claude Code users from expensive ones. Vague requests trigger broad scanning, while specific requests let Claude work with minimal file reads.
Here's the contrast:
Expensive prompt: "Improve this codebase." Claude reads dozens of files trying to work out what "improve" means and explores the whole directory tree. You spend thousands of tokens before any real work happens.
Efficient prompt: "Add input validation to the login function in src/auth/login.ts. Validate that email is a valid email format and password is at least 8 characters. Return a 400 error with a specific message for each validation failure." Claude reads one file, makes the change, done.
Two more habits that save usage on complex tasks:
Use plan mode before implementation. Press Shift+Tab to cycle into plan mode. Claude explores the codebase and proposes an approach for your approval.
If the plan is wrong, you fix it before any code gets written, instead of paying for work you'll throw away.
Press Escape early. If Claude heads the wrong way mid-task, stop it with Escape. Then use /rewind (or double-tap Escape) to restore the conversation and code to a checkpoint.
Catching a wrong turn early is far cheaper than letting it finish and then asking for fixes.
11. Track Your Usage and Set Limits

You can't fix what you don't measure. I covered /usage and Settings > Usage earlier, and during long sessions I check both regularly to see where my budget is going.
For API users: set workspace spend limits in the Claude Console so you don't get surprise bills.
For teams on the API: Anthropic's rate limit recommendations scale by team size.
A team of 1-5 people should plan for 200k-300k tokens per minute (TPM) per user. A 20-50 person team needs roughly 50k-75k TPM per user.
For Pro and Max subscribers: the dollar figure in /usage isn't your bill. Anthropic says that session cost is meant for API users.
The token counts are still real, though. Heavy sessions burn through your included usage faster, so these habits decide how much work fits into your plan.
To make this concrete, here's one creator's /context breakdown on a 1-million-token session, before he sent a single real message: system prompt about 10,000 tokens, tools around 16,400, memory files around 10,700 and skills around 7,000.
He also found a dozen Chrome MCP instances he'd left running at once, each holding its own full context and quietly burning money until he checked. So run /context regularly, not just once.
12. Use the Batch API for Non-Urgent Work (50% Off)
This one is for API users, but it's big enough to include.
Claude's Batch API gives you 50% off both input and output tokens for work that doesn't need an instant answer.
The trade-off is time: a batch can take up to 24 hours to finish.
Where I route work to the Batch API:
Interactive work needs the real-time API. For background jobs and non-urgent generation, cutting input and output costs in half is one of the easiest savings there is.
13. Install Code Intelligence Plugins for Typed Languages
This one comes straight from Anthropic's costs page, and I hadn't seen it mentioned anywhere else.
Code intelligence plugins give Claude precise symbol navigation (go-to-definition, find-references, type info) instead of text search.
A single go-to-definition call replaces a grep followed by reading several candidate files. And installed language servers report type errors after edits, so Claude catches mistakes without running a compiler.
On a TypeScript or Go project, the right plugin can noticeably cut how many exploratory file reads Claude makes per task. It doesn't change what Claude does, just how fast it finds what it needs.
14. Compress Everything Before It Reaches Claude With Headroom
I found this one after this guide first went live, and it earns its own section.
Headroom sits between Claude and everything Claude reads: tool outputs, logs, RAG chunks, files, even conversation history.
It compresses that content before it reaches the model, then reverses the compression on demand if Claude needs the original.
These numbers come from Headroom's own published benchmarks, four test scenarios built from real MCP tool output:
Headroom says the savings depend on how repetitive the content is. Repeated JSON arrays and log lines can clear 90%, while prose barely shrinks.
So run headroom savings on your own traffic to see your real number.
Accuracy held steady in the same tests. GSM8K stayed at 0.870, and TruthfulQA went from 0.530 to 0.560, which Headroom itself says is within the margin of error, not a real improvement.
Setting it up with Claude Code
# 1. Install (Python, ships the CLI) pip install "headroom-ai[all]" # 2. Wrap Claude Code directly headroom wrap claude --memory --code-graph --1m --tool-search # 3. Confirm it is working headroom doctor headroom dashboard
The four flags matter for Claude specifically:
--memory shares a compressed context store across Claude, Codex, and Gemini sessions, so you are not re-paying for the same information twice--code-graph gives Claude a structural map of your codebase instead of raw file dumps--1m targets Claude's long-context mode without burning through it on repeated contentRemove it cleanly with headroom unwrap claude at any point.
One more thing worth knowing: Headroom can also trim what Claude writes back, not just what it reads.
Turn it on with export HEADROOM_OUTPUT_SHAPER=1, and it strips the "Great, let me..." preambles and restated code that pad every response. Output tokens cost 5x input tokens on Opus 5.5 and Fable 5.1, so this matters more than it sounds.
Headroom doesn't replace the hooks and MCP habits from Methods 4 and 8. It automates the same idea at the compression layer, so you keep the savings on sessions where you forget to write a custom hook.
15. Compress Tool Calls at the Source With RTK
Headroom compresses what reaches Claude at the proxy layer. RTK (Rust Token Killer) targets something narrower: Claude Code's own tool-call stream, the hundreds or thousands of small reads and writes it fires off in a single session.
Creator Nick Saraev demonstrated this directly. A single vanilla tool call ran 612 lines and 36,700 characters, most of it repeated boilerplate like "standard out, standard out, standard out."
The same call through RTK dropped to 4 lines and 177 characters, a 99% cut on that one call.
He's upfront that this is a best case, since not every tool call is this redundant. His honest average across a full session is 30-50% fewer tokens with no drop in output quality, because only the repeated noise gets stripped.
RTK and Headroom don't compete. Many builders run both, since they compress different parts of the same pipeline.
16. Rewrite CLAUDE.md With Semantic Compression
Method 2 tells you to keep CLAUDE.md under 200 lines. This is a different move: rewriting what's already there so it says the same thing in fewer tokens.
Most CLAUDE.md files carry filler that reads like a voice memo.
A line like "Hello, thank you so much for helping out with this project, we really appreciate all the work that you do" has zero instructional value. It compresses to "Project instructions and guidelines for the AI system" and loses nothing Claude needs.
In one demonstration, an 865-word CLAUDE.md compressed to 211 words. Estimated tokens dropped from 1,125 to 274, better than a 75% cut, without changing a single rule.
The move: open a Claude session, paste in your CLAUDE.md, and ask it to rewrite the file for maximum information density while keeping every instruction. Compare word counts before and after, then commit the compressed version.
17. Route Large Structured Data Through a Query, Not a Read
If Claude needs to dig through a large log file, a CSV or a spreadsheet, the instinct is to let it read the whole thing. That's expensive and unnecessary.
The fix: load the data into SQLite (or any queryable format) and give Claude a script instead of a file.
For a 5,000-line application log, the instruction becomes explicit: "Important, it's 5,000 lines. Do not read that log directly. Instead, use scripts/logq.py with no arguments to see usage."
Claude then queries for exactly the rows it needs, gets back line numbers and timestamps, and reports the root cause. It never loads 5,000 lines of raw text into context.
This works well beyond logs. For any large structured resource, build a small query function and point Claude at that instead of the raw file.
18. Block Full Reads on Oversized Files
Some files are too big to read start to finish, and they don't need to be. A 618-kilobyte, 20,000-line file should never get read in one shot.
The pattern: let Claude sample the beginning and end first to understand the structure. Then use targeted reads (a sed-style jump to a specific line range) to pull only the section that matters.
In one demonstrated case, this dropped a file read from 20,000 lines to roughly 20-30 lines actually loaded, a saving in the high 90% range on that call.
You don't need this on every tool call. It matters on the reads that would otherwise be enormous.
19. Prompt in English for Token Density
This one takes no setup. English packs more meaning per token than most other Western languages.
A test of the same simple query across languages found:
Mandarin is a notable exception. It's a logographic language and stays token-efficient despite not being English.
If you already prompt in English, this costs you nothing. If your team works in another Western language day to day, expect to pay roughly 1.5 to 2 times more tokens for the same request.
20. Add Context-Frugality Rules to Your System Prompt
This goes a step further than the scoped prompts in Method 10. Instead of writing a specific request every time, put the frugality rule in CLAUDE.md so Claude defaults to careful file access on every task.
A working example:
Read only files directly relevant to the task. Ask before expanding scope beyond three files. Prefer glob or grep to locate, then read the region, not the entire directory. Summarize findings before deciding to read more. Never read generated files, build output, or fixtures unless explicitly asked.
One honest caveat, straight from the source: this can cost you some quality.
Claude's built-in search is already good at guessing where to look, and strict boundaries make it harder for the model to find things on its own. This trade only makes sense if you're willing to write more precise prompts in exchange for a smaller context footprint.
Recent Changes to Claude Usage Limits
Anthropic has changed limits a lot this year. Here's what I could confirm from Anthropic's own posts and docs, newest first.
The 50% boost was extended more than once between July and August. I've only listed the extensions I could confirm from Anthropic's own posts.
Frequently Asked Questions
When do Claude usage limits reset?
Your session limit resets on a rolling five-hour window. On paid plans, the weekly limit resets at a fixed day and time assigned to your account, and it stays the same no matter when you subscribed. You can see both reset times in Settings > Usage, and in Claude Code the limit message tells you when the window resets.
Is there a weekly limit on Claude?
Yes, on paid plans. Pro, Max 5x and Max 20x all have a weekly limit on top of the five-hour session limit, and it applies across all models. On Max you can spend up to 50% of your weekly limits on Fable models.
What happens when I hit my Claude limit?
You're paused until the limit resets, and the message shows when that is. On Pro and Max you can turn on usage credits to keep working at standard API rates, or upgrade to a bigger plan. In Claude Code you can run /usage-credits, and recent versions can wait and pick the task back up after the reset.
Do Claude and Claude Code share usage limits?
Yes. Anthropic says all activity in Claude and Claude Code counts against the same usage limits. So a heavy chat session in the morning leaves less room for coding in the afternoon, and the other way around.
How do I check my Claude usage?
In the Claude app, go to Settings > Usage for progress bars on your five-hour session and weekly limits, plus your next reset time. In Claude Code, run /usage to see your plan usage bars and a breakdown of what is using your limits, like skills, subagents and MCP servers.
How much more usage does Max give than Pro?
Max 5x costs $100 a month and gives you five times the Pro plan's per-session usage. Max 20x costs $200 a month and gives you 20 times. Both still have a weekly limit on top.
What is the difference between /clear and /compact in Claude Code?
/clear wipes the conversation history and starts fresh, so only your CLAUDE.md and configuration reload. /compact keeps the session going but swaps the history for a compressed summary. I use /clear when I switch to an unrelated task and /compact when I need to keep going on the same one.
How do I make my Claude limits last longer?
Clear between unrelated tasks, keep CLAUDE.md under 200 lines, disable MCP servers you are not using, match the model and effort level to the task, and push verbose work to subagents. Those habits cover most of the waste I see.
Final Thoughts
I went from burning through my Claude limits in two days to working comfortably within my plan for a full month. None of this needs premium tools or complex setups.
They're habits and settings that take about an hour to put in place and keep paying off.
Start with the three biggest wins today: run /clear between unrelated tasks, keep your CLAUDE.md under 200 lines, and disable MCP servers you aren't using.
Then check /usage after a few days to see what's still eating your limits, and build from there.




