Most of a Claude Code bill comes from context, not from the code it writes.
Contents
Claude Code is one of the most capable coding agents available, and also one of the easiest to overspend on. The reason is structural rather than mysterious: every turn of a session is a fresh API call that carries the conversation so far, plus every file the agent has read and every tool result it has seen. The tenth turn resends the first nine. Input tokens dominate the bill, and they grow with the length of the session.
Once you see it that way, the habits that save money are obvious — and none of them make the agent worse.
# 1. Compact at every change of direction
/compact asks the model to summarise the session and replaces the history with that summary. Run it whenever the task changes: after a bug is fixed and before the next one, after exploration and before implementation. You can steer it: /compact keep the decisions about the auth module, drop the exploration. Use /clear between unrelated tasks — a fresh context costs nothing.
# 2. Let subagents do the big reads
A subagent has its own context window. If it reads forty files to answer “where is the retry logic?”, those forty files never enter the main conversation — the parent receives a paragraph. Without a subagent, every one of those files would be resent on every later turn. For large codebases this single habit often matters more than model choice.
# 3. Default to the middle tier
Sonnet-class models handle the day-to-day edit–test loop well at a fraction of the flagship price. Switch up with /model for the one genuinely hard decision, then switch back. Point the Haiku tier at a cheap model too: Claude Code uses it for small internal calls on every session.
# 4. Write a real CLAUDE.md
An agent that has to rediscover your build commands, test runner and conventions at the start of every session burns tokens doing it. /init drafts a CLAUDE.md; spend ten minutes editing it. It is the highest-leverage file in the repository.
# 5. Plan first, then edit
Ask for a plan with no edits, correct it in plain language, then say “go”. Plan mode (Shift+Tab in the terminal) enforces this. Exploratory editing — change, test, undo, try again — is the most expensive way to reach a solution.
# 6. Cap headless runs
In scripts and CI, always pass --max-turns and pre-approve tools with --allowedTools. A stuck loop in print mode can quietly run for an hour; a cap turns that into a failed job you can read.
# 7. Check the meter before a long task
/cost shows the session’s spend so far. If a session is already large, compact before starting the next big step rather than after.
# A bonus eighth: let prompt caching work for you
The Messages API can cache a stable prefix — a long system prompt, a large reference document, the unchanging start of a conversation — so that later requests read it from cache at a fraction of the normal input price. Claude Code already uses caching internally, which is one more reason to keep sessions coherent: a session that keeps changing its opening context keeps invalidating the cache. If you build your own agents on the API, put the stable material first and mark it with cache_control.
# What a typical saving looks like
Picture a two-hour refactor on a medium-sized repository. Without any of these habits, the session grows to hundreds of thousands of tokens of history and every turn resends it. With a compact after each sub-task, subagents for the codebase search and a CLAUDE.md that answers the usual questions up front, the same work is typically done with a much smaller running context — the agent writes the same code, it just stops rereading the same files. Teams that adopt the habits often report bills dropping by half or more without any change in output.
# What not to do
Don’t paste entire log files or build outputs into the chat; pipe the relevant lines instead (tail -n 50 build.log | claude -p “…”).
Don’t keep one session open for a whole day across unrelated tasks.
Don’t run a frontier model for mechanical edits such as renames or formatting.
Don’t turn off permission prompts just to save turns — pre-approve the specific safe commands instead.
# Where the price per token comes from
The habits above reduce how many tokens you use. The other half of the bill is what each token costs, and that depends on what sits behind ANTHROPIC_BASE_URL: Anthropic’s API at list price, a Pro or Max subscription with a rolling cap, or a gateway. A breakdown of those options — and why sessions cost what they do — is in this Claude Code pricing guide.
A Claude Code guide covering install, pricing, models and errors.
Whichever backend you use, the order of operations is the same: fix the habits first, then shop for the cheapest tokens. Teams that do both usually find the agent costs a fraction of what their first month suggested. AI Prime Tech is one of the independent gateways that sells the same Claude models below list price, if you want to compare.