September 3, 2026

When Claude Code reports 'Model overloaded', try forking the session

When Claude Code reports 'Model overloaded', try forking the session

A Claude Code session can reach a state where every attempt to continue it fails the same way. The session shows Service is busy, offers Try again, and the retry counter climbs to Retrying (10/10) before stopping. According to Anthropic's documentation, "Claude Code retries transient failures up to 10 times with exponential backoff before showing you an error." The error card's own advice — try again in a moment, or switch to a different model — is a good first move. On several sessions across 9/3/2026, that approach did not recover the session: repeated retries and repeated model switches, minutes apart, kept failing. We forked one of those sessions instead, and the fork picked up the same conversation and continued normally.

What made that worth paying attention to was what happened next. After the fork was already working, we went back to the original session and pressed Try again. It failed. We switched models there as well, and it failed again, while the fork continued running the same work. That does not tell us why the original session remained stuck, but it does make the result more interesting than a fork that happened to coincide with capacity returning.

A Claude Code error card headed 'Service is busy', reading 'Try again in a moment, or switch to a different model', with a View details button on the left and a Try again button beside a model picker showing Opus 5 on the right. The error message that prompted the creation of this article. Anthropic recommends retrying or switching models; across several of our 9/3 sessions, neither recovered the original session.

What a 529 Overloaded error is

A 529 is the Claude API's overloaded_error, which Anthropic's documentation defines as "The API is temporarily overloaded" and attributes to "high traffic across all users." The Claude Code error reference is blunter about what it is not: "A 529 is not your usage limit and doesn't count against your quota." Anthropic’s guidance is straightforward: check status.claude.com for a capacity incident, wait a few minutes, and use /model to switch to another model.

So the shortage sits on Anthropic's side, and nothing you do in the client creates capacity.

The expanded 'Service was busy' detail in Claude Code, showing the text: API Error: 529 Overloaded. This is a server-side issue, usually temporary — try again in a moment. If it persists, check https://status.claude.com. The same failure with its detail open.

When switching models isn’t enough

Switching models is Anthropic's own advice and it did not revive our stalled sessions. The reason Anthropic gives for the advice is that "capacity is tracked per model" — the error reference says so, and adds that Claude Code raises the suggestion itself when a single model is carrying heavy load, in the form "Opus is experiencing high load, please use /model to switch to Sonnet". We switched models from the error card's dropdown several times, waited minutes between attempts, and got the same failure on the same session each time.

We cannot tell you why a particular session stops accepting requests while the account keeps working elsewhere, because Anthropic does not document a per-session capacity behavior and we have not measured one. What we can tell you is what we do instead of continuing to press Try again.

A Claude Code status line reading 'Model overloaded · Retrying (4/10) · 23s', with a Try again button above it. Four attempts in, 23 seconds elapsed. Claude Code is retrying on its own before it shows you anything.

Fork the session

A fork writes everything said so far to a second session and puts you in that second session. The session you were in is left exactly as it was. In the Claude desktop app and at claude.ai/code, right-click the session in the sidebar and choose Fork session, or press Cmd+Option+O on macOS and Ctrl+Alt+O on Windows and Linux. From a message rather than a session, the same menu offers Fork from here.

From the terminal, Anthropic documents two routes on the session management page: run /branch inside the session, optionally with a name, or combine a resume flag with --fork-session.

Either route ends with two independent sessions. Anthropic's architecture overview describes the split: forking "copies the history into a new session ID, leaving the original unchanged." The /branch confirmation prints both session IDs, and the original stays in the session picker, so you lose nothing if the fork does not help.

What a fork carries over, and what it does not

Forking carries the conversation into a new session. When you branch from inside the running session with /branch, Anthropic documents several things that carry over with it:

  • Conversation history is copied into the fork up to the moment you branch.
  • Permission grants you approved with "Allow for this session" carry over, because /branch keeps the same running process. Start the fork as a separate process with --fork-session and those approvals are requested again.
  • Anything already running in the background — a subagent, a background Bash command — carries on, and you see its output in the fork rather than in the session you left.
  • A Remote Control connection follows you into the fork, so a phone or browser watching the session keeps receiving messages.

Two limits matter. A fork branches the conversation, not the filesystem — edits a forked session makes are real, and visible to anything else working in that directory. And the fork carries the history, so the next request still has to process it: Anthropic notes that a session over roughly 100,000 tokens that has been idle for about an hour has an expired prompt cache and will process the full history once regardless.

What we can claim, and what we cannot

Forking is not documented by Anthropic as a remedy for capacity errors, and we are not presenting it as one. Anthropic's documented remedies for a 529 are the status page, waiting, and /model.

Our claim is first-hand and has one piece of evidence in it worth separating from the rest. The weak version — "we forked and it worked" — is consistent with the shortage ending at that moment on its own. The strong version is what we actually saw: the original session kept failing after the fork was already working. Same account, same machine, same directory, same conversation, minutes apart, and one of them ran while the other returned the same error to both Try again and a model switch. The original session kept failing after the fork was already working. That sequence is harder to explain as simply “capacity happened to return while we were forking,” although it does not tell us what differed between the two sessions.

Anthropic does not document any per-session behavior for a 529, and we have not instrumented one — we have a repeated observation, not a mechanism. We also have not run a controlled comparison, cannot give you a success rate, and would not expect a fork to help at all during a platform-wide incident, which the status page is there to tell you about.

We reach for a fork early anyway because it costs almost nothing to try. A fork takes one shortcut, leaves the original session on disk, and either works or leaves you exactly where you were. Ten more minutes of pressing Try again on a session that has already spent a ten-attempt retry budget costs ten minutes and has already failed once.

A Claude Code status line reading 'Model overloaded · Retrying (10/10) · 3m 42s'. Anthropic documents the limit as ten retries with exponential backoff

What forking costs

A fork is cheap, but it is not free. Anthropic documents three of the costs below; the fourth is a limit of what forking is for.

A fork carries the whole conversation, so it does not make later requests smaller. Anthropic's cost guide explains why that matters. Claude Code continues to account for the conversation context on later requests, so a trivial follow-up in a long-running session can still incur input usage based on that much larger history. Forking a long session gives you a second long session. If what you actually want is a smaller context, reach for /compact or a fresh session instead — forking gives you the opposite.

Two sessions now exist on disk and in the picker. /branch copies the transcript, both sessions keep their own IDs, and Anthropic's session page notes that forked sessions appear as separate rows. A second row is the whole point when you want to go back to the original, and clutter when you never do.

The cost that surprised us is the one attached to the other remedy. Switching models forces a cache miss; a same-directory fork can reuse the parent's warm cache. Anthropic's prompt-caching page explains that prompt caches are model-specific, so changing models with /model prevents the next request from reusing the cache built for the previous model. A fork behaves differently. Because it begins with the parent's existing context, Anthropic documents that its first request can reuse the prompt cache already associated with that session history.

So the advice everyone reaches for first is the expensive one, and we took it several times on a long session before forking. Cached input is billed at roughly a tenth of the standard input rate, which is the gap each model switch gives up.

One limit on that. Anthropic scopes the cache per machine and per directory, and says explicitly that this "includes worktrees of the same repository" — so a fork you continue in a different worktree builds its own prefix instead of reading the original's, and the saving above does not apply to it.

None of this makes forking expensive. Forking is cheap in exactly the way people assume a model switch is.

Two settings that change how often a session stalls

Two configuration options change how Claude Code behaves when a model is overloaded. Neither one changes the capacity behind the error.

A fallback model list answers overload without your involvement. Claude Code's model configuration guide documents --fallback-model for a single session and a fallbackModel array in settings to keep the list across sessions. Claude Code tries the listed models in order whenever the model the session is on is overloaded or unavailable, and accepts at most three.

{
  "fallbackModel": ["claude-sonnet-5", "claude-haiku-4-5"]
}

The retry watchdog suits unattended work. Anthropic's error reference documents CLAUDE_CODE_RETRY_WATCHDOG, which, set to 1, keeps retrying capacity errors — both 429 throttles and 529 overloads — rather than giving up at the configured attempt limit. Setting it in an interactive session means a session can sit retrying instead of telling you to intervene, so it belongs in continuous integration jobs and scheduled runs, not on your laptop.

Questions we get

Does a 529 count against my usage limit? No. Anthropic's Claude Code error reference states that "A 529 is not your usage limit and doesn't count against your quota." A 429 is the rate-limit and spend-cap error, and Anthropic's API error page documents the two separately.

How many times does Claude Code retry before showing me the error? Up to ten, with exponential backoff, per the error reference — which is the (10/10) in the retry label. Some failure classes get a smaller budget or none; that page lists which.

Will forking lose my conversation? No. A fork writes the conversation to a second session under a new ID and does not touch the session you forked from, per Anthropic's architecture overview. Both sessions stay in the session picker with their own IDs, so you can go back.

Is forking the same as /compact or rewinding? No, and Anthropic's checkpointing page draws the line explicitly: a summarize action stays inside one session and shrinks its context, whereas /branch and claude --continue --fork-session create a second session and leave the session you started in untouched. A rewind is different from both: a rewind puts code or conversation back to an earlier point inside a single session.

Does forking cost extra tokens? Not for the fork itself. Anthropic's prompt-caching page explains that a fork can reuse cached context from its parent on its first request. The switch that does cost is /model: every model is cached separately, so a model switch pays to process the entire history again. What a fork will not do is make the conversation smaller — it carries the full history, so every later request in it is as large as it was before.

Should I switch models when a session is overloaded? Yes, try a model switch first — a switch is Anthropic's documented advice, and capacity is tracked per model. Just check which kind of switch you got: a /model change holds for the rest of the session, while a switch made by a fallback chain applies to a single turn and then reverts to your primary model.

Thank you for your time

If you have any questions or want to connect on anything that I wrote about above, please email me or book some time on my calendar. Any and all feedback is of course so appreciated.