All articles
AI & LLMsSeptember 20, 2026 6 min read

Context Window Budgeting: How We Actually Split 200K Tokens in Production

A 200K context window sounds like freedom until latency spikes and answers get worse. Here's how we allocate tokens across system, retrieval, history, and scratch in real apps.

The first time a team gets access to a 200K or 1M token context window, they tend to do the same thing: stuff it. Then latency doubles, cost triples, and answer quality quietly drops. Long context is a budget, not a bucket, and the teams shipping good LLM products treat it that way.

This is how we plan token allocation on real projects — the split between system prompt, retrieved context, conversation history, tool schemas, and generation headroom — and the tradeoffs we've made when the numbers stop cooperating.

Why bigger windows don't automatically mean better answers

Anthropic, OpenAI, and Google all publish context lengths that look generous on paper. Claude Sonnet family models advertise 200K tokens, GPT-4.1 class models sit around 1M, and Gemini 1.5/2.x Pro pushes into the multi-million range for select tiers (see each vendor's model docs for current limits). What none of the marketing pages emphasize:

  • Attention degrades unevenly across position. The "lost in the middle" effect is well-documented in research and reproducible in practice — content buried between a large preamble and a large tail gets weighted less than content near the edges.
  • Latency scales with input tokens, not just output. Prefill is cheaper per token than decode, but a 150K-token prompt still adds real wall-clock time before the first token streams.
  • Input cost is not free. Even at input rates 3–5x cheaper than output, a chatty RAG app pushing 80K tokens per turn will out-cost the generation itself within a few messages.

The practical read: treat the window like a memory budget in an embedded system. Every section has a cap, and something has to be evicted when a new claim shows up.

A default budget we start from

For a typical assistant-style product with RAG and 3–6 tools, this is roughly where we land before tuning:

SectionShare of windowNotes
System prompt + policies2–5%Stable, cache-friendly
Tool schemas2–8%Grows fast with poorly-designed tools
Retrieved context (RAG)30–50%The workhorse
Conversation history15–30%Compressed, not raw
Scratch / agent reasoning5–15%Only if using agents
Generation headroom10–20%Reserve, don't share

For a 200K window, that puts retrieval around 60–100K tokens, history around 30–60K, and a firm ceiling on system + tools around 20K combined. If you're at 1M tokens, the ratios shift — retrieval usually stays capped by relevance rather than budget, and history gets more room.

The generation headroom rule

One mistake we see repeatedly: teams don't reserve output tokens explicitly. If your model supports 8K output and you fill the window to 199K on input, the request either fails or the model truncates its own answer awkwardly. Reserve max_tokens up front and treat the remaining space as your real input budget.

MODEL_WINDOW = 200_000
MAX_OUTPUT = 8_000
SAFETY_MARGIN = 2_000  # tokenizer drift, tool call overhead

INPUT_BUDGET = MODEL_WINDOW - MAX_OUTPUT - SAFETY_MARGIN
# 190,000 tokens to spend across system, tools, RAG, history

Allocating retrieval: quality over quantity

The temptation with a big window is to pass 40 chunks instead of 8. In our experience this usually hurts. More retrieved chunks means:

  1. More chance a distractor chunk pulls the model off-topic.
  2. Higher input cost per turn, compounding across a session.
  3. Slower prefill, which users feel as a lag before streaming starts.

A better pattern: retrieve wide, rerank hard, pass narrow. Grab 40–100 candidates from your vector store, rerank with a cross-encoder or a small LLM, then pass the top 6–15 into the final prompt. The window is there as insurance for edge cases (a legal doc that legitimately needs 30 chunks), not as the default.

Position matters — put the important stuff at the ends

When you do pass multiple chunks, order them by relevance score in a U-shape: highest-relevance chunks at the top and bottom of the retrieved block, medium relevance in the middle. It's a small change that measurably improves answer quality on longer contexts, and it costs nothing.

def arrange_chunks(ranked_chunks):
    # ranked_chunks sorted by relevance desc
    top, bottom, middle = [], [], []
    for i, chunk in enumerate(ranked_chunks):
        if i % 2 == 0:
            top.append(chunk)
        else:
            bottom.append(chunk)
    bottom.reverse()
    return top + middle + bottom

Conversation history: summarize, don't accumulate

Raw chat history is the section that quietly eats your budget. A 30-turn support conversation with tool calls and citations can easily hit 40K tokens on its own. Two patterns we rely on:

Rolling summaries. Keep the last N turns verbatim (usually 4–8), and replace older turns with a compact summary generated by a cheaper model. The summary lives in the system-prompt-adjacent area and gets refreshed every M turns.

Structured memory. For agents that need to remember facts (user preferences, prior decisions), extract them into a JSON blob rather than relying on the model to re-derive them from history. This is far more token-efficient and more reliable.

{
  "user_facts": {
    "timezone": "Europe/Warsaw",
    "preferred_stack": ["Next.js", "Postgres"],
    "open_tickets": ["INC-2041"]
  },
  "summary_of_earlier_turns": "User debugged a webhook signature mismatch..."
}

This approach also plays well with prompt caching on Claude and Gemini — the structured block is stable enough to hit the cache on most turns.

Tool schemas: the silent budget killer

If you're using function calling or tool use, every tool definition sits in the prompt on every request. A verbose schema for 12 tools can push past 15K tokens before you've said hello to the user.

Things we do to keep this under control:

  • Tool routing. A cheap classifier (small model or embedding similarity) picks the 3–5 relevant tools for a given query and only those get sent to the main model.
  • Trim descriptions. Tool descriptions get written like technical docs, not marketing. One or two sentences, an example, done.
  • Collapse similar tools. Three tools that all query the same database with different filters should be one tool with a mode parameter.

Measuring what you're actually sending

You cannot budget what you don't measure. Every LLM call in our production apps logs a token breakdown, not just a total:

usage_log = {
    "request_id": req_id,
    "model": "claude-sonnet-x",
    "tokens": {
        "system": count(system_prompt),
        "tools": count(tool_schemas),
        "retrieval": count(rag_context),
        "history": count(chat_history),
        "user_message": count(user_msg),
        "output_reserved": MAX_OUTPUT,
        "output_actual": response.usage.output_tokens,
    },
    "latency_ms": elapsed,
    "cache_hit_tokens": response.usage.cache_read_input_tokens,
}

Once this lands in your observability stack, budget violations become obvious. We alert when any single section exceeds its allocated share by more than 20% for more than 5% of requests in a rolling window. Usually it's a runaway history buffer or a new tool someone added without trimming its schema.

When to reach for a bigger window vs. a smarter pipeline

A useful heuristic: if you're consistently using less than 40% of your current window, a bigger window won't help you. If you're bumping the ceiling and quality is still good, either move to a longer-context model or invest in better retrieval — the answer depends on where the pain is.

Bigger window makes sense when:

  • The task genuinely requires reasoning over a long, coherent document (contract analysis, codebase Q&A on a single service).
  • You've already tried tighter retrieval and lost real information.
  • Latency budget can absorb longer prefill.

Smarter pipeline makes sense when:

  • Answers degrade as you add more context (classic distractor problem).
  • Cost per session is the constraint, not per-request capability.
  • Users expect fast first-token time.

Where we'd start

If you're inheriting an LLM app that's slow, expensive, or subtly wrong, do this in order:

  1. Add per-section token logging to every call. You'll be surprised where the tokens go.
  2. Set a hard max_tokens and back into a real input budget from the model's window.
  3. Cap retrieved chunks at 10–15 and add a reranker before you touch anything else.
  4. Replace raw chat history with a rolling summary plus a structured memory blob.
  5. Route tools instead of sending all of them every turn.

Most of the wins come from discipline, not from a bigger model. The teams we work with who treat context as a budget end up with cheaper, faster, and more accurate systems than the ones who keep upgrading to longer windows and hoping the model figures it out. If you want a hand auditing an existing setup, our AI engineering services page has more on how we approach this.

#LLMs#RAG#Claude#Gemini#Prompt Engineering#Cost Optimization

Want a team like ours?

72Technologies builds production software for the kind of teams who actually read this blog.

Start a project