Semantic Chunking vs Fixed-Size Chunking: What Actually Wins in Production RAG
Semantic chunking sounds smarter than splitting on 512 tokens, but in production it often loses to boring fixed-size chunks. Here's when each one wins, and what we ship by default.

Every RAG project hits the same fork in the road around week two: do we split documents on a fixed token count, or do we get clever with semantic chunking? The clever option feels obviously better. In practice, we've shipped enough retrieval systems to know it usually isn't — and when it is, it's for reasons nobody talks about in the tutorials.
This is the tradeoff breakdown we wish we'd had before we spent a sprint tuning breakpoint thresholds on a 40k-document corpus.
What each approach actually does
Fixed-size chunking splits text every N tokens, usually with some overlap. That's it. You pick 400 tokens with 50 overlap, run it through a tokenizer, and move on.
Semantic chunking tries to split on meaning boundaries. The common recipe, popularized by Greg Kamradt and now baked into LangChain and LlamaIndex, is:
- Split the document into sentences.
- Embed each sentence (or small sentence groups).
- Compute cosine distance between adjacent embeddings.
- Cut wherever the distance exceeds a percentile threshold (say, the 95th percentile of all distances in the doc).
The pitch: chunks end where topics end, so retrieval returns coherent ideas instead of half-sentences.
Why the pitch is only half true
Semantic chunking optimizes for intra-chunk coherence. Retrieval quality depends on something different: whether the chunk contains the specific span that answers the query, and whether your embedding model can find it. Those are related, but not the same thing.
A semantically pure chunk about "Q3 revenue drivers" is useless if the user asks "how much did we spend on AWS in July" and that number sits three paragraphs later in what the splitter decided was a different topic.
The costs nobody mentions
Before we get to quality, look at what semantic chunking costs you at ingest time.
- Embedding calls per sentence, not per chunk. A 50-page PDF might have 1,200 sentences. That's 1,200 embedding calls just to decide where to cut, before you embed the final chunks.
- Non-deterministic chunk sizes. Some chunks come out at 80 tokens, some at 1,500. That breaks downstream assumptions about context budgets and reranker input sizes.
- Threshold tuning per corpus. The 95th-percentile trick works differently on legal contracts (uniform prose) than on Slack exports (bursty, choppy). You end up hand-tuning per document type.
- Reindexing is expensive. If you decide to change the threshold six months in, you re-embed everything twice.
In our experience, semantic chunking adds roughly 3–8x to ingest cost and 2–4x to ingest wall-clock time compared to fixed-size on the same corpus. That's not a dealbreaker for a 500-document knowledge base. It is one for 5 million support tickets.
When fixed-size actually wins
Fixed-size chunking wins more often than the RAG discourse suggests. Specifically:
- Homogeneous prose. Product docs, wikis, textbooks, transcripts. The topic shifts are gentle, and a 400-token window with 15% overlap catches most answers.
- High query volume, tight latency budgets. You're going to add a reranker anyway (see our note on rerankers in RAG). The reranker cleans up the noise from slightly awkward chunk boundaries.
- Corpora that change constantly. Nightly re-ingest of a support inbox is fine with fixed-size. It's painful with semantic.
- You need predictable token accounting. If your prompt template assumes "top 6 chunks fit in 3k tokens," variable-length chunks blow that up.
Here's the boring baseline we start almost every project with:
from tiktoken import get_encoding
enc = get_encoding("cl100k_base")
def fixed_chunks(text: str, size: int = 400, overlap: int = 60):
tokens = enc.encode(text)
chunks = []
step = size - overlap
for i in range(0, len(tokens), step):
window = tokens[i : i + size]
if len(window) < 40: # drop tail scraps
break
chunks.append(enc.decode(window))
return chunks
That's the whole thing. It runs in milliseconds per document, produces uniform chunks, and gives you a solid baseline to beat.
When semantic chunking earns its keep
We reach for semantic chunking in three specific situations:
1. Mixed-format documents
A 10-K filing has narrative sections, tables, footnotes, and boilerplate all glued together. Cutting every 400 tokens will slice a table in half or bury a footnote reference. Semantic chunking (or better, structure-aware chunking that respects headings and tables first, then falls back to semantic within sections) meaningfully improves retrieval here.
2. Long-form reasoning content
Research papers, legal opinions, RFCs. Documents where the answer to a question depends on a chain of argument that spans several paragraphs. Fixed-size chunks often split the premise from the conclusion. Semantic chunks are more likely to keep them together.
3. When you're not using a reranker
If you're going straight from vector search to the LLM with no reranking stage, chunk quality matters much more. Semantic chunking gives the top-k results a better shot at being self-contained. Add a reranker and this advantage shrinks.
Structure first, semantics second
Even in these cases, we rarely use pure semantic chunking. The pattern that works best is layered:
def layered_chunks(doc):
# 1. Split on document structure (H1/H2, page breaks, table boundaries)
sections = split_on_structure(doc)
chunks = []
for section in sections:
if token_count(section) <= 600:
chunks.append(section) # keep intact
else:
# 2. Within oversized sections, use semantic splits
chunks.extend(semantic_split(section, max_tokens=600))
return chunks
Structure gives you 80% of the coherence benefit for free, because the author already decided where the topic boundaries are — they're called headings. Semantic splitting handles the leftover oversized sections.
What the eval numbers actually look like
We don't publish specific benchmark numbers because they're misleading out of context, but the shape we see across projects is consistent:
- On clean prose corpora, fixed-size + reranker matches semantic chunking on recall@10, within noise.
- On mixed-format corpora, structure-aware chunking beats both by a clear margin on answer accuracy in human eval.
- Pure semantic chunking rarely wins outright once a reranker is in the pipeline. It wins on ingest simplicity of the retrieval prompt, not on retrieval quality.
If you want to measure this yourself, build a small eval set of 100–200 real user questions with graded answers, and run each chunking strategy through the same retriever and generator. Don't trust vibes. We've written before about eval discipline on the 72Technologies blog.
Practical guardrails regardless of strategy
A few things we now do on every RAG project, independent of chunking approach:
- Store the parent document ID and character offset with every chunk. You will need this for citations, deduping, and expansion at query time.
- Keep a small overlap even in semantic chunks. Pure semantic splits with zero overlap lose context at boundaries. 30–50 tokens of overlap costs almost nothing.
- Cap max chunk size hard. Even in semantic mode, enforce a ceiling (say, 800 tokens). Otherwise a rambling section becomes a single giant chunk that dominates retrieval scores unfairly.
- Log which chunk got retrieved for which query. Retrieval debugging is impossible without this, and it's the fastest way to spot chunking pathologies.
- Version your chunking config with your index. When you change chunk size or strategy, you need a new index, not a mutated one. Treat the config as part of the index identity.
Where we'd start
On a new RAG project, we ship fixed-size chunking (400 tokens, 60 overlap) with a reranker on day one. It's cheap, deterministic, and easy to reason about. We only move to structure-aware or semantic chunking when the eval set tells us we have to — usually because the corpus has real structure that fixed-size is destroying, or because the questions require multi-paragraph context that keeps getting split.
The rule we follow: don't optimize chunking until retrieval eval scores are bad and you've confirmed the reranker isn't papering over the problem. Most teams jump to semantic chunking because it sounds sophisticated. The sophisticated move is measuring first, and picking the boring option when it works. If you want help building that eval loop or deciding what your corpus actually needs, that's the kind of work our AI engineering team does most weeks.
Want a team like ours?
72Technologies builds production software for the kind of teams who actually read this blog.
Start a projectKeep reading

Agent Loops That Don't Spiral: Hard Limits, Budgets, and Escape Hatches
Autonomous agents fail in two ways: they stop too early, or they never stop. Here's how we design loops that terminate cleanly, stay under budget, and tell you why they quit.

Fallback Chains for LLM Calls: Designing Around Provider Outages Without Breaking Latency SLAs
When Anthropic goes down for two hours, your product still needs to answer. Here's how we design fallback chains across Claude, GPT, and Gemini without blowing latency budgets or shipping garbage responses.
Evals That Actually Catch Regressions: Building an LLM Test Suite You Trust
Most LLM eval suites are theatre. Here's how we build ones that actually block bad deploys — with graders that agree with humans, deterministic seeds, and a scoring rubric your PM can read.
