The Cutting Room·Part 6 of 6·4 min read

The Findability Test: How to Know You Cut It Right

Chunking is editing, and editing needs a reader

Pritish Maheta·

Five parts of craft: the missing pages, the Goldilocks zone, the author's seams, the tables that fight back, the meaning seam. Every one of them ends at the same cliff — you've cut the library, the chunks are floating in place, and a question is coming. Did you cut it right?

Here's the uncomfortable truth the finale exists to fix: almost nobody checks. Teams tune chunk sizes by vibes, ship, and discover their cutting failures the way the HR bot's client did — from a director quoting the interns' notice period. The feedback arrives in production, wearing a user's face.

There's a cheaper reader. The Cutting Room closes with the discipline that separates chunking-as-guesswork from chunking-as-craft: the Findability Test.

The test, in one afternoon

The principle first, because it reframes the job: a chunk isn't good because it looks clean — it's good because real questions find it. Findability isn't a property of the pieces; it's a property of the pieces meeting the questions. So the test needs both sides.

Step 1: Collect fifty real questions. Not questions you invent — questions your users actually ask, harvested from support tickets, search logs, chat history, or a walk around the team. Invented questions share your vocabulary and your assumptions; real ones arrive sideways, in the customer's words, which is precisely what makes them a fair exam. (The Vocabulary Wall's lesson applies to test design too.)

Step 2: For each question, mark where the answer lives. The document and passage a knowledgeable human would point to. This — your answer key — is the tedious part, and it's an afternoon, not a project. Fifty question-to-passage pairs is a golden set, and it will outlive every other artifact in your pipeline.

Step 3: Run the questions against your chunks. For each one, look at the top handful retrieval returns and ask the only question that matters: is the passage from the answer key in there? Count the hits. That percentage — call it your findability rate — is the number this whole series has been building toward. Not answer quality, not model brilliance: did the right piece even make it onto the genius's desk? (Everything downstream is capped by it — the Open-Book series' funnel rule.)

Reading the failures like an editor

The rate tells you whether you have a problem. The misses tell you which one — and after five parts, you can read them like a diagnostician. Pull each missed question and look at the chunks that contain the answer:

Answer straddling two chunks → split failure: revisit boundaries, add overlap (Part 2). Answer inside a sprawling multi-topic chunk → phone book: cut finer, follow the seams (Parts 2–3). Chunk contains the answer but never names its topic → the answer lost its name: stamp lineage labels (Part 3). Rows of context-free numbers → decapitated table: back to Part 4's extraction bench. Transcript cut mid-topic → bring the meaning-scissors (Part 5).

Fix the dominant failure, re-run the same fifty questions, watch the rate move. That loop — test, diagnose, re-cut, re-test — is the entire discipline. Chunking stops being a config value you set once and becomes what it always really was: editing. And editing has never been judged by the editor. It's judged by whether the reader finds what they came for.

The golden set compounds

One more return on the afternoon: the golden set becomes infrastructure. Considering a new embedding model? Run the fifty. New chunk strategy? Run the fifty. Documents re-ingested after a policy update? Run the fifty. Every change to the pipeline gets a before-and-after number instead of a vibe. When the Open-Book series said RAG converts hallucination "from a mystery into a maintenance schedule," this is the maintenance schedule — and it costs one afternoon plus discipline.

The catch

The finale's honest limit: findability is necessary, not sufficient. A perfect rate means the right pieces reach the desk — it says nothing about whether the documents were current (the expired-medicine cabinet), whether the model reads them faithfully, or whether the answer synthesized from them is true. Those failures belong to other benches in other workshops (the parent series' "When RAG Lies" tours them all). The Cutting Room's promise was narrower and it's now kept: the failures the scissors own, you can now see, name, test, and fix.

The one line to remember

You don't know if you cut well until real questions try to find the pieces — fifty real questions and an honest count beat every chunking opinion on the internet.

That's The Cutting Room. It began as one article in The Open-Book AI and became six, because this is where the systems I get called about actually break. If your findability rate just turned out to be a number you don't want to say out loud, my inbox is open.

Facing this problem in production?

I help teams make AI systems smaller, faster, and cheaper — from distillation to full MLOps pipelines.

Work with me