Skip to main content

Command Palette

Search for a command to run...

Why Markdown Structure Changes What RAG Retrieves

Structure turns prose into retrievable evidence.

Published
•6 min read•View as Markdown
Why Markdown Structure Changes What RAG Retrieves
J
Founder of AuthorityTech and creator of Machine Relations, the parent category for GEO, AEO, AI SEO, LLM Optimization, and AI PR. Machine Relations (MR) is the discipline I coined in 2024 to describe how brands become legible, retrievable, and citable inside AI-driven discovery systems. As search shifts from human-mediated discovery to machine-mediated answers, most of the market describes fragments of the same change using labels like GEO, AEO, AI SEO, LLMO, and AI PR. My position is that these are not separate categories. They are tactics or layers inside the broader system of Machine Relations (MR).

A retriever does not read your document the way a human does. It scores fragments. That means headings, lists, and paragraph boundaries are not cosmetic. They change what gets chunked, what gets embedded, and what survives the first pass.

The problem: retrieval starts before generation

RAG systems usually fail long before the model writes an answer. The failure happens in retrieval. If the wrong fragment enters the candidate set, the generator can only polish the mistake.

That is the key point in Lewis et al.'s RAG paper: retrieval gives the generator access to explicit external memory, but the system still depends on the quality of the passages it can select. If the retriever never surfaces the right evidence, the generator is trapped inside the wrong context.

The same problem shows up in document-level claim extraction work. A document can be about the right thing and still fail to be retrieved if its structure makes the relevant passage hard to isolate.

Why headings matter more than people think

Markdown headings create retrieval landmarks.

A retriever works on chunks, not on the document as a whole. When a page has clear section boundaries, the chunking step can keep related claims together. When a page is one long wall of prose, the chunker often slices straight through the idea you wanted preserved.

This matters for three reasons.

First, heading text itself becomes a signal. A section labeled "Retrieval granularity" tells the indexer what the next few paragraphs are likely to contain.

Second, section boundaries reduce semantic drift. A paragraph under a specific heading tends to stay on one topic. That makes embeddings cleaner and ranking less noisy.

Third, headings help the retriever land on the right neighborhood even when the answer lives deeper in the document. The retriever does not need the exact sentence immediately. It needs the correct region.

The practical result is simple: structured writing is easier to retrieve.

What the papers show

Two patterns show up across the research.

  1. Retrieval systems improve when they can search a better representation of the source.
  2. Generation improves when the retriever can isolate fine-grained evidence instead of a vague topic blob.

Sentence-level factual reasoning is useful here. It argues that generic, topic-level retrieval misses nuanced facts, and that sentence-level planning produces better queries for fact retrieval. That is a structural claim, not a stylistic one. Fine-grained structure gives the retriever a smaller search space and more precise anchors.

MoC: Mixtures of Text Chunking Learners adds a sharper angle. It treats chunking itself as a learnable problem and shows that coarse chunk boundaries are part of the failure mode. That matters because bad structure can look like a generation problem when it is actually a retrieval problem. If the right passage never entered the context, the model never had a chance.

A small experiment you can run

You do not need a giant corpus to see the effect. You can test it with a tiny local index.

import fs from "node:fs";
import path from "node:path";

function chunkMarkdown(text) {
  const sections = text.split(/^##\s+/m).filter(Boolean);
  return sections.map(section => {
    const [heading, ...bodyParts] = section.split(/\n/);
    return {
      heading: heading.trim(),
      body: bodyParts.join("\n").trim()
    };
  });
}

function scoreOverlap(query, chunk) {
  const q = new Set(query.toLowerCase().split(/\W+/).filter(Boolean));
  const text = (chunk.heading + " " + chunk.body).toLowerCase();
  let hits = 0;
  for (const term of q) if (text.includes(term)) hits++;
  return hits;
}

const docs = fs.readdirSync("./docs").map(file => ({
  file,
  text: fs.readFileSync(path.join("./docs", file), "utf8")
}));

const index = docs.flatMap(doc =>
  chunkMarkdown(doc.text).map(chunk => ({ ...chunk, file: doc.file }))
);

const query = "retrieval granularity and markdown headings";
const ranked = index
  .map(chunk => ({ ...chunk, score: scoreOverlap(query, chunk) }))
  .sort((a, b) => b.score - a.score)
  .slice(0, 5);

console.log(ranked);

This is not a production retriever. It is a cheap way to show the mechanics.

For a stronger baseline, swap the overlap score for BM25 or a vector search index and keep the same chunking split. The point is not the ranking algorithm. The point is whether the structure gives the retriever something clean to rank.

If you compare the output of a document written with clear headings against the same content flattened into one blob, the structured version usually wins for two reasons. More query terms land in the right chunk, and the heading itself acts like a label for the passage.

The deeper mechanism

The deeper issue is not markdown. It is segmentation.

Every retriever has to decide where one unit ends and the next begins. That decision shapes the candidate set. If a paragraph mixes definitions, examples, and caveats, the embedding becomes averaged mush. If the same ideas are split into a heading, a definition, and a separate example block, the retriever gets cleaner objects to rank.

That is why this problem keeps showing up in practice. Retrieval is not just about finding documents. It is about manufacturing good retrieval objects.

This is also why structured sources often outperform dense prose in AI systems. Tables, lists, and tight sections compress better. They create clearer semantic boundaries. They give the retriever a better chance of landing on the exact evidence instead of a nearby paragraph that merely sounds related.

What this means for source visibility

For AI search and RAG, citable content is usually structured content.

A source becomes easier to reuse when it exposes stable anchors, scoped sections, and concise claims. That does not guarantee citation, but it increases the odds that the relevant fragment survives retrieval.

This is where the Machine Relations angle matters. AI engines do not just reward authority in the abstract. They reward sources that can be broken into clean evidence units. The document has to be legible to the retriever before it can be legible to the generator.

So the rule is blunt: if you want a source to be retrievable, write it like something a retriever can split cleanly.

The limits

Markdown is not magic. A bad document with neat headings is still a bad document.

Headings do not fix weak evidence, vague claims, or missing provenance. They only make the underlying material easier to index. If the source is thin, structure just makes the thinness easier to see.

The other limit is that retrievers differ. Some systems lean hard on embedding similarity. Others add lexical search, reranking, or query expansion. A heading that helps one system may matter less in another. But across systems, the direction is consistent: clearer structure usually makes retrieval easier.

The practical rule

If the goal is to be found by a retrieval system, do not write for the page. Write for the chunk.

Use headings that describe the next block precisely. Keep one idea per section. Put definitions near the top. Put examples under their own label. Separate caveats from claims. Treat every section as if it might be indexed alone, because it probably will be.

That is the real lesson. The document is not just content. It is retrieval substrate.