Skip to content

Pipeline Overview

Every indx run is a DirectoryPipeline: six ordered stages that turn a directory or ZIP into a KnowledgeSpace and, when output is enabled, a sealed .indx archive.

Each stage receives and returns one shared, mutable SpaceContext. Work accumulates in that context as the run proceeds.

The pipeline runs in a fixed canonical order:

01 Walk → 02 Parse → 03 Chunk → 04 Relate → 05 Enrich → 06 Embed+Pack

Three stages are pure built-ins; the other three drive one or more swappable components. Use this table as the stage map:

| # | Stage | Responsibility | Primary component | |---|-------|----------------|-------------------| | 01 | Walk | Traverse folder/ZIP, build the directory graph, detect file types | — (built-in) | | 02 | Parse | Run each file through the chosen parser (text, tables, layout, images) | Parser | | 03 | Chunk | Split content with structure intact; keep lineage and neighbor links | — (built-in) | | 04 | Relate | Resolve references, siblings, parents, duplicates into typed Relations | — (built-in) | | 05 | Enrich | LLM/VLM add detected type, topics, tags, summaries | LLM, VLM | | 06 | Embed+Pack | Vectorize chunks, write vectors to the store, seal the .indx archive | Embedder, Store, OutputWriter |

A single SpaceContext is threaded through the whole run. The command-line entry point constructs the pipeline, binds components, then hands the context to each stage in order.

indx ./docs --out ./ai-ready --config indx.toml
┌─────────────────────────┐
│ DirectoryPipeline │
└─────────────────────────┘
┌───────────────────────────────┴───────────────────────────────┐
│ SpaceContext (one object, threaded through) │
│ ctx.root · ctx.seed · ctx.parsed (doc_id → ParsedDoc) · │
│ ctx.errors · ctx.space (KnowledgeSpace: .documents_ · .chunks │
│ · .relations · .manifest) │
└───────────────────────────────┬───────────────────────────────┘
│ run(ctx) -> ctx (same object)
01 Walk 02 Parse 03 Chunk 04 Relate 05 Enrich 06 Embed+Pack
┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌──────────────┐
│ folder │ dir │ Parser │ Parsed │ split │ Chunk │ resolve│ Rel- │ LLM/VLM│ meta │ Embedder→Store│
│ /ZIP │ graph │ .parse │ Doc │ +line │ list │ refs/ │ ations│ topics │ data │ + OutputWriter│
│ walk │────────▶│ │───────▶│ -age │───────▶│ dups │──────▶│ tags │──────▶│ seal .indx │
└────────┘ └────────┘ └────────┘ └────────┘ └────────┘ └──────┬───────┘
detect types text/table/ │
layout/image ▼
┌───────────────────┐
│ KnowledgeSpace │
│ + ./ai-ready/ │
│ handbook.indx │
│ index.json │
│ chunks/ │
│ embeddings/ │
└───────────────────┘

A fresh DirectoryPipeline registers the six built-in stages in canonical order, binds the resolved components onto the context, then runs each stage. Every stage obeys the same contract — run(ctx: SpaceContext) -> SpaceContext — returning the same object it received, mutated in place.

This gives stages a uniform, append-only view of accumulated work:

  • Walk adds the directory graph.
  • Parse adds ParsedDocs.
  • Chunk adds Chunks.
  • Relate adds Relations.
  • Enrich adds metadata.
  • Embed+Pack adds vectors.

See the Stage protocol for the exact contract.

| Stage | Reads from ctx | Writes to ctx | |-------|------------------|-----------------| | 01 Walk | root | space.documents_ (one Document per file, with folder lineage + detected type) | | 02 Parse | space.documents_ | parsed (doc_id → ParsedDoc) | | 03 Chunk | space.documents_, parsed | space.chunks (with lineage + neighbor links) | | 04 Relate | space.documents_, parsed | space.relations, plus per-object references | | 05 Enrich | space.documents_, parsed | enriched metadata on each Document (type, topics, tags, summaries) | | 06 Embed+Pack | space.chunks | embedding on each Chunk + space.manifest, the store, and the sealed .indx archive |

The pipeline is a list you control. Swap the component a stage drives, insert a custom stage, replace a stage by name, or drop a stage entirely.

  • drop("enrich") skips all LLM/VLM work. Use it when no model is available, or when you don’t want enrichment. This is a common, supported operation.
  • drop("embed-pack") (or the --no-embed flag) produces a graph-only space with no vectors.
  • Custom stages — anything that satisfies the Stage protocol and returns the same context can be inserted, such as a PII-redaction pass after Chunk.

The example below swaps the parser and drops enrichment, then runs the pipeline.

from indx import DirectoryPipeline
# requires: pip install "indx[local]" (docling + bge-m3 + qdrant)
pipeline = (
DirectoryPipeline(embedder="bge-m3", store="qdrant")
.use(parser="docling") # swap a component
.drop("enrich") # skip LLM enrichment entirely
)
space = pipeline.run("./notes", "./out")