1:1 mentoring with Big Tech AI engineers
RAG & MCP

Feeds, Changes & Deletes

How documents reach the index and how it learns one changed or died: delivery channels, SFTP failure modes, a document identity that survives an update, tombstones, snapshot vs delta feeds, and a restart-safe batch loop.

Last updated

Production10 min readFirst readDocument ProcessingMetadata & Filters

After this section you can

  • Choose how each channel reports changes and deletes, and say what its failure looks like
  • Give each document an identity that survives an update, so it replaces its chunks and a delete retires them
  • Make an ingest loop safe to kill and restart, and prove a change and a delete reached the answers
07

Keeping the Index True: Feeds, Changes & Deletes

How documents reach the index, and how it learns that one changed or died.

Key idea

An index stays true only if it learns what changed and what died. Give each document an identity that survives an update (not its hash), replace all of a changed document’s chunks, delete on an explicit record or a verified full snapshot, and make every run safe to restart. This is gate 1: in the index, whole and current.

The ledger row decides: skip, patch, replace or retire
1 Arrives file, event or row tenant from the credential Compare hash and ACL hash with the ledger row Ledger (tenant, doc_id) → hash, ACL hash Skip same hash, same ACL: no work Patch ACL same hash, new ACL Replace all chunks new hash, or no row yet Tombstone delete record or snapshot gap revokes now Staged batch queries cannot see it Live index whole and current one switch

Content reaches queries only at the switch; a revocation goes live at once (swipe on a phone).

Related

More in RAG & MCP

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium