SurfSense/docs/adr/0001-git-native-knowledge-base.md
CREDO23 0b34db9d93 feat(knowledge-store): git-native knowledge base
Restores the work from #1649, which was reverted on main in #1650 to keep
it out of the last release. Content is identical to that merge.

This is a revert of the revert (a8292f5dc) rather than a merge of
kb_git_mvp, deliberately. Merging the branch would make its commits an
ancestor of both dev and main; main's side deleted those files, so the
next merge between the two branches would silently delete them again. A
revert carries the content without the history, so dev and main share no
ancestor that knows about these files, and the eventual dev -> main merge
sees them as added on one side only and keeps them.

Git becomes the source of truth for knowledge base content; Postgres and
pgvector become a derived, rebuildable index. Both switches guarding the
new path default to off: the KNOWLEDGE_STORE_ENABLED env var and the
per-workspace knowledge_store_enabled column, so merging this changes no
runtime behaviour.

Migrations 175 and 176 only add columns and use ADD COLUMN IF NOT EXISTS.
They are already applied on production, where alembic_version was moved
back to 174 during the revert, so they will re-run harmlessly.

Verified on this branch: 0 conflicts against dev, every app.* import in
the restored files resolves, 2251 unit tests pass. The one collection
error (platforms/google_maps) is missing a fixture that is untracked on
both dev and main, and predates this change.
2026-07-31 02:33:38 +02:00

9.4 KiB
Raw Permalink Blame History

ADR 0001: Git-native Knowledge Base (Git as source of truth, Postgres as derived index)

  • Status: Proposed (brainstorm outcome — for team review)
  • Date: 2026-07-24
  • Origin: Rohan Verma's meeting proposal to pivot from the custom-built KB "file system" to a Git-based system due to persistent maintenance issues; Thierry Bakera to investigate.

Context

What we have today

SurfSense does not actually have a file system. It has a virtual filesystem façade mapped onto Postgres rows, used by the chat agent. The moving parts:

  • Virtual /documents/ namespace computed from DB rows — surfsense_backend/app/agents/chat/runtime/path_resolver.py
  • Read-side backend faking ls/read/glob/grep over Postgres — .../filesystem/backends/kb_postgres.py
  • Write-side "commit at end of turn" layer — .../main_agent/middleware/kb_persistence/middleware.py
  • Three separate hand-rolled versioning/audit systems:
    • DocumentVersion (user history) — app/utils/document_versioning.py
    • DocumentRevision / FolderRevision (agent revert snapshots) — app/services/revert_service.py
    • AgentActionLog (tool-call audit)
  • Supporting machinery: fractional indexing for ordering, move tracking, content_hash change detection, chunk reconciliation.

The problem

The team has been hand-implementing — on top of a relational DB never designed for it — the exact primitives Git provides natively. That re-implementation is the source of the "persistent maintenance issues."

Hand-rolled today Git provides natively
path_resolver + folder tree tree objects
end-of-turn staged commit atomic commits
DocumentVersion snapshots commit history
DocumentRevision + revert_service git revert / reset
AgentActionLog commit log / blame
content_hash dedup content-addressed blobs (SHA)
fractional indexing / move tracking tree diff / rename detection

Search stack (unchanged by this ADR)

Hybrid chunk search: pgvector (HNSW) + Postgres FTS + RRF, optional reranking. Chunking via Chonkie; incremental via chunk_reconciler.py. See .../shared/retrieval/hybrid_search.py.


Decision

Adopt Git as the single source of truth for all indexed KB content. Postgres becomes a derived, rebuildable index holding only chunks + embeddings.

Core model

agent notes ─┐
editor saves ─┤
uploads ──────┤→  Git commit (source of truth)  →  indexer  →  Postgres (chunks + embeddings)
Notion ───────┤
Drive ────────┘
(indexable connectors only)

Slack / Gmail (live connectors) ──→ queried at chat time, bypass storage entirely
  • Git = truth for everything that gets stored/indexed (agent/editor notes, uploads, and indexable connectors like Notion, Drive, Obsidian — the is_indexable connectors).
  • Postgres = derived index only (chunks + embeddings). It is a cache: it can be wiped and rebuilt from Git at any time via a single reindex(workspace) function.
  • Live connectors (Slack, Gmail) are never stored or indexed — they are queried live at chat time and are entirely out of scope for this design.
  • Binary blobs (original PDF/DOCX) stay in the existing local/Azure blob store (or Git-LFS later); Git holds the extracted markdown, not raw binaries.

What changes for the agent

The agent's tools are unchanged (ls, read, write, edit, mv, rm). Only the backend behind them changes:

Agent action Backed by
File ops (ls/read/write/edit/mv/rm) Git working tree (real files)
Semantic search Postgres (derived chunk/embedding index)
  • Before: KBPostgresBackend fakes files over Postgres rows.
  • After: a Git-working-tree backend operates on real files. path_resolver largely disappears (paths are real).
  • Write flow: agent edits → git working tree → one commit at end of turn (replaces the kb_persistence commit-to-Postgres step) → indexer refreshes Postgres chunks.

What we delete

  • The virtual-FS façade (path_resolver, kb_postgres staging).
  • DocumentVersion, DocumentRevision / FolderRevision, revert_service → replaced by git history + git revert.
  • Fractional indexing / bespoke move tracking → git tree operations.

This pivot is mostly deletion, which is the point.


We are borrowing from authoritative sources, not inventing

Every decision traces to a proven, battle-tested reference. The only SurfSense-specific work is the adaptation glue.

Decision Borrowed from
Git = truth, Postgres = rebuildable cache Fossil SCM — canonical artifacts + SQL tables as a pure cache, recomputed via fossil rebuild (production since 2007). https://fossil-scm.org/home/doc/trunk/www/theory1.wiki
Content in git, metadata/index in a DB Gollum (GitHub/GitLab wikis): "storage abstraction layer... only some data in the DB". https://github.com/gollum/gollum · https://docs.gitlab.com/17.5/development/wikis/
Silent commit-per-save, hide git from users kherad. https://github.com/mohammadmaso/kherad
Embeddings keyed by blob SHA, incremental (not rebuild) Coregit LLM Wiki (https://coregit.dev/blog/llm-wiki-launch) + vector-index-as-cache best practices (LangChain RecordManager, LlamaIndex docstore). Cache key = (model_version, content_hash).
Python git engine dulwich — pure Python, deploy-friendly, real wire protocol. https://github.com/jelmer/dulwich
"Don't put a firehose in git" / git limits "Git is not a database" critiques — degrades past ~500k1M files, no query engine, no concurrency control.
Future content model (raw/+wiki/, lint, contradiction-flagging) Karpathy's LLM Wiki. https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
Future real-time collab Yjs (Notion, Linear). https://github.com/yjs/yjs
Future fact-level temporal memory Zep / Graphiti (bi-temporal, "invalidate don't delete"). https://arxiv.org/html/2501.13956

The counter-example we explicitly reject

Wiki.js git module = DB is truth, git is a two-way mirror. This creates two sources of truth and a reconciliation cursor; it has a known class of silent-sync bugs (see requarks/wiki discussion #7860). We use one-way derivation (git → Postgres), never two-way sync.


Scope: v1 (keep it simple)

Ship the smallest thing that removes the maintenance pain:

  1. Git repo per workspace (dulwich) holding indexed content as markdown.
  2. One commit per agent turn / editor save. No branches, no merge, no review workflow.
  3. Delete the three versioning systems; history/undo = git log + git revert.
  4. pgvector stays, rebuilt from git, keyed by blob SHA (point existing chunk_reconciler at blob SHA so unchanged files skip re-embedding).
  5. Per-workspace lock/queue around commits (git is single-writer; this is a data-integrity boundary, not a feature — non-negotiable).
  6. One reindex(workspace) function that wipes and rebuilds Postgres chunks from the git repo (the Fossil rebuild discipline — makes the system safe to ship early).

Deferred (v2+, explicitly out of v1)

  • Connect-your-own-remote (GitHub/GitLab) — free later because the repo is real git.
  • CRDT / Yjs real-time collaboration.
  • Review / merge workflows (kherad's reviewer layer).
  • Graphiti / bi-temporal fact graph.
  • Karpathy raw/ + wiki/ content model, contradiction-flagging, lint.

Consequences

Positive

  • Large net deletion of bespoke code (the maintenance win).
  • Storage and search decouple → the team can improve semantic search independently (Rohan's stated goal).
  • Postgres becomes disposable/rebuildable → simpler recovery, fewer consistency bugs.
  • Unlocks future "user owns their KB as a real git repo" differentiator.

Negative / risks

  • Concurrency: git is single-writer per repo → requires the per-workspace lock (mitigated in v1).
  • Repo hygiene: many small commits → periodic git gc/repack (operational, manageable).
  • Migration: existing Postgres KBs must be exported into git repos once, preserving unique_identifier_hash mapping.
  • Real-time UI (Zero): currently driven by Postgres logical replication; still needs a git → Postgres projection to keep the web client live. This is new code that partially offsets deletions.

Open questions (for team discussion)

  1. Zero / real-time UI: confirm the git → Postgres projection path and whether Zero stays as-is.
  2. Binaries: keep blob store vs. adopt Git-LFS.
  3. Migration cutover: big-bang vs. per-workspace feature flag (recommend feature-flag rollout).
  4. Merge UX later: CRDT (Yjs) vs. review-gate (kherad) when multi-writer becomes a requirement.

Appendix: key file index (current implementation)

Topic Path
Virtual path resolver surfsense_backend/app/agents/chat/runtime/path_resolver.py
Virtual FS read backend .../filesystem/backends/kb_postgres.py
Virtual FS write commit .../main_agent/middleware/kb_persistence/middleware.py
Hybrid search .../shared/retrieval/hybrid_search.py
Chunk reconciliation surfsense_backend/app/indexing_pipeline/chunk_reconciler.py
User version history surfsense_backend/app/utils/document_versioning.py
Agent revert surfsense_backend/app/services/revert_service.py
ORM models surfsense_backend/app/db.py