Restores the work from #1649, which was reverted on main in #1650 to keep
it out of the last release. Content is identical to that merge.
This is a revert of the revert (a8292f5dc) rather than a merge of
kb_git_mvp, deliberately. Merging the branch would make its commits an
ancestor of both dev and main; main's side deleted those files, so the
next merge between the two branches would silently delete them again. A
revert carries the content without the history, so dev and main share no
ancestor that knows about these files, and the eventual dev -> main merge
sees them as added on one side only and keeps them.
Git becomes the source of truth for knowledge base content; Postgres and
pgvector become a derived, rebuildable index. Both switches guarding the
new path default to off: the KNOWLEDGE_STORE_ENABLED env var and the
per-workspace knowledge_store_enabled column, so merging this changes no
runtime behaviour.
Migrations 175 and 176 only add columns and use ADD COLUMN IF NOT EXISTS.
They are already applied on production, where alembic_version was moved
back to 174 during the revert, so they will re-run harmlessly.
Verified on this branch: 0 conflicts against dev, every app.* import in
the restored files resolves, 2251 unit tests pass. The one collection
error (platforms/google_maps) is missing a fixture that is untracked on
both dev and main, and predates this change.
9.4 KiB
ADR 0001: Git-native Knowledge Base (Git as source of truth, Postgres as derived index)
- Status: Proposed (brainstorm outcome — for team review)
- Date: 2026-07-24
- Origin: Rohan Verma's meeting proposal to pivot from the custom-built KB "file system" to a Git-based system due to persistent maintenance issues; Thierry Bakera to investigate.
Context
What we have today
SurfSense does not actually have a file system. It has a virtual filesystem façade mapped onto Postgres rows, used by the chat agent. The moving parts:
- Virtual
/documents/namespace computed from DB rows —surfsense_backend/app/agents/chat/runtime/path_resolver.py - Read-side backend faking
ls/read/glob/grepover Postgres —.../filesystem/backends/kb_postgres.py - Write-side "commit at end of turn" layer —
.../main_agent/middleware/kb_persistence/middleware.py - Three separate hand-rolled versioning/audit systems:
DocumentVersion(user history) —app/utils/document_versioning.pyDocumentRevision/FolderRevision(agent revert snapshots) —app/services/revert_service.pyAgentActionLog(tool-call audit)
- Supporting machinery: fractional indexing for ordering, move tracking,
content_hashchange detection, chunk reconciliation.
The problem
The team has been hand-implementing — on top of a relational DB never designed for it — the exact primitives Git provides natively. That re-implementation is the source of the "persistent maintenance issues."
| Hand-rolled today | Git provides natively |
|---|---|
path_resolver + folder tree |
tree objects |
| end-of-turn staged commit | atomic commits |
DocumentVersion snapshots |
commit history |
DocumentRevision + revert_service |
git revert / reset |
AgentActionLog |
commit log / blame |
content_hash dedup |
content-addressed blobs (SHA) |
| fractional indexing / move tracking | tree diff / rename detection |
Search stack (unchanged by this ADR)
Hybrid chunk search: pgvector (HNSW) + Postgres FTS + RRF, optional reranking. Chunking via Chonkie; incremental via chunk_reconciler.py. See .../shared/retrieval/hybrid_search.py.
Decision
Adopt Git as the single source of truth for all indexed KB content. Postgres becomes a derived, rebuildable index holding only chunks + embeddings.
Core model
agent notes ─┐
editor saves ─┤
uploads ──────┤→ Git commit (source of truth) → indexer → Postgres (chunks + embeddings)
Notion ───────┤
Drive ────────┘
(indexable connectors only)
Slack / Gmail (live connectors) ──→ queried at chat time, bypass storage entirely
- Git = truth for everything that gets stored/indexed (agent/editor notes, uploads, and indexable connectors like Notion, Drive, Obsidian — the
is_indexableconnectors). - Postgres = derived index only (chunks + embeddings). It is a cache: it can be wiped and rebuilt from Git at any time via a single
reindex(workspace)function. - Live connectors (Slack, Gmail) are never stored or indexed — they are queried live at chat time and are entirely out of scope for this design.
- Binary blobs (original PDF/DOCX) stay in the existing local/Azure blob store (or Git-LFS later); Git holds the extracted markdown, not raw binaries.
What changes for the agent
The agent's tools are unchanged (ls, read, write, edit, mv, rm). Only the backend behind them changes:
| Agent action | Backed by |
|---|---|
File ops (ls/read/write/edit/mv/rm) |
Git working tree (real files) |
| Semantic search | Postgres (derived chunk/embedding index) |
- Before:
KBPostgresBackendfakes files over Postgres rows. - After: a Git-working-tree backend operates on real files.
path_resolverlargely disappears (paths are real). - Write flow: agent edits → git working tree → one commit at end of turn (replaces the
kb_persistencecommit-to-Postgres step) → indexer refreshes Postgres chunks.
What we delete
- The virtual-FS façade (
path_resolver,kb_postgresstaging). DocumentVersion,DocumentRevision/FolderRevision,revert_service→ replaced by git history +git revert.- Fractional indexing / bespoke move tracking → git tree operations.
This pivot is mostly deletion, which is the point.
We are borrowing from authoritative sources, not inventing
Every decision traces to a proven, battle-tested reference. The only SurfSense-specific work is the adaptation glue.
| Decision | Borrowed from |
|---|---|
| Git = truth, Postgres = rebuildable cache | Fossil SCM — canonical artifacts + SQL tables as a pure cache, recomputed via fossil rebuild (production since 2007). https://fossil-scm.org/home/doc/trunk/www/theory1.wiki |
| Content in git, metadata/index in a DB | Gollum (GitHub/GitLab wikis): "storage abstraction layer... only some data in the DB". https://github.com/gollum/gollum · https://docs.gitlab.com/17.5/development/wikis/ |
| Silent commit-per-save, hide git from users | kherad. https://github.com/mohammadmaso/kherad |
| Embeddings keyed by blob SHA, incremental (not rebuild) | Coregit LLM Wiki (https://coregit.dev/blog/llm-wiki-launch) + vector-index-as-cache best practices (LangChain RecordManager, LlamaIndex docstore). Cache key = (model_version, content_hash). |
| Python git engine | dulwich — pure Python, deploy-friendly, real wire protocol. https://github.com/jelmer/dulwich |
| "Don't put a firehose in git" / git limits | "Git is not a database" critiques — degrades past ~500k–1M files, no query engine, no concurrency control. |
Future content model (raw/+wiki/, lint, contradiction-flagging) |
Karpathy's LLM Wiki. https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f |
| Future real-time collab | Yjs (Notion, Linear). https://github.com/yjs/yjs |
| Future fact-level temporal memory | Zep / Graphiti (bi-temporal, "invalidate don't delete"). https://arxiv.org/html/2501.13956 |
The counter-example we explicitly reject
Wiki.js git module = DB is truth, git is a two-way mirror. This creates two sources of truth and a reconciliation cursor; it has a known class of silent-sync bugs (see requarks/wiki discussion #7860). We use one-way derivation (git → Postgres), never two-way sync.
Scope: v1 (keep it simple)
Ship the smallest thing that removes the maintenance pain:
- Git repo per workspace (dulwich) holding indexed content as markdown.
- One commit per agent turn / editor save. No branches, no merge, no review workflow.
- Delete the three versioning systems; history/undo = git log +
git revert. - pgvector stays, rebuilt from git, keyed by blob SHA (point existing
chunk_reconcilerat blob SHA so unchanged files skip re-embedding). - Per-workspace lock/queue around commits (git is single-writer; this is a data-integrity boundary, not a feature — non-negotiable).
- One
reindex(workspace)function that wipes and rebuilds Postgres chunks from the git repo (the Fossilrebuilddiscipline — makes the system safe to ship early).
Deferred (v2+, explicitly out of v1)
- Connect-your-own-remote (GitHub/GitLab) — free later because the repo is real git.
- CRDT / Yjs real-time collaboration.
- Review / merge workflows (kherad's reviewer layer).
- Graphiti / bi-temporal fact graph.
- Karpathy
raw/+wiki/content model, contradiction-flagging, lint.
Consequences
Positive
- Large net deletion of bespoke code (the maintenance win).
- Storage and search decouple → the team can improve semantic search independently (Rohan's stated goal).
- Postgres becomes disposable/rebuildable → simpler recovery, fewer consistency bugs.
- Unlocks future "user owns their KB as a real git repo" differentiator.
Negative / risks
- Concurrency: git is single-writer per repo → requires the per-workspace lock (mitigated in v1).
- Repo hygiene: many small commits → periodic
git gc/repack (operational, manageable). - Migration: existing Postgres KBs must be exported into git repos once, preserving
unique_identifier_hashmapping. - Real-time UI (Zero): currently driven by Postgres logical replication; still needs a git → Postgres projection to keep the web client live. This is new code that partially offsets deletions.
Open questions (for team discussion)
- Zero / real-time UI: confirm the git → Postgres projection path and whether Zero stays as-is.
- Binaries: keep blob store vs. adopt Git-LFS.
- Migration cutover: big-bang vs. per-workspace feature flag (recommend feature-flag rollout).
- Merge UX later: CRDT (Yjs) vs. review-gate (kherad) when multi-writer becomes a requirement.
Appendix: key file index (current implementation)
| Topic | Path |
|---|---|
| Virtual path resolver | surfsense_backend/app/agents/chat/runtime/path_resolver.py |
| Virtual FS read backend | .../filesystem/backends/kb_postgres.py |
| Virtual FS write commit | .../main_agent/middleware/kb_persistence/middleware.py |
| Hybrid search | .../shared/retrieval/hybrid_search.py |
| Chunk reconciliation | surfsense_backend/app/indexing_pipeline/chunk_reconciler.py |
| User version history | surfsense_backend/app/utils/document_versioning.py |
| Agent revert | surfsense_backend/app/services/revert_service.py |
| ORM models | surfsense_backend/app/db.py |