kimi-code/packages/minidb/README.md
Haozhe 119a33f7f1
Some checks are pending
CI / test (2) (push) Waiting to run
CI / test (3) (push) Waiting to run
CI / test (4) (push) Waiting to run
CI / test (5) (push) Waiting to run
CI / test-pi-tui (push) Waiting to run
CI / test-windows (push) Waiting to run
CI / lint (push) Waiting to run
CI / build (push) Waiting to run
CI / test (1) (push) Waiting to run
CI / typecheck (push) Waiting to run
Nix Build / Check flake.nix workspace sync (push) Waiting to run
Nix Build / nix build .#kimi-code (push) Blocked by required conditions
Release / Release (push) Waiting to run
Release / Deploy docs (push) Blocked by required conditions
Release / Native release artifact (push) Blocked by required conditions
Release / Publish native release assets (push) Blocked by required conditions
feat(minidb): persistent index generations and lifecycle hardening (#2604)
* perf(minidb): skip idle everysec fsyncs and add lifecycle stats

- everysec WAL now fsyncs on the timer only while dirty (tracked by a
  write/sync generation watermark); close() keeps its unconditional
  final sync, and background sync failures surface via walFsyncErrors
  plus a sticky lastWalFsyncError instead of being silently swallowed
- add WAL queue/group-commit counters (walQueuedBytes,
  walMaxQueuedBytes, walGroupCommits, walGroupCommitFrames) and
  lifecycle phase stats: recovery bytes/frames/duration, index/text
  rebuild durations, compaction total/snapshot/rotation/postings
  durations, rotation pause, and query candidates/decoded/sorted
  rows; add a syncIntervalMs open option threaded through compaction
  WAL rotation
- rewrite the bench on fixed-seed synthetic data with a stable
  machine-readable JSON report (cold open 10k/50k/100k, word/ngram
  search, idle-fsync acceptance, 100k compaction, event-loop delay,
  peak heap/RSS per scenario) and pin the schema in
  test/bench-json.test.ts; add app-side baselines with loose
  complexity budgets in sessionIndex and searchService tests
- fix ClusterDb lock-pool closeAll() leaking in-flight shard opens
  and drain the query store's async close on server shutdown,
  eliminating the ENOTEMPTY directory-teardown race

* perf(minidb): bound startup rebuild and steady-state hot paths

- rebuild all derived indexes in one shared store walk: a single decode
  per record fans out to staged builders, dt rebuild reads record
  metadata only, and index-less opens no longer decode at all
- rank full-text results with a bounded min-heap plus a stable key
  tie-break instead of sorting every candidate
- remove/overwrite text docs via a docID -> delta-terms reverse map
  instead of scanning the whole delta vocabulary
- validate unique batches incrementally against touched postings
  instead of copying the full per-index owner map
- reap due TTL entries from the expiry heap on the write path instead
  of a full-store sweep per write

* feat(session-index): add minidb read model with keyset pagination

- add ISessionIndex read-model lifecycle (prepare/status, ready/degraded
  states) behind the persistence_minidb_readmodel experimental flag
- add ISessionIndexMirror write side recording fresh summaries into a
  bounded, coalescing queue after the authoritative document is durable
- replace the offset cursor with before/after keyset pagination; rename
  list/countActive to listRecent/count
- extend IQueryStore with ordered columns and pageByColumn, plus
  getMany/listKeys/dropCollection
- wire the read model through kap-server routes and start, and update the
  klient sessions contract
- index every session for global search instead of the 500 most recent

* feat(kap-server): bound search sync lifecycle, pagination, and query budgets

- split search requests from sync work: searchIndex() no longer awaits
  runSync/reopen/reindex; a single-flight sync coordinator with debounce
  and backpressure runs in the background, stale generations keep
  serving with explicit stale/degraded state, and refresh/sync/reindex
  failures surface via lastRefreshError instead of being swallowed
- scope file-meta keys by session id (\0meta\file\<sessionId>\<hash>)
  with lazy + one-shot background migration from the legacy hash-only
  keys, so one session sync only touches its own meta rows
- make authoritative scans incremental (mtime/ino/size rescan
  conditions, unchanged files no longer rewrite meta) and read wire
  deltas in 1 MiB chunks instead of whole-file buffer + split
- replace offset pagination with versioned v2 keyset page tokens
  (fingerprint + index generation + sort boundary); generation changes
  fail old tokens with invalid_page_token, legacy v1 offset tokens are
  served once and upgraded, and pages collect via bounded top-K instead
  of full sort + offset skip
- add query budgets enforced at the postings/score stage: max query
  terms, literal length cap, postings visit budget (minidb
  searchBounded/maxVisits with prefix decoding that never fabricates
  hits and skips the postings LRU), candidate caps, deadline and text
  budget; truncation is reported via incomplete reasons
  candidate_cap/postings_budget/deadline
- reopen read-only dbs by opening the next handle before closing the
  previous one so a failed refresh keeps the old generation serving;
  failed opens now self-heal through search traffic

100k-message bench: first page p95 < 300ms and page-100 cost on par
with page 1; event-loop delay during queries stays sub-millisecond.

* fix(minidb): poison, roll back, and recover the WAL on write failures

- give WAL writes a commit point: a failed flushBatch poisons the WAL
  (WAL_POISONED, tracked separately as walWriteErrors vs walFsyncErrors),
  rejects queued frames in reverse enqueue order, and stops scheduling
  further batches; everysec background sync failures stay non-rejecting
  per stage-1 semantics
- recover in place to a known-safe point: a serialized recovery chain
  truncates the WAL back to the first un-acked frame, rebuilds
  size/nextOffset, and clears the poison; writes queue behind the
  recovery gate (zero-cost when idle), a failed truncate flips the
  instance into an explicit writeDisabled state, and a stale truncate
  offset (WAL file replaced by a rotation) skips the truncate
- roll failed flush groups back as a unit: frames are stamped with
  their batchId, MiniDb keeps per-group earliest pre-state, and the
  first rejection restores every key of the group (rejected writes no
  longer reappear after reopen, and in-memory state matches reopen for
  any failure interleaving); the per-op seq guard remains for
  cross-group and rotation-retry races
- wrap applyOp and the following in-memory mutations so a contract
  violation poisons the WAL and rolls the group back instead of
  escaping as a half-commit; frames never enqueued (seal race) roll
  back per-op without poisoning
- tag errors past the commit point with ambiguous: true so callers can
  distinguish "definitely not applied" from "maybe applied but revoked"
- close() waits for the recovery chain to go idle and backup() fences
  behind in-flight recovery before copying files

Controlled A/B bench (22 alternating iterations, 100k concurrent sets):
write-path throughput regression is within the 2% budget.

* fix(minidb): turn the file lock into an instance-owned serialized lease

- distinguish lock ownership by instance instead of pid: every acquire
  mints a pid:uuid token carried by lock/bid/watch files, inspect().mine
  compares tokens, liveness still follows pid, tokenless legacy files
  keep the old stale-takeover path, and hasLiveForeignWatch excludes
  self by token so same-process contenders see each other (closing the
  double-win takeover and the cross-instance release); a live same-pid
  lock is still respected, and re-acquiring a held lock is idempotent
- serialize acquire/renew/release through a per-instance promise-chain
  mutex: renew re-checks held inside the chain and release waits for an
  in-flight renew, eliminating the renew/rename-after-unlink ghost lock
- make MiniDb.close() a state machine (open/closing/closed) with a
  shared closePromise: cleanup runs per-resource try/catch in
  dependency order (text indexes, store, valueReader, WAL, lock),
  aggregates every cleanup error into an AggregateError, stays in
  'closing' on failure so a retry finishes the cleanup, and no longer
  leaks the lock when the WAL close fails; a rejected in-flight
  compaction no longer escapes the cleanup pass

* fix(minidb): keep readers on one consistent file generation

- add an internal persistent-files module as the single source of truth
  for the persisted file set (snapshot, WAL, sidecars, postings
  pattern, fingerprint subset); lock-pool fingerprints, persistentFiles,
  open stale-tmp cleanup, and backup/restore filtering all derive from
  it, and fingerprints upgrade to dev:ino:size:mtimeMs so compound
  sidecar changes can no longer hide from cluster readers
- pair snapshot and WAL generations during recovery (transitional
  stat-pairing until stage-5 manifests): each pass anchors the fds it
  scans, re-stats afterwards, tolerates append-only WAL growth, retries
  bounded times on generation churn with a clean store reset, and
  throws RECOVERY_GENERATION_CHURN when churn exceeds the budget; the
  disk-mode ValueReader attach re-validates inodes so stale offsets
  never read a replaced file
- make the rotation directory fsyncs strict: failures abort the
  rotation through the existing rollback path instead of being
  swallowed, while platforms without directory fsync degrade once with
  a warn and stats.dirFsyncUnsupported

* fix(minidb): serialize index-definition sidecar mutations, persist before publish

- extract the promise-chain mutex into a shared createSerializer() and
  give each sidecar family (secondary/compound/text) its own chain:
  create/drop run uninterruptibly (memory change + rebuild + persist),
  different families stay independent, and the data write path never
  shares these chains
- reverse the publication order to staged -> persist -> publish: a
  create stages the definition, rebuilds via the staged builder,
  persists the sidecar including the new definition, then publishes
  atomically; any failure discards the staged state leaving live and
  sidecar untouched (no phantom indexes, retry-safe); a drop persists
  the sidecar without the definition before removing it live; text
  index create/drop adopt the same pattern, replacing the hand-rolled
  unwind, and a dropping marker keeps compaction postings rebuilds out
  of the persist window
- feed staged indexes from the incremental write path (add/remove/
  checkUnique/checkUniqueBatch visit live+staged) so writes landing in
  the persist window are not lost at publish; queries still see live
  only
- harden writeFileAtomic: instance-unique tmp names (.tmp-pid-seq),
  a strict fsyncDir after rename so a successful persist is crash
  durable, and whitelist-based stale-tmp cleanup that never touches
  lock tmp files

* fix(minidb): validate writes before any side effect, canonicalize values once

- canonical value at the write boundary: the json codec re-parses the
  encoded bytes once and every downstream consumer (unique checks,
  secondary/compound/text indexes, dt extraction) sees exactly the
  persisted representation, so getter/toJSON/Proxy documents can no
  longer diverge between the index view and the storage view
- reorder the set/batch pipeline so every fallible check happens before
  any visible side effect: prepare (key/ttl checks, encoding, canonical
  decode, index field extraction, tokenization) -> unique checks ->
  ensureMemoryFor eviction -> commit; a constraint failure now leaves
  the database untouched (no more evicted victims on rejected inserts),
  and applyOp is structurally pure against pre-validated data
- tokenize at the prepare boundary: TextIndex gains prepareAdd/
  addPrepared and the buildQueue carries validated key+tokens mutations
  instead of raw docs, so a throwing custom tokenizer can no longer
  poison the live view or the queue, and custom-tokenizer output is
  rejected per token over 0xffff bytes before it can permanently break
  postings rebuilds; prepared tokens are keyed by index instance so a
  same-name drop+create mid-write re-tokenizes instead of crossing
  tokenizers
- strict batch structure validation: scanBatchOpRefs/decodeBatchOps
  reject unknown op types, out-of-bounds lengths, and trailing bytes
  (offset must equal body length), so a valid-CRC but malformed batch is
  skipped as a unit and counted via RecoveryInfo.corruptBatches instead
  of being partially applied

Bench vs the stage-1 baseline: json write throughput regression is
within the 5% budget (median ~2-4% depending on the measurement).

* feat(minidb): add OpTracker drain primitive and atomic backup, harden tests

- introduce the internal OpTracker (close gate + in-flight counter with
  enter/leave/close/whenIdle and reference-counted pause/resume) and
  drive every shutdown/drain path from it: WAL background syncs are
  tracked so close() waits out an in-flight sync before closing the fd,
  cluster lock-pool closeAll() closes the gates and drains busy
  callbacks before closing handles, and MiniDb writes pass a write gate
- make backup() atomic with a defined linearization point: pause the
  write gate, drain in-flight writes (every acknowledged write is now
  included), copy to a sibling temp dir with per-file fsyncs, write the
  manifest last as the commit marker, and rename into place; failures
  clean up and leave no partial backup, and concurrent writes are
  rejected with BACKUP_IN_PROGRESS
- reap emptied compound-index groups on remove (the groups map no
  longer grows monotonically), move the open-time mkdir behind the
  readOnly check so a read-only open of a missing directory fails with
  ENOENT instead of creating it, and never run a destructive rebuild
  for a read-only open failure (explicit or onLockFail fallback)
- consolidate every review fault-injection repro into the formal suite
  behind deterministic barrier helpers (programmable writev/sync/
  rename/tokenize hooks) and convert the six timing-based tests to
  barrier/tick-driven assertions; the .tmp repro scripts are removed

The converted timing tests and the full suite pass 50 repeat runs
(including under CPU load injection) with zero flakes.

* feat(minidb): persist derived indexes as atomic generations, open from WAL delta

- checkpoint the store, dt/secondary/compound indexes, and text
  dictionary/postings/docs into immutable generations under
  generations/g-NNNNNN published atomically (tmp build, per-file
  checksums and fsyncs, dir rename, CURRENT swap, strict dir fsyncs);
  the manifest records the format version, WAL/snapshot checkpoint
  anchors, per-index definition hashes, and codec/value-mode
  compatibility
- open now loads the published generation and replays only the WAL
  delta after its checkpoint: no full value decode, corpus
  tokenization, or postings rewrite on a normal reopen (warm opens are
  3.5-13.8x faster at 100k/1M records); a definition change rebuilds
  only the affected index, and corrupt generation files fall back to
  the previous generation or the legacy full recovery without ever
  touching the authoritative snapshot/WAL
- build generations transactionally with compaction (rotation plus
  derived state publish as one unit, replacing the synchronous
  rebuildTextPostings tail), capture concurrent writes through a sealed
  op queue with byte/op caps, hard-link clean postings and the snapshot
  into the new generation, and repoint every live text base into the
  CURRENT generation after publish
- cluster/read-only refresh watches CURRENT and the WAL watermark:
  pure generation publishes keep readers on incremental catch-up while
  rotations reopen onto the new generation; writers building the next
  generation never disturb readers of the current one
- legacy databases open through the old path unchanged and gain their
  first generation in the background; OpenOptions.indexGenerations:
  false fully restores the pre-generation behavior

* feat(minidb): workerize text-index builds and split MiniDb into facets

- split the monolithic src/index.ts into facet modules (mini-db, types,
  value-codec, memory-guard, backup, query-engine, text-registry,
  wal-group, generation-builder/loader, write-path, read-path,
  index-admin, lifecycle, stats) and move text-index.ts to text-index/
- run corpus-scale text-index builds off the main thread via the bounded
  worker engine (src/worker/), exported through the new worker-runtime
  subpath, with inline fallback for small corpora and rollback switches
- defer the open-time fallback text rebuild into a maintenance task;
  searches on a not-yet-committed base raise TextIndexBuildingError
- add the unified maintenance scheduler, bounded async read surface,
  and a maintenance bench
- kap-server search: switch to searchBoundedAsync and serve the
  building page while the index base rebuilds after fallback recovery
- kimi-code: install the SEA-bundled minidb text-build worker at
  startup, bundle it via the native asset scripts, and add the
  startup-trace util plus the KIMI_TUI_INPUT_LATENCY debug probe

* fix(minidb): treat win32 EPERM as unsupported directory fsync

- extract isUnsupportedDirectoryFsyncError and cover win32 EPERM
- drop the one-shot console.warn; stats.dirFsyncUnsupported carries the degraded state

* fix(kap-server): harden search-index dispose and drain lifecycle

- dispose() now closes an OpTracker gate and drains in-flight sync/refresh
  passes before closing the db, so no background write can hit a closed
  handle; the deleteSessionDocs loop and trailing stats write skip once
  the gate closes (review #20)
- drainGlobalSearchDisposals loops to a fixpoint so disposals registered
  while a drain is in flight are also awaited (review #21)
- pin the post-open failure semantics with a regression test: a failed
  text-index setup closes the handle and the next open reacquires the
  writer lock instead of self-locking read-only (review #19)
- export OpTracker from the minidb root for the search service's drain

* chore: fix oxlint type-aware lint errors
2026-08-04 20:38:42 +08:00

24 KiB
Raw Permalink Blame History

minidb

A pure-Node.js (zero native addons) embedded key-value database that mixes ideas from Redis (in-memory KV speed, data structures, AOF-style rewrite) and SQLite (durable single-file persistence, WAL, indexed queries). Written in TypeScript, with strict types and zero runtime dependencies.

Built by studying the real source code of Redis, SQLite, NeDB, Bitcask, and a tiny SQLite clone — see DESIGN_NOTES.md for what each one taught us.

Features

  • O(1) GET/SET from an in-memory index (Map), values as bytes.
  • Optional larger-than-RAM values: valueMode: 'disk' keeps only value pointers in RAM and reads value bulk from the snapshot/WAL with synchronous positioned reads.
  • Durable: append-only WAL with group commit + three fsync policies (always / everysec / no), exactly like Redis AOF.
  • Crash-safe recovery: CRC-framed records; a torn tail from a crash is detected and truncated automatically.
  • Snapshots + compaction (WAL rewrite) bound disk growth and speed up restart.
  • TTL: lazy + active expiration (Redis-style, heap-backed).
  • Secondary indexes: equality and range indexes over JSON documents, backed by a Redis-style skip list with O(log N) rank/range; unique and sparse supported; array fields indexed per element.
  • RESP server (optional): talk to it with redis-cli / ioredis.
  • ClusterDb (optional sharding layer): multi-process concurrent read/write over N sharded MiniDb directories, scaling write throughput with shard count.
  • Value codecs: buffer, string, or json.
  • Zero dependencies, pure node:* built-ins.

Install / layout

Written in TypeScript, zero runtime dependencies. Source is in src/; build output goes to dist/ (generated by npm run build).

src/
  index.ts          public API barrel (re-exports; no logic)
  mini-db.ts        the MiniDb class (open/close, KV reads, index admin, delegates)
  types.ts          public option/record types (+ the internal PreparedOp)
  value-codec.ts    value codecs + key byte-string / atomic sidecar helpers
  memory-guard.ts   maxMemory facet (LRU access set, evict/reject enforce)
  backup.ts         online backup (write-gate fence + atomic copy)
  query-engine.ts   unified query facet (candidate collection, dt fast path)
  text-registry.ts  text-index registry facet (defs, create/drop, search)
  wal-group.ts      WAL group-commit + in-place recovery gating facet
  generation-builder.ts  persistent index generation build facet
  generation-loader.ts   generation load facet (image loads + WAL delta replay)
  write-path.ts     write path facet (set/del/batch/expire commit machinery)
  read-path.ts      read path facet (KV reads, scans, live-record feeds)
  index-admin.ts    secondary/compound index admin facet (+ open-time rebuild)
  lifecycle.ts      open/close/openOrRebuild/renewLock flows
  stats.ts          the stats object factory
  codec.ts          binary frame encode/decode + CRC parser + resync
  crc32.ts          table-based CRC-32
  wal.ts            buffered append + group commit + fsync policies
  store.ts          in-memory index + TTL (lazy + active) + value refs
  value-reader.ts   synchronous positioned reads for disk-backed values
  snapshot.ts       chunked snapshot writer
  recovery.ts       load snapshot + replay WAL + resync/truncate
  compaction.ts     WAL rewrite / rotation
  skiplist.ts       comparator-driven skip list with span
  index-manager.ts  secondary indexes (equality + range)
  dt-index.ts       datetime-column indexes
  query.ts          jq-like value query engine
  text-index/       full-text index (in-RAM dictionary + on-disk postings)
  text-postings.ts  on-disk postings file for the full-text index
  lockfile.ts       exclusive file lock + read-only
  server.ts         optional RESP TCP server
npm run build      # compile src/ -> dist/ (JS + .d.ts)
npm run typecheck  # tsc --noEmit (strict)

Quick start (embedded)

import { MiniDb } from 'minidb';

const db = await MiniDb.open({
  dir: './data',
  valueCodec: 'string',     // 'buffer' | 'string' | 'json'
  fsyncPolicy: 'everysec',  // 'always' | 'everysec' | 'no'
});

await db.set('hello', 'world');
db.get('hello');            // 'world'

await db.set('temp', 'x', { ttl: 5000 }); // expires in 5s
db.ttl('temp');             // remaining ms (-1 none, -2 missing)

await db.del('hello');
await db.close();           // flush + fsync + close

Within this repo, import from ./src/index.ts (run with tsx). As an installed package, import from 'minidb' (the built dist/ output).

Re-open the same dir and your data is recovered from the snapshot + WAL.

Typed usage

MiniDb is generic over the value type:

interface User { name: string; age: number; city: string }
const db = await MiniDb.open<User>({ dir: './data', valueCodec: 'json' });
await db.set('u1', { name: 'Ann', age: 30, city: 'Paris' });
const u = db.get('u1'); // User | undefined

JSON + secondary indexes

const db = await MiniDb.open({ dir: './data', valueCodec: 'json' });

await db.createIndex('byCity', { field: 'city' });                 // equality
await db.createIndex('byAge',  { field: 'age',  type: 'range' });  // range (skip list)
await db.createIndex('byMail', { field: 'email', unique: true });  // unique

await db.set('u1', { name: 'Ann', city: 'Paris', age: 30, email: 'ann@example.com' });
await db.set('u2', { name: 'Bob', city: 'Paris', age: 41, email: 'bob@example.com' });

db.findEq('byCity', 'Paris');          // [{ key:'u1', value:{...} }, { key:'u2', ... }]
db.findRange('byAge', { min: 30, max: 40, count: 10 });

Index definitions are persisted and rebuilt from the store on startup.

Document model & querying (MongoDB-like subset)

A record is:

{ key: 'user:1234', value: { ...any JSON... }, dt1..dtN: <datetime columns> }
  • key — string ≤ 128 chars, unique, ordered (range / prefix / ordered scan).
  • value — any JSON object (jq-like filter + projection + full-text).
  • dt1..dtN — any number of top-level datetime columns, each indexed for O(log N) range queries. Pass them as { dt: { created: Date.now() } } (ISO strings or epoch ms).

Key scans (ordered)

db.scan({ gte: 'user:1', lte: 'user:9', limit: 100 }); // range in key order
db.prefix('user:');                                    // prefix scan

Datetime column queries

await db.set('a', { n: 1 }, { dt: { created: '2024-01-01' } });
db.dtColumns();                              // ['created']
db.dtRange('created', { gte: t0, lte: t1 }); // O(log N), [{ key, value, dt, dtValue }]

Unified query

db.query(q) composes key range, dt range, full-text, and a value filter:

db.query({
  key:    { prefix: 'post:' },
  dt:     { created: { gte: t0, lte: t1 } },
  text:   { index: 'body', q: '北京', op: 'AND' },
  filter: { age: { $gte: 18 }, $or: [{ city: 'Paris' }, { city: 'London' }] },
  sort:   { age: -1 },
  project: ['name', 'age'],
  skip: 0, limit: 20,
});

Filter operators: $eq $ne $gt $gte $lt $lte $in $nin $regex $exists $contains $type, plus $and $or $nor $not. Paths use dot/bracket notation ("address.city", "tags[0]").

Indexed dimensions (key / dt / text / a createIndex'd field) are fast; db.query() automatically uses matching equality/range value indexes for simple top-level predicates (including $and), while a value filter with no matching index falls back to a full collection scan, exactly like MongoDB without an index.

await db.createTextIndex('body', { fields: ['bio'] }); // fields optional (default: all strings)
db.search('body', 'hello 世界', { op: 'AND' });         // [{ key, value, score }]

Latin words + CJK unigram/bigram tokenization (no dictionary, zero deps), with TF-IDF ranking.

API

Method Description
MiniDb.open({ dir, valueCodec?, valueMode?, fsyncPolicy?, compactThresholdBytes?, autoCompact?, activeExpireIntervalMs?, recovery?, readOnly?, onLockFail?, maxMemoryBytes?, maxMemoryPolicy? }) Open / create a database
MiniDb.restore(srcDir, destDir, opts?) Restore a db.backup() directory and open it
MiniDb.openOrRebuild(opts, { onRebuild? }) Open; on corruption, discard + reopen empty (cache use). Never deletes a live-locked db
get(key) Decoded value or undefined
set(key, value, { ttl? }) Set; ttl in ms. Resolves per fsync policy
del(key) Delete; returns true if it existed
has(key), size Membership / live key count
mget([keys]), mset([[k,v],...]) Batch read / write
`batch([{op:'set' 'del', key, value?, ttl?, dt?}])`
expire(key, ttlMs), ttl(key) Set / read TTL
createIndex(name, { field, type?, unique?, sparse? }) Secondary index (json codec)
dropIndex(name), listIndexes() Manage indexes
findEq(name, value), findRange(name, opts) Query value indexes
createCompoundIndex(name, { groupBy, orderBy, orderType? }) Compound index (group + order, e.g. workspace + updatedAt)
compoundRange(name, groupValue, opts) Ordered range within a group, O(log N + limit), no full sort
dropCompoundIndex(name), listCompoundIndexes() Manage compound indexes
scan({ gte, gt, lte, lt, limit, reverse }) Ordered key range scan
prefix(p, limit?) Key prefix scan
dtColumns(), dtRange(col, opts) Datetime column range query
query({ key, dt, text, filter, project, sort, skip, limit }) Unified Mongo-like query
createTextIndex(name, { fields? }), dropTextIndex(name) Full-text index
search(name, q, { op?, limit? }) Full-text search
compact() Force a snapshot + WAL rewrite now
backup(destDir, { compact? }) Write a consistent online backup directory
close() Flush, fsync, close

Durability (fsyncPolicy)

Policy Guarantees Relative speed
always fsync after every flush slowest (safest)
everysec fsync once per second (default) fast, ≤1s loss window
no OS decides when to flush fastest, may lose seconds on power loss

Value storage mode (valueMode)

valueMode controls where value bulk lives:

const db = await MiniDb.open({
  dir: './data',
  valueCodec: 'json',
  valueMode: 'memory', // 'memory' | 'disk' | 'auto'
});
  • memory (default): values are kept in RAM, Redis-style. Reads are memory-bound.
  • disk: values stay inline in the durable snapshot/WAL, while RAM keeps only keys, metadata, indexes, and small { file, off, len } value pointers. Reads use synchronous positioned reads, so the public API stays synchronous. This allows value bulk to exceed RAM, at the cost of disk reads on cold get() paths.
  • auto: at startup, compare the current db.snapshot + db.wal size with maxMemoryBytes. If the persisted files exceed the budget, open as disk; otherwise open as memory. If maxMemoryBytes is not set, auto falls back to memory.

Memory limits

With valueMode: 'memory' the budget bounds keys plus values. With valueMode: 'disk' it bounds the in-RAM key/metadata/index footprint, not the value bulk. You can bound writes with an approximate memory budget:

const db = await MiniDb.open({
  dir: './data',
  valueCodec: 'json',
  maxMemoryBytes: 512 * 1024 * 1024,
  maxMemoryPolicy: 'reject', // or 'evict-lru'
});

reject throws when a write would exceed the budget; evict-lru durably deletes least-recently-used keys first. In valueMode: 'disk', evict-lru trims keys to bound the metadata/index footprint; it does not delete value bytes from disk because values are already stored outside RAM. db.stats.evictions and db.stats.maxMemoryRejections track the behavior.

The budget is tracked in approximate logical bytes (key + value + dt metadata), not real JS heap usage. Actual per-key overhead is higher — on the order of several hundred bytes per key for the Map entry, the ordered index node, and bookkeeping — so a database of many small keys needs far more RAM than store.bytes suggests. (Measured: ~0.6 KB of heap per key for tiny records, so 1M small keys ≈ 600 MB of heap.) For millions of small keys, prefer valueMode: 'disk' and/or budget accordingly.

Online backup / restore

await db.backup('./backup');
const restored = await MiniDb.restore('./backup', './restored', { valueCodec: 'json' });

backup() optionally compacts, then fences writes at a linearization point — writes submitted while the backup runs reject with a BACKUP_IN_PROGRESS error, and every write acknowledged before the fence is included — and copies the snapshot, WAL, index definitions, and text postings into a sibling temp directory that is fsync'd and atomically renamed over the destination (the manifest, written last, is the commit marker). A failed backup leaves no partial directory behind, and restore() can reopen the result.

Write throughput & SSD endurance

MiniDb's write path is append-only and sequential, which is already SSD-friendly. A few choices have an outsized effect on write amplification and flash wear:

  • Use batch writes. set() resolves after its own flush, so for (const x of xs) await db.set(...) issues one flush (and, under fsyncPolicy: 'always', one fsync) per key. Coalesce writes with db.batch([...]) / db.mset([[k, v], ...]) (single atomic WAL frame) or await Promise.all(xs.map(x => db.set(x))) (group commit collapses the tick into one writev + one fsync).
  • Keep the default fsyncPolicy: 'everysec'. It bounds fsyncs to ~1/s. always is the most durable but the hardest on flash (one fsync per flush); reserve it for data that must survive any crash.
  • Tune compactThresholdBytes for large datasets. Compaction rewrites all live data, so steady-state write amplification is roughly 1 + liveDataBytes / compactThresholdBytes. Raising the threshold (e.g. 256 MiB1 GiB) trades a larger WAL / longer recovery for less rewrite.
  • Observe it via db.stats. walBytesWritten, walFsyncs, snapshotBytesWritten, compactions, evictions, maxMemoryRejections, and queryIndexHits let you measure real write volume, fsync rate, memory pressure, and index usage under your workload.

RESP server (redis-cli compatible)

npm run server -- --dir ./data --port 6379
# or: node --import tsx src/server.ts --dir ./data --port 6379

Then in another shell:

redis-cli -p 6379 SET foo bar
redis-cli -p 6379 GET foo

Supported commands: PING ECHO GET SET DEL EXISTS MGET MSET TTL DBSIZE COMPACT INFO QUIT.

Benchmarks

Measured on Node v24, 100-byte values (npm run bench, N=100000):

Operation Throughput
Raw Map set (baseline) ~8.6 M ops/s
Raw Map get (baseline) ~20 M ops/s
DB get (in-memory) ~8.0 M ops/s
DB set, fsync=everysec (concurrent, group commit) ~725 k ops/s
DB set, fsync=no (concurrent, group commit) ~396 k ops/s
DB set, fsync=always (sequential, 1 fsync/op) ~328 ops/s
Compact snapshot of 100k keys ~69 ms (12 MiB)

Reads are memory-bound and track the raw Map closely. Writes are disk-bound; group commit makes concurrent writes very fast, while synchronous (always) writes pay the fsync cost one would expect from any database. Numbers vary by machine/disk.

Query benchmarks (N=50k docs, 200 iters each, node bench/query.js):

Query Throughput
dt range (small result) ~183 k ops/s
key prefix scan (~100 rows) ~23 k ops/s
full-text search (latin / CJK) ~350490 ops/s
value filter, no index (full scan of 50k) ~23 ops/s

Indexed dimensions (key / dt / text) are fast; an unindexed value filter scans the whole collection — create a secondary index (createIndex) for hot value predicates.

For ordered pagination within a group (e.g. "sessions in a workspace ordered by updatedAt"), use a compound index instead of fetching the whole group and sorting in memory. On 10k sessions in one workspace, paginating with compoundRange is ~2040× faster than "fetch-all + sort", and stays sub-millisecond at any offset.

Multi-process ClusterDb scaling across process counts × shard counts (the headline: shard-affinity writes scale ~linearly until the single-shard speed is reached; uniform cross-shard traffic pays lock handoffs):

pnpm bench:cluster   # bench/cluster.ts — spawns real writer/reader processes

Testing

Three layers of tests, run with the built-in node:test runner (no deps):

npm test            # unit tests (fast)
npm run test:e2e    # end-to-end stability suite
npm run test:all    # both

The unit tests (test/*.test.js) cover each module: frame codec/CRC, WAL group commit, store TTL, snapshot/compaction, recovery truncation, skip list, secondary/full-text indexes, and the RESP server.

The E2E stability suite (test/e2e/*.test.js) covers crash-safety and long-run behavior:

File What it verifies
fuzz-model.test.js thousands of random ops match a reference model (seeded, reproducible)
crash-recovery.test.js kill -9 mid-write and mid-compaction → recovery is always consistent
index-consistency.test.js key/dt/secondary/full-text indexes never drift from the store
compaction-race.test.js heavy concurrent writes during compaction lose nothing
recovery-matrix.test.js WAL corruption at head/mid/tail under resync vs strict
durability.test.js always/everysec/no close-durability + many open/close cycles
boundary.test.js key-length limits, large values, many keys, empty db
soak.test.js sustained ops + heap stability (opt-in: SOAK=30 npm run test:e2e)

The cluster suite (test/cluster/*.test.ts) covers the ClusterDb sharding layer: topology/routing, merged scans, lock contention and lease renewal in-process, cross-shard indexes/compaction, true multi-process scenarios (concurrent writers on disjoint and shared shards, live cross-process read visibility, read/write storms) and crash takeovers (kill -9 → contiguous recovery + stale-lock handoff).

Design in one paragraph

Log-structured engine: all writes append to a CRC-framed WAL (group committed, configurable fsync); all reads hit an in-memory Map. When the WAL grows past a threshold, a snapshot of the live keys is written and the WAL is rotated — non-blocking: writers keep appending to the WAL (which doubles as a Redis-style rewrite buffer) while the snapshot is written, and pause only for a brief final rotation. Recovery loads the latest snapshot then replays the WAL, truncating any torn tail. Secondary indexes use a Redis-style skip list for range queries. See DESIGN_NOTES.md for the full rationale and the source-code study behind each choice.

Concurrency & multi-process

minidb is single-writer. Opening a directory for writing acquires an exclusive lock file (db.lock); a second writer is rejected with a LockError. A lock is taken over only when its owner PID is dead (stale-lock recovery), never merely because it is old.

// second process: throws LockError
await MiniDb.open({ dir: './data' });

// degrade to read-only instead of throwing
const ro = await MiniDb.open({ dir: './data', onLockFail: 'readonly' });
ro.get('k');        // ok (point-in-time view)
await ro.set('k');  // throws "read-only mode"

// open read-only alongside a writer
const r = await MiniDb.open({ dir: './data', readOnly: true });

For many clients, run the RESP server (single minidb process, many TCP clients) — that is the intended concurrent-access model, like Redis.

ClusterDb: multi-process sharding

When several processes must read and write the same logical database without a server, ClusterDb shards the key space over N ordinary minidb directories (each with its own WAL, snapshot, and db.lock):

import { ClusterDb } from '@moonshot-ai/minidb/cluster';

const db = await ClusterDb.open({ dir: './data', shardCount: 16, valueCodec: 'json' });
await db.set('user:1', { name: 'alice' });      // routed by hash to one shard
await db.mset([['a:1', v1], ['b:2', v2]]);      // grouped per shard, atomic per shard
const all = await db.scan({ prefix: 'user:' }); // merged over all shards
await db.close();
  • Writes never meet a single global writer: each process acquires shard write locks on demand (with retry up to lockAcquireTimeoutMs), caches them briefly (lockPoolMaxShards, LRU), and yields a held shard after lockHoldMs (default 250ms) so other processes are never starved. Throughput scales with the number of distinct shards being written — up to the single-shard speed of MiniDb per shard.
  • Reads never take locks: they use the cached writer when the local process holds the shard, else a read-only instance that is revalidated against the shard files on every use, so a read that starts after another process's commit always observes it.
  • Consistency: single-key and same-shard batch ops are strongly consistent (single writer per shard, atomic WAL frames). Cross-shard mset/mdel/batch are best-effort (atomic per shard, not globally; crossShard: 'none' rejects them instead). Scans merge per-shard snapshots, so entries from different shards may reflect different points in time.
  • Indexes (createIndex/createTextIndex) are recorded in a cluster-wide registry and applied by every shard writer on open; findEq/findRange/ search merge per-shard results (text scores are per-shard). Index management acquires every shard writer — run it from one process, off the hot path.
  • Crash recovery is per shard: a lock left by a dead PID is taken over by the next opener, exactly like single MiniDb.

crossShard: '2pc' is reserved for a future two-phase commit and is rejected today. Performance numbers across process/shard counts: run pnpm bench:cluster (see bench/cluster.ts).

For a rebuildable cache, use openOrRebuild: a corrupt cache is discarded and reopened empty, while a live-locked db is never destroyed:

const db = await MiniDb.openOrRebuild(
  { dir: cacheDir, valueCodec: 'json' },
  { onRebuild: (err) => log.warn('cache rebuilt:', err.message) },
);

Caveats / roadmap

  • In the default valueMode: 'memory', the dataset must fit in RAM Redis-style. Use valueMode: 'disk' for larger-than-RAM value bulk; cold reads then perform synchronous positioned reads against the snapshot/WAL.
  • Full-text index postings are stored on disk (larger-than-RAM); only the term dictionary and per-doc metadata stay in memory. Postings are rebuilt from the store on open and on compaction. Search reads postings synchronously, so a cold, very large postings list can briefly block the event loop.
  • Compaction is non-blocking for writes: the WAL itself acts as a BGREWRITEAOF-style rewrite buffer, so the (slow) snapshot is written while writers keep appending. Writes pause only for the final rotation (a flush, a tail copy, and two atomic renames). A pre-copy drains most of the tail beforehand when writes are slow enough; under sustained writes that outrun the pre-copy, the rotation absorbs a larger tail — the same bounded end-of-rewrite pause Redis accepts for its AOF diff flush — so compaction always terminates. Mid-compaction crashes leave db.*.tmp files behind, which the next writer open removes automatically.
  • Snapshot encoding runs on the main thread (chunked + yielding); offloading to a worker_thread is a planned optimization.
  • Single minidb directory = single process / single writer. For multi-process access use ClusterDb (above): sharding scales writes, but hash routing means whole-range scans fan out to all shards, and uniform cross-shard write patterns pay shard-lock handoff costs (lockHoldMs per yield) — workloads with per-process shard affinity scale best.

Credits

Design distilled from reading: Redis (references/redis), the SQLite WAL paper, NeDB (references/nedb), Bitcask (references/bitcask), and the cstack SQLite tutorial (references/db_tutorial).