* perf(minidb): skip idle everysec fsyncs and add lifecycle stats - everysec WAL now fsyncs on the timer only while dirty (tracked by a write/sync generation watermark); close() keeps its unconditional final sync, and background sync failures surface via walFsyncErrors plus a sticky lastWalFsyncError instead of being silently swallowed - add WAL queue/group-commit counters (walQueuedBytes, walMaxQueuedBytes, walGroupCommits, walGroupCommitFrames) and lifecycle phase stats: recovery bytes/frames/duration, index/text rebuild durations, compaction total/snapshot/rotation/postings durations, rotation pause, and query candidates/decoded/sorted rows; add a syncIntervalMs open option threaded through compaction WAL rotation - rewrite the bench on fixed-seed synthetic data with a stable machine-readable JSON report (cold open 10k/50k/100k, word/ngram search, idle-fsync acceptance, 100k compaction, event-loop delay, peak heap/RSS per scenario) and pin the schema in test/bench-json.test.ts; add app-side baselines with loose complexity budgets in sessionIndex and searchService tests - fix ClusterDb lock-pool closeAll() leaking in-flight shard opens and drain the query store's async close on server shutdown, eliminating the ENOTEMPTY directory-teardown race * perf(minidb): bound startup rebuild and steady-state hot paths - rebuild all derived indexes in one shared store walk: a single decode per record fans out to staged builders, dt rebuild reads record metadata only, and index-less opens no longer decode at all - rank full-text results with a bounded min-heap plus a stable key tie-break instead of sorting every candidate - remove/overwrite text docs via a docID -> delta-terms reverse map instead of scanning the whole delta vocabulary - validate unique batches incrementally against touched postings instead of copying the full per-index owner map - reap due TTL entries from the expiry heap on the write path instead of a full-store sweep per write * feat(session-index): add minidb read model with keyset pagination - add ISessionIndex read-model lifecycle (prepare/status, ready/degraded states) behind the persistence_minidb_readmodel experimental flag - add ISessionIndexMirror write side recording fresh summaries into a bounded, coalescing queue after the authoritative document is durable - replace the offset cursor with before/after keyset pagination; rename list/countActive to listRecent/count - extend IQueryStore with ordered columns and pageByColumn, plus getMany/listKeys/dropCollection - wire the read model through kap-server routes and start, and update the klient sessions contract - index every session for global search instead of the 500 most recent * feat(kap-server): bound search sync lifecycle, pagination, and query budgets - split search requests from sync work: searchIndex() no longer awaits runSync/reopen/reindex; a single-flight sync coordinator with debounce and backpressure runs in the background, stale generations keep serving with explicit stale/degraded state, and refresh/sync/reindex failures surface via lastRefreshError instead of being swallowed - scope file-meta keys by session id (\0meta\file\<sessionId>\<hash>) with lazy + one-shot background migration from the legacy hash-only keys, so one session sync only touches its own meta rows - make authoritative scans incremental (mtime/ino/size rescan conditions, unchanged files no longer rewrite meta) and read wire deltas in 1 MiB chunks instead of whole-file buffer + split - replace offset pagination with versioned v2 keyset page tokens (fingerprint + index generation + sort boundary); generation changes fail old tokens with invalid_page_token, legacy v1 offset tokens are served once and upgraded, and pages collect via bounded top-K instead of full sort + offset skip - add query budgets enforced at the postings/score stage: max query terms, literal length cap, postings visit budget (minidb searchBounded/maxVisits with prefix decoding that never fabricates hits and skips the postings LRU), candidate caps, deadline and text budget; truncation is reported via incomplete reasons candidate_cap/postings_budget/deadline - reopen read-only dbs by opening the next handle before closing the previous one so a failed refresh keeps the old generation serving; failed opens now self-heal through search traffic 100k-message bench: first page p95 < 300ms and page-100 cost on par with page 1; event-loop delay during queries stays sub-millisecond. * fix(minidb): poison, roll back, and recover the WAL on write failures - give WAL writes a commit point: a failed flushBatch poisons the WAL (WAL_POISONED, tracked separately as walWriteErrors vs walFsyncErrors), rejects queued frames in reverse enqueue order, and stops scheduling further batches; everysec background sync failures stay non-rejecting per stage-1 semantics - recover in place to a known-safe point: a serialized recovery chain truncates the WAL back to the first un-acked frame, rebuilds size/nextOffset, and clears the poison; writes queue behind the recovery gate (zero-cost when idle), a failed truncate flips the instance into an explicit writeDisabled state, and a stale truncate offset (WAL file replaced by a rotation) skips the truncate - roll failed flush groups back as a unit: frames are stamped with their batchId, MiniDb keeps per-group earliest pre-state, and the first rejection restores every key of the group (rejected writes no longer reappear after reopen, and in-memory state matches reopen for any failure interleaving); the per-op seq guard remains for cross-group and rotation-retry races - wrap applyOp and the following in-memory mutations so a contract violation poisons the WAL and rolls the group back instead of escaping as a half-commit; frames never enqueued (seal race) roll back per-op without poisoning - tag errors past the commit point with ambiguous: true so callers can distinguish "definitely not applied" from "maybe applied but revoked" - close() waits for the recovery chain to go idle and backup() fences behind in-flight recovery before copying files Controlled A/B bench (22 alternating iterations, 100k concurrent sets): write-path throughput regression is within the 2% budget. * fix(minidb): turn the file lock into an instance-owned serialized lease - distinguish lock ownership by instance instead of pid: every acquire mints a pid:uuid token carried by lock/bid/watch files, inspect().mine compares tokens, liveness still follows pid, tokenless legacy files keep the old stale-takeover path, and hasLiveForeignWatch excludes self by token so same-process contenders see each other (closing the double-win takeover and the cross-instance release); a live same-pid lock is still respected, and re-acquiring a held lock is idempotent - serialize acquire/renew/release through a per-instance promise-chain mutex: renew re-checks held inside the chain and release waits for an in-flight renew, eliminating the renew/rename-after-unlink ghost lock - make MiniDb.close() a state machine (open/closing/closed) with a shared closePromise: cleanup runs per-resource try/catch in dependency order (text indexes, store, valueReader, WAL, lock), aggregates every cleanup error into an AggregateError, stays in 'closing' on failure so a retry finishes the cleanup, and no longer leaks the lock when the WAL close fails; a rejected in-flight compaction no longer escapes the cleanup pass * fix(minidb): keep readers on one consistent file generation - add an internal persistent-files module as the single source of truth for the persisted file set (snapshot, WAL, sidecars, postings pattern, fingerprint subset); lock-pool fingerprints, persistentFiles, open stale-tmp cleanup, and backup/restore filtering all derive from it, and fingerprints upgrade to dev:ino:size:mtimeMs so compound sidecar changes can no longer hide from cluster readers - pair snapshot and WAL generations during recovery (transitional stat-pairing until stage-5 manifests): each pass anchors the fds it scans, re-stats afterwards, tolerates append-only WAL growth, retries bounded times on generation churn with a clean store reset, and throws RECOVERY_GENERATION_CHURN when churn exceeds the budget; the disk-mode ValueReader attach re-validates inodes so stale offsets never read a replaced file - make the rotation directory fsyncs strict: failures abort the rotation through the existing rollback path instead of being swallowed, while platforms without directory fsync degrade once with a warn and stats.dirFsyncUnsupported * fix(minidb): serialize index-definition sidecar mutations, persist before publish - extract the promise-chain mutex into a shared createSerializer() and give each sidecar family (secondary/compound/text) its own chain: create/drop run uninterruptibly (memory change + rebuild + persist), different families stay independent, and the data write path never shares these chains - reverse the publication order to staged -> persist -> publish: a create stages the definition, rebuilds via the staged builder, persists the sidecar including the new definition, then publishes atomically; any failure discards the staged state leaving live and sidecar untouched (no phantom indexes, retry-safe); a drop persists the sidecar without the definition before removing it live; text index create/drop adopt the same pattern, replacing the hand-rolled unwind, and a dropping marker keeps compaction postings rebuilds out of the persist window - feed staged indexes from the incremental write path (add/remove/ checkUnique/checkUniqueBatch visit live+staged) so writes landing in the persist window are not lost at publish; queries still see live only - harden writeFileAtomic: instance-unique tmp names (.tmp-pid-seq), a strict fsyncDir after rename so a successful persist is crash durable, and whitelist-based stale-tmp cleanup that never touches lock tmp files * fix(minidb): validate writes before any side effect, canonicalize values once - canonical value at the write boundary: the json codec re-parses the encoded bytes once and every downstream consumer (unique checks, secondary/compound/text indexes, dt extraction) sees exactly the persisted representation, so getter/toJSON/Proxy documents can no longer diverge between the index view and the storage view - reorder the set/batch pipeline so every fallible check happens before any visible side effect: prepare (key/ttl checks, encoding, canonical decode, index field extraction, tokenization) -> unique checks -> ensureMemoryFor eviction -> commit; a constraint failure now leaves the database untouched (no more evicted victims on rejected inserts), and applyOp is structurally pure against pre-validated data - tokenize at the prepare boundary: TextIndex gains prepareAdd/ addPrepared and the buildQueue carries validated key+tokens mutations instead of raw docs, so a throwing custom tokenizer can no longer poison the live view or the queue, and custom-tokenizer output is rejected per token over 0xffff bytes before it can permanently break postings rebuilds; prepared tokens are keyed by index instance so a same-name drop+create mid-write re-tokenizes instead of crossing tokenizers - strict batch structure validation: scanBatchOpRefs/decodeBatchOps reject unknown op types, out-of-bounds lengths, and trailing bytes (offset must equal body length), so a valid-CRC but malformed batch is skipped as a unit and counted via RecoveryInfo.corruptBatches instead of being partially applied Bench vs the stage-1 baseline: json write throughput regression is within the 5% budget (median ~2-4% depending on the measurement). * feat(minidb): add OpTracker drain primitive and atomic backup, harden tests - introduce the internal OpTracker (close gate + in-flight counter with enter/leave/close/whenIdle and reference-counted pause/resume) and drive every shutdown/drain path from it: WAL background syncs are tracked so close() waits out an in-flight sync before closing the fd, cluster lock-pool closeAll() closes the gates and drains busy callbacks before closing handles, and MiniDb writes pass a write gate - make backup() atomic with a defined linearization point: pause the write gate, drain in-flight writes (every acknowledged write is now included), copy to a sibling temp dir with per-file fsyncs, write the manifest last as the commit marker, and rename into place; failures clean up and leave no partial backup, and concurrent writes are rejected with BACKUP_IN_PROGRESS - reap emptied compound-index groups on remove (the groups map no longer grows monotonically), move the open-time mkdir behind the readOnly check so a read-only open of a missing directory fails with ENOENT instead of creating it, and never run a destructive rebuild for a read-only open failure (explicit or onLockFail fallback) - consolidate every review fault-injection repro into the formal suite behind deterministic barrier helpers (programmable writev/sync/ rename/tokenize hooks) and convert the six timing-based tests to barrier/tick-driven assertions; the .tmp repro scripts are removed The converted timing tests and the full suite pass 50 repeat runs (including under CPU load injection) with zero flakes. * feat(minidb): persist derived indexes as atomic generations, open from WAL delta - checkpoint the store, dt/secondary/compound indexes, and text dictionary/postings/docs into immutable generations under generations/g-NNNNNN published atomically (tmp build, per-file checksums and fsyncs, dir rename, CURRENT swap, strict dir fsyncs); the manifest records the format version, WAL/snapshot checkpoint anchors, per-index definition hashes, and codec/value-mode compatibility - open now loads the published generation and replays only the WAL delta after its checkpoint: no full value decode, corpus tokenization, or postings rewrite on a normal reopen (warm opens are 3.5-13.8x faster at 100k/1M records); a definition change rebuilds only the affected index, and corrupt generation files fall back to the previous generation or the legacy full recovery without ever touching the authoritative snapshot/WAL - build generations transactionally with compaction (rotation plus derived state publish as one unit, replacing the synchronous rebuildTextPostings tail), capture concurrent writes through a sealed op queue with byte/op caps, hard-link clean postings and the snapshot into the new generation, and repoint every live text base into the CURRENT generation after publish - cluster/read-only refresh watches CURRENT and the WAL watermark: pure generation publishes keep readers on incremental catch-up while rotations reopen onto the new generation; writers building the next generation never disturb readers of the current one - legacy databases open through the old path unchanged and gain their first generation in the background; OpenOptions.indexGenerations: false fully restores the pre-generation behavior * feat(minidb): workerize text-index builds and split MiniDb into facets - split the monolithic src/index.ts into facet modules (mini-db, types, value-codec, memory-guard, backup, query-engine, text-registry, wal-group, generation-builder/loader, write-path, read-path, index-admin, lifecycle, stats) and move text-index.ts to text-index/ - run corpus-scale text-index builds off the main thread via the bounded worker engine (src/worker/), exported through the new worker-runtime subpath, with inline fallback for small corpora and rollback switches - defer the open-time fallback text rebuild into a maintenance task; searches on a not-yet-committed base raise TextIndexBuildingError - add the unified maintenance scheduler, bounded async read surface, and a maintenance bench - kap-server search: switch to searchBoundedAsync and serve the building page while the index base rebuilds after fallback recovery - kimi-code: install the SEA-bundled minidb text-build worker at startup, bundle it via the native asset scripts, and add the startup-trace util plus the KIMI_TUI_INPUT_LATENCY debug probe * fix(minidb): treat win32 EPERM as unsupported directory fsync - extract isUnsupportedDirectoryFsyncError and cover win32 EPERM - drop the one-shot console.warn; stats.dirFsyncUnsupported carries the degraded state * fix(kap-server): harden search-index dispose and drain lifecycle - dispose() now closes an OpTracker gate and drains in-flight sync/refresh passes before closing the db, so no background write can hit a closed handle; the deleteSessionDocs loop and trailing stats write skip once the gate closes (review #20) - drainGlobalSearchDisposals loops to a fixpoint so disposals registered while a drain is in flight are also awaited (review #21) - pin the post-open failure semantics with a regression test: a failed text-index setup closes the handle and the next open reacquires the writer lock instead of self-locking read-only (review #19) - export OpTracker from the minidb root for the search service's drain * chore: fix oxlint type-aware lint errors
24 KiB
minidb
A pure-Node.js (zero native addons) embedded key-value database that mixes ideas from Redis (in-memory KV speed, data structures, AOF-style rewrite) and SQLite (durable single-file persistence, WAL, indexed queries). Written in TypeScript, with strict types and zero runtime dependencies.
Built by studying the real source code of Redis, SQLite, NeDB, Bitcask, and a tiny SQLite clone — see
DESIGN_NOTES.mdfor what each one taught us.
Features
- O(1) GET/SET from an in-memory index (
Map), values as bytes. - Optional larger-than-RAM values:
valueMode: 'disk'keeps only value pointers in RAM and reads value bulk from the snapshot/WAL with synchronous positioned reads. - Durable: append-only WAL with group commit + three fsync policies
(
always/everysec/no), exactly like Redis AOF. - Crash-safe recovery: CRC-framed records; a torn tail from a crash is detected and truncated automatically.
- Snapshots + compaction (WAL rewrite) bound disk growth and speed up restart.
- TTL: lazy + active expiration (Redis-style, heap-backed).
- Secondary indexes: equality and range indexes over JSON documents, backed by a Redis-style skip list with O(log N) rank/range; unique and sparse supported; array fields indexed per element.
- RESP server (optional): talk to it with
redis-cli/ioredis. - ClusterDb (optional sharding layer): multi-process concurrent read/write over N sharded MiniDb directories, scaling write throughput with shard count.
- Value codecs:
buffer,string, orjson. - Zero dependencies, pure
node:*built-ins.
Install / layout
Written in TypeScript, zero runtime dependencies. Source is in src/; build
output goes to dist/ (generated by npm run build).
src/
index.ts public API barrel (re-exports; no logic)
mini-db.ts the MiniDb class (open/close, KV reads, index admin, delegates)
types.ts public option/record types (+ the internal PreparedOp)
value-codec.ts value codecs + key byte-string / atomic sidecar helpers
memory-guard.ts maxMemory facet (LRU access set, evict/reject enforce)
backup.ts online backup (write-gate fence + atomic copy)
query-engine.ts unified query facet (candidate collection, dt fast path)
text-registry.ts text-index registry facet (defs, create/drop, search)
wal-group.ts WAL group-commit + in-place recovery gating facet
generation-builder.ts persistent index generation build facet
generation-loader.ts generation load facet (image loads + WAL delta replay)
write-path.ts write path facet (set/del/batch/expire commit machinery)
read-path.ts read path facet (KV reads, scans, live-record feeds)
index-admin.ts secondary/compound index admin facet (+ open-time rebuild)
lifecycle.ts open/close/openOrRebuild/renewLock flows
stats.ts the stats object factory
codec.ts binary frame encode/decode + CRC parser + resync
crc32.ts table-based CRC-32
wal.ts buffered append + group commit + fsync policies
store.ts in-memory index + TTL (lazy + active) + value refs
value-reader.ts synchronous positioned reads for disk-backed values
snapshot.ts chunked snapshot writer
recovery.ts load snapshot + replay WAL + resync/truncate
compaction.ts WAL rewrite / rotation
skiplist.ts comparator-driven skip list with span
index-manager.ts secondary indexes (equality + range)
dt-index.ts datetime-column indexes
query.ts jq-like value query engine
text-index/ full-text index (in-RAM dictionary + on-disk postings)
text-postings.ts on-disk postings file for the full-text index
lockfile.ts exclusive file lock + read-only
server.ts optional RESP TCP server
npm run build # compile src/ -> dist/ (JS + .d.ts)
npm run typecheck # tsc --noEmit (strict)
Quick start (embedded)
import { MiniDb } from 'minidb';
const db = await MiniDb.open({
dir: './data',
valueCodec: 'string', // 'buffer' | 'string' | 'json'
fsyncPolicy: 'everysec', // 'always' | 'everysec' | 'no'
});
await db.set('hello', 'world');
db.get('hello'); // 'world'
await db.set('temp', 'x', { ttl: 5000 }); // expires in 5s
db.ttl('temp'); // remaining ms (-1 none, -2 missing)
await db.del('hello');
await db.close(); // flush + fsync + close
Within this repo, import from
./src/index.ts(run withtsx). As an installed package, import from'minidb'(the builtdist/output).
Re-open the same dir and your data is recovered from the snapshot + WAL.
Typed usage
MiniDb is generic over the value type:
interface User { name: string; age: number; city: string }
const db = await MiniDb.open<User>({ dir: './data', valueCodec: 'json' });
await db.set('u1', { name: 'Ann', age: 30, city: 'Paris' });
const u = db.get('u1'); // User | undefined
JSON + secondary indexes
const db = await MiniDb.open({ dir: './data', valueCodec: 'json' });
await db.createIndex('byCity', { field: 'city' }); // equality
await db.createIndex('byAge', { field: 'age', type: 'range' }); // range (skip list)
await db.createIndex('byMail', { field: 'email', unique: true }); // unique
await db.set('u1', { name: 'Ann', city: 'Paris', age: 30, email: 'ann@example.com' });
await db.set('u2', { name: 'Bob', city: 'Paris', age: 41, email: 'bob@example.com' });
db.findEq('byCity', 'Paris'); // [{ key:'u1', value:{...} }, { key:'u2', ... }]
db.findRange('byAge', { min: 30, max: 40, count: 10 });
Index definitions are persisted and rebuilt from the store on startup.
Document model & querying (MongoDB-like subset)
A record is:
{ key: 'user:1234', value: { ...any JSON... }, dt1..dtN: <datetime columns> }
- key — string ≤ 128 chars, unique, ordered (range / prefix / ordered scan).
- value — any JSON object (jq-like filter + projection + full-text).
- dt1..dtN — any number of top-level datetime columns, each indexed for
O(log N) range queries. Pass them as
{ dt: { created: Date.now() } }(ISO strings or epoch ms).
Key scans (ordered)
db.scan({ gte: 'user:1', lte: 'user:9', limit: 100 }); // range in key order
db.prefix('user:'); // prefix scan
Datetime column queries
await db.set('a', { n: 1 }, { dt: { created: '2024-01-01' } });
db.dtColumns(); // ['created']
db.dtRange('created', { gte: t0, lte: t1 }); // O(log N), [{ key, value, dt, dtValue }]
Unified query
db.query(q) composes key range, dt range, full-text, and a value filter:
db.query({
key: { prefix: 'post:' },
dt: { created: { gte: t0, lte: t1 } },
text: { index: 'body', q: '北京', op: 'AND' },
filter: { age: { $gte: 18 }, $or: [{ city: 'Paris' }, { city: 'London' }] },
sort: { age: -1 },
project: ['name', 'age'],
skip: 0, limit: 20,
});
Filter operators: $eq $ne $gt $gte $lt $lte $in $nin $regex $exists $contains $type, plus $and $or $nor $not. Paths use dot/bracket notation
("address.city", "tags[0]").
Indexed dimensions (key / dt / text / a
createIndex'd field) are fast;db.query()automatically uses matching equality/range value indexes for simple top-level predicates (including$and), while a valuefilterwith no matching index falls back to a full collection scan, exactly like MongoDB without an index.
Full-text search
await db.createTextIndex('body', { fields: ['bio'] }); // fields optional (default: all strings)
db.search('body', 'hello 世界', { op: 'AND' }); // [{ key, value, score }]
Latin words + CJK unigram/bigram tokenization (no dictionary, zero deps), with TF-IDF ranking.
API
| Method | Description |
|---|---|
MiniDb.open({ dir, valueCodec?, valueMode?, fsyncPolicy?, compactThresholdBytes?, autoCompact?, activeExpireIntervalMs?, recovery?, readOnly?, onLockFail?, maxMemoryBytes?, maxMemoryPolicy? }) |
Open / create a database |
MiniDb.restore(srcDir, destDir, opts?) |
Restore a db.backup() directory and open it |
MiniDb.openOrRebuild(opts, { onRebuild? }) |
Open; on corruption, discard + reopen empty (cache use). Never deletes a live-locked db |
get(key) |
Decoded value or undefined |
set(key, value, { ttl? }) |
Set; ttl in ms. Resolves per fsync policy |
del(key) |
Delete; returns true if it existed |
has(key), size |
Membership / live key count |
mget([keys]), mset([[k,v],...]) |
Batch read / write |
| `batch([{op:'set' | 'del', key, value?, ttl?, dt?}])` |
expire(key, ttlMs), ttl(key) |
Set / read TTL |
createIndex(name, { field, type?, unique?, sparse? }) |
Secondary index (json codec) |
dropIndex(name), listIndexes() |
Manage indexes |
findEq(name, value), findRange(name, opts) |
Query value indexes |
createCompoundIndex(name, { groupBy, orderBy, orderType? }) |
Compound index (group + order, e.g. workspace + updatedAt) |
compoundRange(name, groupValue, opts) |
Ordered range within a group, O(log N + limit), no full sort |
dropCompoundIndex(name), listCompoundIndexes() |
Manage compound indexes |
scan({ gte, gt, lte, lt, limit, reverse }) |
Ordered key range scan |
prefix(p, limit?) |
Key prefix scan |
dtColumns(), dtRange(col, opts) |
Datetime column range query |
query({ key, dt, text, filter, project, sort, skip, limit }) |
Unified Mongo-like query |
createTextIndex(name, { fields? }), dropTextIndex(name) |
Full-text index |
search(name, q, { op?, limit? }) |
Full-text search |
compact() |
Force a snapshot + WAL rewrite now |
backup(destDir, { compact? }) |
Write a consistent online backup directory |
close() |
Flush, fsync, close |
Durability (fsyncPolicy)
| Policy | Guarantees | Relative speed |
|---|---|---|
always |
fsync after every flush | slowest (safest) |
everysec |
fsync once per second (default) | fast, ≤1s loss window |
no |
OS decides when to flush | fastest, may lose seconds on power loss |
Value storage mode (valueMode)
valueMode controls where value bulk lives:
const db = await MiniDb.open({
dir: './data',
valueCodec: 'json',
valueMode: 'memory', // 'memory' | 'disk' | 'auto'
});
memory(default): values are kept in RAM, Redis-style. Reads are memory-bound.disk: values stay inline in the durable snapshot/WAL, while RAM keeps only keys, metadata, indexes, and small{ file, off, len }value pointers. Reads use synchronous positioned reads, so the public API stays synchronous. This allows value bulk to exceed RAM, at the cost of disk reads on coldget()paths.auto: at startup, compare the currentdb.snapshot+db.walsize withmaxMemoryBytes. If the persisted files exceed the budget, open asdisk; otherwise open asmemory. IfmaxMemoryBytesis not set,autofalls back tomemory.
Memory limits
With valueMode: 'memory' the budget bounds keys plus values. With
valueMode: 'disk' it bounds the in-RAM key/metadata/index footprint, not the
value bulk. You can bound writes with an approximate memory budget:
const db = await MiniDb.open({
dir: './data',
valueCodec: 'json',
maxMemoryBytes: 512 * 1024 * 1024,
maxMemoryPolicy: 'reject', // or 'evict-lru'
});
reject throws when a write would exceed the budget; evict-lru durably deletes
least-recently-used keys first. In valueMode: 'disk', evict-lru trims keys to
bound the metadata/index footprint; it does not delete value bytes from disk
because values are already stored outside RAM. db.stats.evictions and
db.stats.maxMemoryRejections track the behavior.
The budget is tracked in approximate logical bytes (key + value + dt
metadata), not real JS heap usage. Actual per-key overhead is higher — on the
order of several hundred bytes per key for the Map entry, the ordered index
node, and bookkeeping — so a database of many small keys needs far more RAM
than store.bytes suggests. (Measured: ~0.6 KB of heap per key for tiny
records, so 1M small keys ≈ 600 MB of heap.) For millions of small keys,
prefer valueMode: 'disk' and/or budget accordingly.
Online backup / restore
await db.backup('./backup');
const restored = await MiniDb.restore('./backup', './restored', { valueCodec: 'json' });
backup() optionally compacts, then fences writes at a linearization point —
writes submitted while the backup runs reject with a BACKUP_IN_PROGRESS
error, and every write acknowledged before the fence is included — and copies
the snapshot, WAL, index definitions, and text postings into a sibling temp
directory that is fsync'd and atomically renamed over the destination (the
manifest, written last, is the commit marker). A failed backup leaves no
partial directory behind, and restore() can reopen the result.
Write throughput & SSD endurance
MiniDb's write path is append-only and sequential, which is already SSD-friendly. A few choices have an outsized effect on write amplification and flash wear:
- Use batch writes.
set()resolves after its own flush, sofor (const x of xs) await db.set(...)issues one flush (and, underfsyncPolicy: 'always', one fsync) per key. Coalesce writes withdb.batch([...])/db.mset([[k, v], ...])(single atomic WAL frame) orawait Promise.all(xs.map(x => db.set(x)))(group commit collapses the tick into onewritev+ one fsync). - Keep the default
fsyncPolicy: 'everysec'. It bounds fsyncs to ~1/s.alwaysis the most durable but the hardest on flash (one fsync per flush); reserve it for data that must survive any crash. - Tune
compactThresholdBytesfor large datasets. Compaction rewrites all live data, so steady-state write amplification is roughly1 + liveDataBytes / compactThresholdBytes. Raising the threshold (e.g. 256 MiB–1 GiB) trades a larger WAL / longer recovery for less rewrite. - Observe it via
db.stats.walBytesWritten,walFsyncs,snapshotBytesWritten,compactions,evictions,maxMemoryRejections, andqueryIndexHitslet you measure real write volume, fsync rate, memory pressure, and index usage under your workload.
RESP server (redis-cli compatible)
npm run server -- --dir ./data --port 6379
# or: node --import tsx src/server.ts --dir ./data --port 6379
Then in another shell:
redis-cli -p 6379 SET foo bar
redis-cli -p 6379 GET foo
Supported commands: PING ECHO GET SET DEL EXISTS MGET MSET TTL DBSIZE COMPACT INFO QUIT.
Benchmarks
Measured on Node v24, 100-byte values (npm run bench, N=100000):
| Operation | Throughput |
|---|---|
Raw Map set (baseline) |
~8.6 M ops/s |
Raw Map get (baseline) |
~20 M ops/s |
| DB get (in-memory) | ~8.0 M ops/s |
| DB set, fsync=everysec (concurrent, group commit) | ~725 k ops/s |
| DB set, fsync=no (concurrent, group commit) | ~396 k ops/s |
| DB set, fsync=always (sequential, 1 fsync/op) | ~328 ops/s |
| Compact snapshot of 100k keys | ~69 ms (12 MiB) |
Reads are memory-bound and track the raw Map closely. Writes are
disk-bound; group commit makes concurrent writes very fast, while
synchronous (always) writes pay the fsync cost one would expect from any
database. Numbers vary by machine/disk.
Query benchmarks (N=50k docs, 200 iters each, node bench/query.js):
| Query | Throughput |
|---|---|
| dt range (small result) | ~183 k ops/s |
| key prefix scan (~100 rows) | ~23 k ops/s |
| full-text search (latin / CJK) | ~350–490 ops/s |
| value filter, no index (full scan of 50k) | ~23 ops/s |
Indexed dimensions (key / dt / text) are fast; an unindexed value filter scans
the whole collection — create a secondary index (createIndex) for hot value
predicates.
For ordered pagination within a group (e.g. "sessions in a workspace ordered
by updatedAt"), use a compound index instead of fetching the whole group and
sorting in memory. On 10k sessions in one workspace, paginating with
compoundRange is ~20–40× faster than "fetch-all + sort", and stays
sub-millisecond at any offset.
Multi-process ClusterDb scaling across process counts × shard counts (the
headline: shard-affinity writes scale ~linearly until the single-shard speed
is reached; uniform cross-shard traffic pays lock handoffs):
pnpm bench:cluster # bench/cluster.ts — spawns real writer/reader processes
Testing
Three layers of tests, run with the built-in node:test runner (no deps):
npm test # unit tests (fast)
npm run test:e2e # end-to-end stability suite
npm run test:all # both
The unit tests (test/*.test.js) cover each module: frame codec/CRC, WAL
group commit, store TTL, snapshot/compaction, recovery truncation, skip list,
secondary/full-text indexes, and the RESP server.
The E2E stability suite (test/e2e/*.test.js) covers crash-safety and
long-run behavior:
| File | What it verifies |
|---|---|
fuzz-model.test.js |
thousands of random ops match a reference model (seeded, reproducible) |
crash-recovery.test.js |
kill -9 mid-write and mid-compaction → recovery is always consistent |
index-consistency.test.js |
key/dt/secondary/full-text indexes never drift from the store |
compaction-race.test.js |
heavy concurrent writes during compaction lose nothing |
recovery-matrix.test.js |
WAL corruption at head/mid/tail under resync vs strict |
durability.test.js |
always/everysec/no close-durability + many open/close cycles |
boundary.test.js |
key-length limits, large values, many keys, empty db |
soak.test.js |
sustained ops + heap stability (opt-in: SOAK=30 npm run test:e2e) |
The cluster suite (test/cluster/*.test.ts) covers the ClusterDb
sharding layer: topology/routing, merged scans, lock contention and lease
renewal in-process, cross-shard indexes/compaction, true multi-process
scenarios (concurrent writers on disjoint and shared shards, live cross-process
read visibility, read/write storms) and crash takeovers (kill -9 → contiguous
recovery + stale-lock handoff).
Design in one paragraph
Log-structured engine: all writes append to a CRC-framed WAL (group
committed, configurable fsync); all reads hit an in-memory Map. When the WAL
grows past a threshold, a snapshot of the live keys is written and the WAL
is rotated — non-blocking: writers keep appending to the WAL (which doubles
as a Redis-style rewrite buffer) while the snapshot is written, and pause only
for a brief final rotation. Recovery loads the latest
snapshot then replays the WAL, truncating any torn tail. Secondary indexes use
a Redis-style skip list for range queries. See
DESIGN_NOTES.md for the full rationale and the
source-code study behind each choice.
Concurrency & multi-process
minidb is single-writer. Opening a directory for writing acquires an exclusive
lock file (db.lock); a second writer is rejected with a LockError. A lock is
taken over only when its owner PID is dead (stale-lock recovery), never merely
because it is old.
// second process: throws LockError
await MiniDb.open({ dir: './data' });
// degrade to read-only instead of throwing
const ro = await MiniDb.open({ dir: './data', onLockFail: 'readonly' });
ro.get('k'); // ok (point-in-time view)
await ro.set('k'); // throws "read-only mode"
// open read-only alongside a writer
const r = await MiniDb.open({ dir: './data', readOnly: true });
For many clients, run the RESP server (single minidb process, many TCP clients) — that is the intended concurrent-access model, like Redis.
ClusterDb: multi-process sharding
When several processes must read and write the same logical database
without a server, ClusterDb shards the key space over N ordinary minidb
directories (each with its own WAL, snapshot, and db.lock):
import { ClusterDb } from '@moonshot-ai/minidb/cluster';
const db = await ClusterDb.open({ dir: './data', shardCount: 16, valueCodec: 'json' });
await db.set('user:1', { name: 'alice' }); // routed by hash to one shard
await db.mset([['a:1', v1], ['b:2', v2]]); // grouped per shard, atomic per shard
const all = await db.scan({ prefix: 'user:' }); // merged over all shards
await db.close();
- Writes never meet a single global writer: each process acquires shard
write locks on demand (with retry up to
lockAcquireTimeoutMs), caches them briefly (lockPoolMaxShards, LRU), and yields a held shard afterlockHoldMs(default 250ms) so other processes are never starved. Throughput scales with the number of distinct shards being written — up to the single-shard speed of MiniDb per shard. - Reads never take locks: they use the cached writer when the local process holds the shard, else a read-only instance that is revalidated against the shard files on every use, so a read that starts after another process's commit always observes it.
- Consistency: single-key and same-shard batch ops are strongly
consistent (single writer per shard, atomic WAL frames). Cross-shard
mset/mdel/batchare best-effort (atomic per shard, not globally;crossShard: 'none'rejects them instead). Scans merge per-shard snapshots, so entries from different shards may reflect different points in time. - Indexes (
createIndex/createTextIndex) are recorded in a cluster-wide registry and applied by every shard writer on open;findEq/findRange/searchmerge per-shard results (text scores are per-shard). Index management acquires every shard writer — run it from one process, off the hot path. - Crash recovery is per shard: a lock left by a dead PID is taken over by the next opener, exactly like single MiniDb.
crossShard: '2pc' is reserved for a future two-phase commit and is rejected
today. Performance numbers across process/shard counts: run
pnpm bench:cluster (see bench/cluster.ts).
For a rebuildable cache, use openOrRebuild: a corrupt cache is discarded
and reopened empty, while a live-locked db is never destroyed:
const db = await MiniDb.openOrRebuild(
{ dir: cacheDir, valueCodec: 'json' },
{ onRebuild: (err) => log.warn('cache rebuilt:', err.message) },
);
Caveats / roadmap
- In the default
valueMode: 'memory', the dataset must fit in RAM Redis-style. UsevalueMode: 'disk'for larger-than-RAM value bulk; cold reads then perform synchronous positioned reads against the snapshot/WAL. - Full-text index postings are stored on disk (larger-than-RAM); only the term dictionary and per-doc metadata stay in memory. Postings are rebuilt from the store on open and on compaction. Search reads postings synchronously, so a cold, very large postings list can briefly block the event loop.
- Compaction is non-blocking for writes: the WAL itself acts as a
BGREWRITEAOF-style rewrite buffer, so the (slow) snapshot is written while writers keep appending. Writes pause only for the final rotation (a flush, a tail copy, and two atomic renames). A pre-copy drains most of the tail beforehand when writes are slow enough; under sustained writes that outrun the pre-copy, the rotation absorbs a larger tail — the same bounded end-of-rewrite pause Redis accepts for its AOF diff flush — so compaction always terminates. Mid-compaction crashes leavedb.*.tmpfiles behind, which the next writer open removes automatically. - Snapshot encoding runs on the main thread (chunked + yielding); offloading to
a
worker_threadis a planned optimization. - Single minidb directory = single process / single writer. For multi-process
access use
ClusterDb(above): sharding scales writes, but hash routing means whole-range scans fan out to all shards, and uniform cross-shard write patterns pay shard-lock handoff costs (lockHoldMsper yield) — workloads with per-process shard affinity scale best.
Credits
Design distilled from reading: Redis (references/redis), the SQLite WAL
paper, NeDB (references/nedb), Bitcask (references/bitcask), and the
cstack SQLite tutorial (references/db_tutorial).