kimi-code/packages/minidb
Haozhe 1328b32037
feat(acp): add experimental agent-core-v2 ACP server (kimi acp-v2) (#2571)
* feat(acp): add agent-core-v2 ACP server

- add ACP session lifecycle, configuration, permissions, and event bridging
- expose the experimental kimi acp-v2 command with terminal authentication
- add integration coverage and workspace build configuration

* test: use neutral example domains in test fixtures and docs

- replace placeholder hostnames (evil.com, foo.com, internal.corp,
  real.corp) with example.test / example.com in agent-core-v2 and
  kap-server tests
- replace fixture emails (x@y.com, a@x.com) with example addresses in
  minidb tests and README

* fix(acp): align acp-server with agent-core-v2 interfaces and address review

- add missing appendText to AcpHostFileSystem (IHostFileSystem drift)
- replace IAgentPromptService.prompt with inject
- use Turn.cancel() instead of abortController
- gate FS reverse-RPCs on client capabilities, fallback to local FS
- return PROTOCOL_VERSION constant instead of echoing client version
- remove misleading mcpCapabilities from initialize response
- dispose old session wrapper before replacing on load/resume
- fix object stringification lint error in convert.ts
- add acp-v2 to expected CLI sub-command list in test

* fix(acp): use enqueue for prompt submission, stop advertising unimplemented builtins

- replace IAgentPromptService.inject with enqueue so onBeforeSubmitPrompt
  hooks (prompt-blocking policy) are not bypassed
- stop advertising builtin slash commands (/help, /status, etc.) until
  builtin command execution is implemented
- add comment explaining appendText stays local (ACP has no append RPC)
- update skills test to match new availableCommands behavior

* fix(acp): filter turn events by turnId, surface auth failures as auth_required

- track turnId in driveTurn and ignore events from unrelated turns,
  preventing queued prompts from settling on the running turn
- reject prompt requests with auth_required when turn fails with an
  auth-related error code, enabling ACP client re-auth flow

* fix(acp): gate acp-v2 behind experimental flag, filter sessions by cwd

- add acp-v2 experimental flag (KIMI_CODE_EXPERIMENTAL_ACP_V2) and gate
  CLI command registration behind it
- filter session/list results by requested cwd instead of returning
  sessions from all workspaces
- detect hook-blocked prompts via PromptHandle.state and add TODO for
  streaming block messages once the hook context exposes them

* refactor(acp-server): rewire ACP server onto the klient facade

- replace direct agent-core-v2 scope/service access (ISessionLifecycleService,
  ISessionIndex, IEventBus, ISessionInteractionService, etc.) with the Klient
  facade: klient.global.sessions / klient.session(id) / agent('main') handles
- drive turns via agent.prompt() + session-level agent event subscriptions
  instead of per-prompt IEventBus wiring; settle on turn.ended
- route approval/question bridging through session.interactions events
- hide the thinking config option and skill catalog behind KLIENT-GAP markers
  until klient exposes those surfaces
- acp-fs: pass realpath through to the local inner backend
- klient: session.restore() rejects both null and undefined handles

* feat(agent-core-v2): add session delete and ephemeral per-session MCP servers

- add ISessionLifecycleService.delete: close a live session first, then
  remove its persisted data, evict the index read-model entry, and append
  a deleted tombstone to session_index.jsonl; unknown ids raise
  session.not_found
- add CreateSessionOptions/ResumeSessionOptions.mcpServers: session-owned
  MCP overlay merged over the workspace manager via
  MergedMcpConnectionView (an ephemeral name shadows a workspace server),
  never persisted, released when the session scope tears down
- return PromptLaunchResult from activateSkill so callers get the
  launched turn id and activation failures (unknown skill, busy) surface
- add ISessionSkillCatalog.list() as a wire-friendly catalog snapshot
- add ISessionIndex.remove for read-model eviction on delete

* feat(klient): expose session delete, per-session MCP, skills, and stream events

- session lifecycle contract: delete, resume/restore options, and
  CreateSessionOptions.mcpServers (ephemeral per-session MCP servers)
- add the session skills contract and facade accessors for the
  wire-friendly skill catalog snapshot
- register tool.call.delta, tool.progress, and compaction.* agent stream
  events so consumers can subscribe with typed payloads

* feat(acp-server): align ACP v2 server with acp-adapter capabilities

- complete the klient-facade rewire: ACP client connection holder and
  the terminal/* reverse-RPC runner routed through the Agent scope
- negotiate the protocol version on initialize instead of pinning v1
- compress oversized prompt images at the ACP ingestion point with a
  format gate, caption, and persisted originals; a cancel arriving
  mid-compression settles the prompt as cancelled without a turn
- stream tool call args via tool.call.delta (lazy pending create,
  cumulative replace, started upgrade) and refresh titles via
  tool.progress status updates
- report compaction progress and results after /compact via the
  compaction.* events
- answer unknown slash commands locally instead of sending them to the
  model
- accept legacy "<id>,thinking" model ids and legacy approve /
  approve_for_session approval option ids
- keep sessions without cwd metadata in cwd-filtered session/list
- sanitize wire errors: auth codes map to auth_required, turn.agent_busy
  to invalid_request, everything else to a fixed internal-error message
- bump @agentclientprotocol/sdk to ^1.3.0

* fix(cli): drop stale registerServerCommand call and sherif ACP SDK split

- commands.ts called registerServerCommand, which no longer exists on
  current main (the deprecated `kimi server` shim is registered via
  registerWebCommand), breaking typecheck, build, and every CLI test
  that builds the program
- sherif rejects the @agentclientprotocol/sdk major split between
  acp-adapter (^0.23.0, production kimi acp) and acp-server (^1.3.0,
  experimental); the two hosts legitimately target different SDK
  majors, so ignore the dependency in the sherif invocation

* test: update fixtures for acp-v2 flag and domain rename, refresh nix deps hash

- kap-server origin.test: two CORS cases still used foo.com after the
  whitelist moved to foo.example.com, so the origin was no longer
  whitelisted and the expected CORS headers were withheld
- node-sdk config.test: expect the new acp-v2 experimental flag in the
  harness feature metadata
- flake.nix: update the fetchPnpmDeps hash for the
  @agentclientprotocol/sdk 1.3.0 lockfile change

* fix(acp): widen the ACP v2 auth gate beyond OAuth-only providers

The gate consulted only auth.summarize(), which iterates providers
declaring an oauth section — configurations that authenticate with a
plain apiKey or provider env-bag credentials (no OAuth at all) were
rejected with auth_required even though the default model is fully
usable.

- klient: expose authSummaryService.ensureReady on the global auth
  facade (the contract already declared it)
- acp-server: gate on the engine's own readiness probe for the default
  model — config apiKey / env-bag / OAuth token all count, matching how
  the model is actually used — and fall back to "any logged-in OAuth
  provider" (the legacy adapter's first branch)
- test: an apiKey-only config passes the gate with auth enforcement on;
  the OAuth logout regression is unchanged

* fix(acp): reject concurrent prompts instead of displacing the in-flight turn

A second session/prompt while a turn is running overwrote the session's
only TurnDriver: the engine quietly queues plain prompts submitted
during an active turn (the launch resolves undefined, indistinguishable
from a hook-blocked launch), so the first prompt never settled and both
turns' events went unattributed.

Guard both model-bound launch paths (plain prompt and skill activation)
with a synchronous in-flight check and reject with invalid_request
(turn.agent_busy), matching the legacy adapter's busy semantics. Local
slash handling (builtins, unknown-command answers) is unaffected.
2026-08-04 10:20:24 +08:00
..
bench feat(minidb): ClusterDb sharding, incremental reader catch-up, and engine hardening (#1816) 2026-07-17 15:57:10 +08:00
src feat(kap-server): add session-less POST /workspace/fs:search route (#2437) 2026-07-31 12:11:26 +08:00
test feat(acp): add experimental agent-core-v2 ACP server (kimi acp-v2) (#2571) 2026-08-04 10:20:24 +08:00
.gitignore feat(v2): land agent-core-v2 engine and kap-server behind experimental flag (#1441) 2026-07-12 21:44:04 +08:00
CHANGELOG.md ci: release packages (#1785) 2026-07-17 20:45:40 +08:00
DESIGN_NOTES.md feat(minidb): ClusterDb sharding, incremental reader catch-up, and engine hardening (#1816) 2026-07-17 15:57:10 +08:00
package.json ci: release packages (#1785) 2026-07-17 20:45:40 +08:00
README.md feat(acp): add experimental agent-core-v2 ACP server (kimi acp-v2) (#2571) 2026-08-04 10:20:24 +08:00
tsconfig.json feat(v2): land agent-core-v2 engine and kap-server behind experimental flag (#1441) 2026-07-12 21:44:04 +08:00
tsdown.config.ts feat(minidb): ClusterDb sharding, incremental reader catch-up, and engine hardening (#1816) 2026-07-17 15:57:10 +08:00
vitest.config.ts feat(minidb): ClusterDb sharding, incremental reader catch-up, and engine hardening (#1816) 2026-07-17 15:57:10 +08:00

minidb

A pure-Node.js (zero native addons) embedded key-value database that mixes ideas from Redis (in-memory KV speed, data structures, AOF-style rewrite) and SQLite (durable single-file persistence, WAL, indexed queries). Written in TypeScript, with strict types and zero runtime dependencies.

Built by studying the real source code of Redis, SQLite, NeDB, Bitcask, and a tiny SQLite clone — see DESIGN_NOTES.md for what each one taught us.

Features

  • O(1) GET/SET from an in-memory index (Map), values as bytes.
  • Optional larger-than-RAM values: valueMode: 'disk' keeps only value pointers in RAM and reads value bulk from the snapshot/WAL with synchronous positioned reads.
  • Durable: append-only WAL with group commit + three fsync policies (always / everysec / no), exactly like Redis AOF.
  • Crash-safe recovery: CRC-framed records; a torn tail from a crash is detected and truncated automatically.
  • Snapshots + compaction (WAL rewrite) bound disk growth and speed up restart.
  • TTL: lazy + active expiration (Redis-style, heap-backed).
  • Secondary indexes: equality and range indexes over JSON documents, backed by a Redis-style skip list with O(log N) rank/range; unique and sparse supported; array fields indexed per element.
  • RESP server (optional): talk to it with redis-cli / ioredis.
  • ClusterDb (optional sharding layer): multi-process concurrent read/write over N sharded MiniDb directories, scaling write throughput with shard count.
  • Value codecs: buffer, string, or json.
  • Zero dependencies, pure node:* built-ins.

Install / layout

Written in TypeScript, zero runtime dependencies. Source is in src/; build output goes to dist/ (generated by npm run build).

src/
  index.ts          public API (MiniDb) + types
  codec.ts          binary frame encode/decode + CRC parser + resync
  crc32.ts          table-based CRC-32
  wal.ts            buffered append + group commit + fsync policies
  store.ts          in-memory index + TTL (lazy + active) + value refs
  value-reader.ts   synchronous positioned reads for disk-backed values
  snapshot.ts       chunked snapshot writer
  recovery.ts       load snapshot + replay WAL + resync/truncate
  compaction.ts     WAL rewrite / rotation
  skiplist.ts       comparator-driven skip list with span
  index-manager.ts  secondary indexes (equality + range)
  dt-index.ts       datetime-column indexes
  query.ts          jq-like value query engine
  text-index.ts     full-text index (in-RAM dictionary + on-disk postings)
  text-postings.ts  on-disk postings file for the full-text index
  lockfile.ts       exclusive file lock + read-only
  server.ts         optional RESP TCP server
npm run build      # compile src/ -> dist/ (JS + .d.ts)
npm run typecheck  # tsc --noEmit (strict)

Quick start (embedded)

import { MiniDb } from 'minidb';

const db = await MiniDb.open({
  dir: './data',
  valueCodec: 'string',     // 'buffer' | 'string' | 'json'
  fsyncPolicy: 'everysec',  // 'always' | 'everysec' | 'no'
});

await db.set('hello', 'world');
db.get('hello');            // 'world'

await db.set('temp', 'x', { ttl: 5000 }); // expires in 5s
db.ttl('temp');             // remaining ms (-1 none, -2 missing)

await db.del('hello');
await db.close();           // flush + fsync + close

Within this repo, import from ./src/index.ts (run with tsx). As an installed package, import from 'minidb' (the built dist/ output).

Re-open the same dir and your data is recovered from the snapshot + WAL.

Typed usage

MiniDb is generic over the value type:

interface User { name: string; age: number; city: string }
const db = await MiniDb.open<User>({ dir: './data', valueCodec: 'json' });
await db.set('u1', { name: 'Ann', age: 30, city: 'Paris' });
const u = db.get('u1'); // User | undefined

JSON + secondary indexes

const db = await MiniDb.open({ dir: './data', valueCodec: 'json' });

await db.createIndex('byCity', { field: 'city' });                 // equality
await db.createIndex('byAge',  { field: 'age',  type: 'range' });  // range (skip list)
await db.createIndex('byMail', { field: 'email', unique: true });  // unique

await db.set('u1', { name: 'Ann', city: 'Paris', age: 30, email: 'ann@example.com' });
await db.set('u2', { name: 'Bob', city: 'Paris', age: 41, email: 'bob@example.com' });

db.findEq('byCity', 'Paris');          // [{ key:'u1', value:{...} }, { key:'u2', ... }]
db.findRange('byAge', { min: 30, max: 40, count: 10 });

Index definitions are persisted and rebuilt from the store on startup.

Document model & querying (MongoDB-like subset)

A record is:

{ key: 'user:1234', value: { ...any JSON... }, dt1..dtN: <datetime columns> }
  • key — string ≤ 128 chars, unique, ordered (range / prefix / ordered scan).
  • value — any JSON object (jq-like filter + projection + full-text).
  • dt1..dtN — any number of top-level datetime columns, each indexed for O(log N) range queries. Pass them as { dt: { created: Date.now() } } (ISO strings or epoch ms).

Key scans (ordered)

db.scan({ gte: 'user:1', lte: 'user:9', limit: 100 }); // range in key order
db.prefix('user:');                                    // prefix scan

Datetime column queries

await db.set('a', { n: 1 }, { dt: { created: '2024-01-01' } });
db.dtColumns();                              // ['created']
db.dtRange('created', { gte: t0, lte: t1 }); // O(log N), [{ key, value, dt, dtValue }]

Unified query

db.query(q) composes key range, dt range, full-text, and a value filter:

db.query({
  key:    { prefix: 'post:' },
  dt:     { created: { gte: t0, lte: t1 } },
  text:   { index: 'body', q: '北京', op: 'AND' },
  filter: { age: { $gte: 18 }, $or: [{ city: 'Paris' }, { city: 'London' }] },
  sort:   { age: -1 },
  project: ['name', 'age'],
  skip: 0, limit: 20,
});

Filter operators: $eq $ne $gt $gte $lt $lte $in $nin $regex $exists $contains $type, plus $and $or $nor $not. Paths use dot/bracket notation ("address.city", "tags[0]").

Indexed dimensions (key / dt / text / a createIndex'd field) are fast; db.query() automatically uses matching equality/range value indexes for simple top-level predicates (including $and), while a value filter with no matching index falls back to a full collection scan, exactly like MongoDB without an index.

await db.createTextIndex('body', { fields: ['bio'] }); // fields optional (default: all strings)
db.search('body', 'hello 世界', { op: 'AND' });         // [{ key, value, score }]

Latin words + CJK unigram/bigram tokenization (no dictionary, zero deps), with TF-IDF ranking.

API

Method Description
MiniDb.open({ dir, valueCodec?, valueMode?, fsyncPolicy?, compactThresholdBytes?, autoCompact?, activeExpireIntervalMs?, recovery?, readOnly?, onLockFail?, maxMemoryBytes?, maxMemoryPolicy? }) Open / create a database
MiniDb.restore(srcDir, destDir, opts?) Restore a db.backup() directory and open it
MiniDb.openOrRebuild(opts, { onRebuild? }) Open; on corruption, discard + reopen empty (cache use). Never deletes a live-locked db
get(key) Decoded value or undefined
set(key, value, { ttl? }) Set; ttl in ms. Resolves per fsync policy
del(key) Delete; returns true if it existed
has(key), size Membership / live key count
mget([keys]), mset([[k,v],...]) Batch read / write
`batch([{op:'set' 'del', key, value?, ttl?, dt?}])`
expire(key, ttlMs), ttl(key) Set / read TTL
createIndex(name, { field, type?, unique?, sparse? }) Secondary index (json codec)
dropIndex(name), listIndexes() Manage indexes
findEq(name, value), findRange(name, opts) Query value indexes
createCompoundIndex(name, { groupBy, orderBy, orderType? }) Compound index (group + order, e.g. workspace + updatedAt)
compoundRange(name, groupValue, opts) Ordered range within a group, O(log N + limit), no full sort
dropCompoundIndex(name), listCompoundIndexes() Manage compound indexes
scan({ gte, gt, lte, lt, limit, reverse }) Ordered key range scan
prefix(p, limit?) Key prefix scan
dtColumns(), dtRange(col, opts) Datetime column range query
query({ key, dt, text, filter, project, sort, skip, limit }) Unified Mongo-like query
createTextIndex(name, { fields? }), dropTextIndex(name) Full-text index
search(name, q, { op?, limit? }) Full-text search
compact() Force a snapshot + WAL rewrite now
backup(destDir, { compact? }) Write a consistent online backup directory
close() Flush, fsync, close

Durability (fsyncPolicy)

Policy Guarantees Relative speed
always fsync after every flush slowest (safest)
everysec fsync once per second (default) fast, ≤1s loss window
no OS decides when to flush fastest, may lose seconds on power loss

Value storage mode (valueMode)

valueMode controls where value bulk lives:

const db = await MiniDb.open({
  dir: './data',
  valueCodec: 'json',
  valueMode: 'memory', // 'memory' | 'disk' | 'auto'
});
  • memory (default): values are kept in RAM, Redis-style. Reads are memory-bound.
  • disk: values stay inline in the durable snapshot/WAL, while RAM keeps only keys, metadata, indexes, and small { file, off, len } value pointers. Reads use synchronous positioned reads, so the public API stays synchronous. This allows value bulk to exceed RAM, at the cost of disk reads on cold get() paths.
  • auto: at startup, compare the current db.snapshot + db.wal size with maxMemoryBytes. If the persisted files exceed the budget, open as disk; otherwise open as memory. If maxMemoryBytes is not set, auto falls back to memory.

Memory limits

With valueMode: 'memory' the budget bounds keys plus values. With valueMode: 'disk' it bounds the in-RAM key/metadata/index footprint, not the value bulk. You can bound writes with an approximate memory budget:

const db = await MiniDb.open({
  dir: './data',
  valueCodec: 'json',
  maxMemoryBytes: 512 * 1024 * 1024,
  maxMemoryPolicy: 'reject', // or 'evict-lru'
});

reject throws when a write would exceed the budget; evict-lru durably deletes least-recently-used keys first. In valueMode: 'disk', evict-lru trims keys to bound the metadata/index footprint; it does not delete value bytes from disk because values are already stored outside RAM. db.stats.evictions and db.stats.maxMemoryRejections track the behavior.

The budget is tracked in approximate logical bytes (key + value + dt metadata), not real JS heap usage. Actual per-key overhead is higher — on the order of several hundred bytes per key for the Map entry, the ordered index node, and bookkeeping — so a database of many small keys needs far more RAM than store.bytes suggests. (Measured: ~0.6 KB of heap per key for tiny records, so 1M small keys ≈ 600 MB of heap.) For millions of small keys, prefer valueMode: 'disk' and/or budget accordingly.

Online backup / restore

await db.backup('./backup');
const restored = await MiniDb.restore('./backup', './restored', { valueCodec: 'json' });

backup() flushes the WAL, optionally compacts, and copies the snapshot, WAL, index definitions, and text postings while writers are briefly parked, producing a consistent directory that restore() can reopen.

Write throughput & SSD endurance

MiniDb's write path is append-only and sequential, which is already SSD-friendly. A few choices have an outsized effect on write amplification and flash wear:

  • Use batch writes. set() resolves after its own flush, so for (const x of xs) await db.set(...) issues one flush (and, under fsyncPolicy: 'always', one fsync) per key. Coalesce writes with db.batch([...]) / db.mset([[k, v], ...]) (single atomic WAL frame) or await Promise.all(xs.map(x => db.set(x))) (group commit collapses the tick into one writev + one fsync).
  • Keep the default fsyncPolicy: 'everysec'. It bounds fsyncs to ~1/s. always is the most durable but the hardest on flash (one fsync per flush); reserve it for data that must survive any crash.
  • Tune compactThresholdBytes for large datasets. Compaction rewrites all live data, so steady-state write amplification is roughly 1 + liveDataBytes / compactThresholdBytes. Raising the threshold (e.g. 256 MiB1 GiB) trades a larger WAL / longer recovery for less rewrite.
  • Observe it via db.stats. walBytesWritten, walFsyncs, snapshotBytesWritten, compactions, evictions, maxMemoryRejections, and queryIndexHits let you measure real write volume, fsync rate, memory pressure, and index usage under your workload.

RESP server (redis-cli compatible)

npm run server -- --dir ./data --port 6379
# or: node --import tsx src/server.ts --dir ./data --port 6379

Then in another shell:

redis-cli -p 6379 SET foo bar
redis-cli -p 6379 GET foo

Supported commands: PING ECHO GET SET DEL EXISTS MGET MSET TTL DBSIZE COMPACT INFO QUIT.

Benchmarks

Measured on Node v24, 100-byte values (npm run bench, N=100000):

Operation Throughput
Raw Map set (baseline) ~8.6 M ops/s
Raw Map get (baseline) ~20 M ops/s
DB get (in-memory) ~8.0 M ops/s
DB set, fsync=everysec (concurrent, group commit) ~725 k ops/s
DB set, fsync=no (concurrent, group commit) ~396 k ops/s
DB set, fsync=always (sequential, 1 fsync/op) ~328 ops/s
Compact snapshot of 100k keys ~69 ms (12 MiB)

Reads are memory-bound and track the raw Map closely. Writes are disk-bound; group commit makes concurrent writes very fast, while synchronous (always) writes pay the fsync cost one would expect from any database. Numbers vary by machine/disk.

Query benchmarks (N=50k docs, 200 iters each, node bench/query.js):

Query Throughput
dt range (small result) ~183 k ops/s
key prefix scan (~100 rows) ~23 k ops/s
full-text search (latin / CJK) ~350490 ops/s
value filter, no index (full scan of 50k) ~23 ops/s

Indexed dimensions (key / dt / text) are fast; an unindexed value filter scans the whole collection — create a secondary index (createIndex) for hot value predicates.

For ordered pagination within a group (e.g. "sessions in a workspace ordered by updatedAt"), use a compound index instead of fetching the whole group and sorting in memory. On 10k sessions in one workspace, paginating with compoundRange is ~2040× faster than "fetch-all + sort", and stays sub-millisecond at any offset.

Multi-process ClusterDb scaling across process counts × shard counts (the headline: shard-affinity writes scale ~linearly until the single-shard speed is reached; uniform cross-shard traffic pays lock handoffs):

pnpm bench:cluster   # bench/cluster.ts — spawns real writer/reader processes

Testing

Three layers of tests, run with the built-in node:test runner (no deps):

npm test            # unit tests (fast)
npm run test:e2e    # end-to-end stability suite
npm run test:all    # both

The unit tests (test/*.test.js) cover each module: frame codec/CRC, WAL group commit, store TTL, snapshot/compaction, recovery truncation, skip list, secondary/full-text indexes, and the RESP server.

The E2E stability suite (test/e2e/*.test.js) covers crash-safety and long-run behavior:

File What it verifies
fuzz-model.test.js thousands of random ops match a reference model (seeded, reproducible)
crash-recovery.test.js kill -9 mid-write and mid-compaction → recovery is always consistent
index-consistency.test.js key/dt/secondary/full-text indexes never drift from the store
compaction-race.test.js heavy concurrent writes during compaction lose nothing
recovery-matrix.test.js WAL corruption at head/mid/tail under resync vs strict
durability.test.js always/everysec/no close-durability + many open/close cycles
boundary.test.js key-length limits, large values, many keys, empty db
soak.test.js sustained ops + heap stability (opt-in: SOAK=30 npm run test:e2e)

The cluster suite (test/cluster/*.test.ts) covers the ClusterDb sharding layer: topology/routing, merged scans, lock contention and lease renewal in-process, cross-shard indexes/compaction, true multi-process scenarios (concurrent writers on disjoint and shared shards, live cross-process read visibility, read/write storms) and crash takeovers (kill -9 → contiguous recovery + stale-lock handoff).

Design in one paragraph

Log-structured engine: all writes append to a CRC-framed WAL (group committed, configurable fsync); all reads hit an in-memory Map. When the WAL grows past a threshold, a snapshot of the live keys is written and the WAL is rotated — non-blocking: writers keep appending to the WAL (which doubles as a Redis-style rewrite buffer) while the snapshot is written, and pause only for a brief final rotation. Recovery loads the latest snapshot then replays the WAL, truncating any torn tail. Secondary indexes use a Redis-style skip list for range queries. See DESIGN_NOTES.md for the full rationale and the source-code study behind each choice.

Concurrency & multi-process

minidb is single-writer. Opening a directory for writing acquires an exclusive lock file (db.lock); a second writer is rejected with a LockError. A lock is taken over only when its owner PID is dead (stale-lock recovery), never merely because it is old.

// second process: throws LockError
await MiniDb.open({ dir: './data' });

// degrade to read-only instead of throwing
const ro = await MiniDb.open({ dir: './data', onLockFail: 'readonly' });
ro.get('k');        // ok (point-in-time view)
await ro.set('k');  // throws "read-only mode"

// open read-only alongside a writer
const r = await MiniDb.open({ dir: './data', readOnly: true });

For many clients, run the RESP server (single minidb process, many TCP clients) — that is the intended concurrent-access model, like Redis.

ClusterDb: multi-process sharding

When several processes must read and write the same logical database without a server, ClusterDb shards the key space over N ordinary minidb directories (each with its own WAL, snapshot, and db.lock):

import { ClusterDb } from '@moonshot-ai/minidb/cluster';

const db = await ClusterDb.open({ dir: './data', shardCount: 16, valueCodec: 'json' });
await db.set('user:1', { name: 'alice' });      // routed by hash to one shard
await db.mset([['a:1', v1], ['b:2', v2]]);      // grouped per shard, atomic per shard
const all = await db.scan({ prefix: 'user:' }); // merged over all shards
await db.close();
  • Writes never meet a single global writer: each process acquires shard write locks on demand (with retry up to lockAcquireTimeoutMs), caches them briefly (lockPoolMaxShards, LRU), and yields a held shard after lockHoldMs (default 250ms) so other processes are never starved. Throughput scales with the number of distinct shards being written — up to the single-shard speed of MiniDb per shard.
  • Reads never take locks: they use the cached writer when the local process holds the shard, else a read-only instance that is revalidated against the shard files on every use, so a read that starts after another process's commit always observes it.
  • Consistency: single-key and same-shard batch ops are strongly consistent (single writer per shard, atomic WAL frames). Cross-shard mset/mdel/batch are best-effort (atomic per shard, not globally; crossShard: 'none' rejects them instead). Scans merge per-shard snapshots, so entries from different shards may reflect different points in time.
  • Indexes (createIndex/createTextIndex) are recorded in a cluster-wide registry and applied by every shard writer on open; findEq/findRange/ search merge per-shard results (text scores are per-shard). Index management acquires every shard writer — run it from one process, off the hot path.
  • Crash recovery is per shard: a lock left by a dead PID is taken over by the next opener, exactly like single MiniDb.

crossShard: '2pc' is reserved for a future two-phase commit and is rejected today. Performance numbers across process/shard counts: run pnpm bench:cluster (see bench/cluster.ts).

For a rebuildable cache, use openOrRebuild: a corrupt cache is discarded and reopened empty, while a live-locked db is never destroyed:

const db = await MiniDb.openOrRebuild(
  { dir: cacheDir, valueCodec: 'json' },
  { onRebuild: (err) => log.warn('cache rebuilt:', err.message) },
);

Caveats / roadmap

  • In the default valueMode: 'memory', the dataset must fit in RAM Redis-style. Use valueMode: 'disk' for larger-than-RAM value bulk; cold reads then perform synchronous positioned reads against the snapshot/WAL.
  • Full-text index postings are stored on disk (larger-than-RAM); only the term dictionary and per-doc metadata stay in memory. Postings are rebuilt from the store on open and on compaction. Search reads postings synchronously, so a cold, very large postings list can briefly block the event loop.
  • Compaction is non-blocking for writes: the WAL itself acts as a BGREWRITEAOF-style rewrite buffer, so the (slow) snapshot is written while writers keep appending. Writes pause only for the final rotation (a flush, a tail copy, and two atomic renames). A pre-copy drains most of the tail beforehand when writes are slow enough; under sustained writes that outrun the pre-copy, the rotation absorbs a larger tail — the same bounded end-of-rewrite pause Redis accepts for its AOF diff flush — so compaction always terminates. Mid-compaction crashes leave db.*.tmp files behind, which the next writer open removes automatically.
  • Snapshot encoding runs on the main thread (chunked + yielding); offloading to a worker_thread is a planned optimization.
  • Single minidb directory = single process / single writer. For multi-process access use ClusterDb (above): sharding scales writes, but hash routing means whole-range scans fan out to all shards, and uniform cross-shard write patterns pay shard-lock handoff costs (lockHoldMs per yield) — workloads with per-process shard affinity scale best.

Credits

Design distilled from reading: Redis (references/redis), the SQLite WAL paper, NeDB (references/nedb), Bitcask (references/bitcask), and the cstack SQLite tutorial (references/db_tutorial).