ruvector/crates/ruvllm/examples
OceanLi 34390efe56
feat(ruvllm): add lattice as an optional macOS Metal LlmBackend (#642)
* feat(ruvllm): add lattice as an optional macOS LlmBackend

Adds LatticeBackend, a pluggable LlmBackend implementation over
lattice-inference's pure-Rust Qwen3.5 Metal GPU forward pass, gated behind a
new default-OFF `lattice` feature (macOS-only: dependency under
[target.'cfg(target_os = "macos")'.dependencies], module gated
#[cfg(all(feature = "lattice", target_os = "macos"))]).

- MetalQwen35State (!Send) is owned by a dedicated worker thread, mirroring
  lattice_serve.rs's spawn_worker/run_worker_loop pattern, but over plain
  std::sync::mpsc (TokenStream is std-mpsc-backed).
- generate_stream_v2 streams every real decoded token via
  generate_streaming_with_cancel, unlike candle's prefill-only stream stub.
- get_embeddings returns RuvLLMError::NotImplemented (honest, per ratified
  O1) rather than a fake zero vector.
- create_backend() precedence: lattice (if enabled) > candle > NoopBackend.

Root Cargo.toml carries an uncommitted dev-only [patch.crates-io] pointing
lattice-inference at a local checkout; not included in this commit.

* fix(ruvllm): enforce stop strings + reject unsupported penalties in LatticeBackend

Codex round-1 fixes:
- MAJOR 1: lattice's Metal generation loops honor EOS/stop_token_ids but not
  GenerateConfig::stop_strings, so callers' stop_sequences were silently
  ignored. Added StopScan: incremental stop-string scanner that holds back
  the longest possible stop prefix (char-boundary safe), excludes the matched
  stop from output, and halts generation through the token callback. Both
  generate (via the streaming loop, so a match actually stops decode) and
  generate_stream_v2 route through it; no stop strings = zero-overhead path.
- MAJOR 2: frequency_penalty/presence_penalty are live ruvllm fields
  (serving/engine.rs:547, mistral_backend.rs:907), not dead ones; nonzero
  values now fail fast with NotImplemented instead of being silently dropped.
- MINOR 3: em dashes removed from all added lines (repo prose lint).
- 6 non-GPU unit tests: StopScan cut/holdback/multi-stop/UTF-8 + penalty
  rejection on both entry points.

* chore(ruvllm): bump lattice-inference to 0.5

* fix(ruvllm): adapt LatticeBackend to lattice-inference 0.5 Result APIs

generate and generate_streaming_with_cancel return Result in lattice
0.5; propagate failures as RuvLLMError::Backend on the once path and
StreamEvent::Error on the stream path instead of unwrapping.

* bench(ruvllm): add lattice_bench example, reproducible backend throughput harness

Measures load time, TTFT, and decode throughput for the lattice backend
(stream and blocking legs), with a BENCH_GREEDY env toggle so results can
be compared against greedy standalone-engine numbers using the same
prefill-canceling slope method. The candle backend is timed via blocking
generate() only; its generate_stream_v2 emits a single token from prefill
logits and is not a decode loop. Feature-gated: builds as a stub without
the lattice feature.

* docs(ruvllm): model-prep guide for lattice_bench + rustfmt

The bench doc header now walks through obtaining a runnable model dir:
f16 safetensors straight from HuggingFace, or quantizing with lattice's
quantize_q4 and copying tokenizer.json + config.json next to the .q4
output (the quantizer writes weights only). Documents all flags and the
BENCH_GREEDY toggle. README points to it from the lattice section.
Also applies rustfmt to lattice_backend.rs (import order, comment
alignment).

* fix(ruvllm): derive safetensors precision label from torch_dtype, not hard-coded Bf16

load_worker_state stamped every safetensors checkpoint as Quantization::Bf16,
so an f16/f32 checkpoint got a false precision label in ModelInfo and a wrong
bytes_per_weight in the num_parameters estimate. Read torch_dtype from the
already-open config.json instead — the same honesty guard lattice_bench.rs
applies to the candle side — falling back to Bf16 (the Qwen3.5 release dtype
and the previous fixed label) when the field is missing or unmapped, since a
label must not fail a load that from_safetensors already accepted.

Verified on macOS arm64 (M4, Metal): cargo test -p ruvllm --features lattice
green, including the new safetensors_precision_label_follows_torch_dtype test.

Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_017sXWL4ox5bhC86FYwJpmyK

---------

Co-authored-by: ruvnet <ruv@ruv.net>
2026-07-05 11:10:37 -04:00
..
benchmark_model.rs fix: update pgrx to 0.12.9 in both CI workflows and fix formatting 2026-02-21 22:34:37 +00:00
download_test_model.rs fix: update pgrx to 0.12.9 in both CI workflows and fix formatting 2026-02-21 22:34:37 +00:00
generate_claude_dataset.rs fix: update pgrx to 0.12.9 in both CI workflows and fix formatting 2026-02-21 22:34:37 +00:00
hub_cli.rs fix: update pgrx to 0.12.9 in both CI workflows and fix formatting 2026-02-21 22:34:37 +00:00
lattice_bench.rs feat(ruvllm): add lattice as an optional macOS Metal LlmBackend (#642) 2026-07-05 11:10:37 -04:00
run_eval.rs fix: update pgrx to 0.12.9 in both CI workflows and fix formatting 2026-02-21 22:34:37 +00:00
train_contrastive.rs fix: update pgrx to 0.12.9 in both CI workflows and fix formatting 2026-02-21 22:34:37 +00:00