Commit graph

7 commits

Author SHA1 Message Date
Helldez
557c63dda0 docs(assets): label the Qwen3.8-Flash-Next hero clip like the DeepSeek one
Some checks are pending
ci / changes (push) Waiting to run
ci / format (push) Blocked by required conditions
ci / host-linux (push) Blocked by required conditions
ci / host-windows (push) Waiting to run
ci / host-macos (push) Waiting to run
ci / android-apk (push) Waiting to run
The refreshed clip went in unlabelled, so the frame said nothing about which
model was running while the DeepSeek hero next to it names its own. Same
treatment now: a black band over the app title bar carrying the model and
quantization, "running on a phone, airplane mode" under it, and a second band
over the navigation bar with the real-time note and the run's tok/s.

Typography is measured off the DeepSeek asset rather than guessed - Arial 14 for
the subtitle (the same string lands on the same pixels, x 72 to 287), Arial Bold
16 for the model name, Arial 15 for the caption, and its purple. The palette is
built with stats_mode=full: on static overlay text the diff mode spends no
colours and quantizes white down to 249.

The caption reads "real time" rather than the DeepSeek clip's "from here on -
real time", which was there because the master it was cut from had sped-up
stretches ahead of that point. This clip is real time end to end.
2026-08-30 17:35:24 +02:00
Helldez
33b9fcf143 docs(readme): refresh the Qwen3.8-Flash-Next hero, 3.48 tok/s on the Q2_K build
The clip in the README was the UD-IQ3_XXS file at 2.03 tok/s. This is a new
in-app run on the 12 GB test phone: 82 tokens at 3.48 tok/s, real time, with
prefill 6.87 s (3.9 tok/s), 13487 MB streamed, a 1000/1000 MiB expert cache at
56 % hit, 70 major faults per token, pinned dense weights, overlap on, four
lanes, cold experts dropped at 100 %.

The file is the DevQuasar Q2_K build the catalog gained in #189, not the one the
old clip used, so the caption names the quantization and the size follows it:
six shards, 80,447,449,856 bytes. The two builds are the same size on disk to
within 2 %, because the 28.8 GB n-gram table dominates both and stays at IQ4_NL
in each. What differs is the dense side — 2.4 GB pinned against 4.3 — which is
the cache room the app README already describes.

The architecture table carried two stale claims: a single ~2 tok/s figure that
belonged to UD-IQ3_XXS alone, and a note that the model was pinned to an
unmerged upstream PR. Upstream support has been in since b10666.
2026-08-30 16:56:11 +02:00
Helldez
927e2d3b31
feat(moe): Qwen3.8-Flash-Next support (#172)
Qwen3.8-Flash-Next (qwen4exp): 125B total, ~6B active, 512 routed experts at
top-10 plus one shared, 48 hybrid gated-delta SSM / sparse attention layers,
and a 51B n-gram embedding table. One registry row streams the experts; a
dense-policy guard keeps the n-gram table (larger than any phone's RAM)
mmap'd under every mode so pinned and anonymous dense weights survive load.
Runs on the 12 GB test phone at ~2 tok/s with pinned dense weights, compute-
bound, and sits in the app catalog as a three-shard download. Submodule
pinned to upstream master b10666, the first with the architecture merged,
with the expert-ready hook on top. README hero clip, changelog and docs
updated. App 0.22.0 (versionCode 37).
2026-08-28 10:07:43 +02:00
Helldez
0091a90a48
feat(app): settings grouped by purpose, and four defects an audit found (#151)
* feat(app): settings grouped by purpose, and four defects an audit found

Settings now show the recommended configuration first and fold everything else
into a collapsed Experimental group per category. That is a statement about
evidence, not about how finished the code is: inside are the levers measured on
one device, measured once, or still owed a measurement. They stay in the
release build, because testing them on other hardware is what this app is for
and a lever nobody can reach is a lever nobody can refute. The caveat is stated
once in the group header instead of leaking into some descriptions and not
others.

Every description was rewritten to say what the setting does for the person
reading it. Out went the measured figures, which need the device, the model and
the day beside them to mean anything and have none of that room under a switch,
and out went the implementation names: O_DIRECT, top-k, dma-buf, mmap and KV
cache are not what someone deciding whether to turn something on needs to know.
The metrics screen keeps the flag names, deliberately: there the reader is
matching the UI against a CSV column and the technical name IS the vocabulary.

Four defects, all found by auditing rather than by anything failing:

The session signature is now derived from the argv instead of being a
hand-written list beside it. Those two had to be kept in step with nothing
enforcing it, and forgetting a field is a silent bug: the setting appears to
change while the engine keeps running the old configuration. Three of four
rebases this week collided on exactly that list.

A malformed end-of-turn summary no longer strands the UI. The whole handler sat
inside a catch with no failure branch, so a parse error left the state in
GENERATING with no turn committed and nothing said. The streamed answer is now
kept, the reason is shown, and the state returns to READY.

MainActivity drops from about 1050 lines to under 700: the model download and
import UI moves to ModelPickerUi.kt, which shares nothing with the chat screen.
No logic moved, only its address.

Dead code removed: a field whose own comment described a use it did not have,
two functions nobody called, and five string resources describing a UI two
rewrites ago.

* docs: record the settings regrouping and the signature fix

Rule 6: the changelog and the docs a change invalidates ship with it. The app README described the settings screen as it was before the regrouping, and explained one experimental lever in terms of predictor accuracy percentages that the UI no longer shows.

* docs: the improved DeepSeek hero recording

* feat(app): keep the flag vocabulary in Settings, and make Experimental read as a boundary

The first pass at rewriting the descriptions went too far: it renamed the controls into consumer phrasing and lost the vocabulary that lets a setting here be matched against the CLI, the CSV preamble and the docs. Labels are back to the flag's own names; the descriptions are shorter than the originals rather than longer, and still carry no measured figures. The Experimental group gets a divider and a tonal bar: collapsed, it is the only thing between the recommended configuration and the levers that can change the reply, so it has to look like a boundary rather than one more row.
2026-08-02 01:50:15 +02:00
Helldez
5f19289288
docs: DeepSeek V4 Flash leads the README, and the gate list stops lying (#150)
The flagship demo is now the 284B model generating on a 12 GB phone at about
1 tok/s, with its own recording, instead of gpt-oss carrying that slot. The
quoted 0.94 tok/s is the app's own reading and the sentence next to it says
what produced it: cold-expert dropping at full strength, which trades quality.
gpt-oss keeps its lossless and knob-on figures one paragraph down, and the
three-model clip stays where it was.

The model is named by its exact release, 0731, everywhere it appears rather
than only in the opening line. Anyone reproducing this needs to know which
DeepSeek V4 Flash it was.

The list of what ctest gates claimed the LRU cache, evictions, overlap,
temporal prefetch, the dense rebind and multi-turn sessions. It omitted
predictive prefetch and the split multi-shard model, both of which are gated,
and it did not mention that a lossy knob only gets machinery gates. Fixed to
match what the suite actually runs today.

Also: docs/ngram.md was missing from the documentation index despite being a
250-line document for a shipped flag, and the methodology caveat carried an
editorialising aside about future storage in the one paragraph that has to
stay dry.
2026-08-02 01:15:33 +02:00
Helldez
643e5ee2a7
docs: professional README and canonical AGENTS.md (#132)
* docs(readme): reposition around large MoE models in general, add logo and TOC

The front page led with gpt-oss-120b and read as a single-stunt repo. It now
leads with what the engine is for: running MoE models past RAM on phones and
PCs, on llama.cpp's public API, with every llama.cpp quantization coming for
free. The 120B stays as the most extreme proof, not the pitch.

Structure follows the usual professional layout: logo (chip-buddy, light and
dark variants under docs/assets/logo/), badges, table of contents, a "Why this
exists" section that covers all three regimes (far past RAM, just past RAM,
and models that barely fit, which streaming keeps inside a chosen budget),
and Features grouped the way the app's Settings groups them. Benchmark tables
and the evidence sections are unchanged in substance; the test device is
phrased generically.

* docs(agents): make AGENTS.md the canonical agent guide, CLAUDE.md points to it

The agent guide lived in CLAUDE.md with AGENTS.md as a stub pointing at it,
which is backwards: AGENTS.md is the cross-tool convention (Codex, Copilot,
Cursor and others read it), CLAUDE.md is one tool's name for the same file.
The full guide now lives in AGENTS.md and CLAUDE.md is a one-line import.

While moving it, the guide gains the release rules that were only tribal
knowledge: release APKs come from the release-apk workflow, never a local
build; every released feature bumps versionCode/versionName in the same PR;
release titles are the bare version; on-device validation before a release.
The style section now names the CI clang-format version (18), since a
mismatched local formatter passes locally and fails the check.

* docs(readme): state per-platform host status honestly

Linux is CI-verified, Windows is where the desktop bench ran (but needs CMake
directly, not the bash script, and MSVC's Release output path), macOS builds
from the same sources but is unvalidated and has no O_DIRECT.
2026-07-28 17:40:36 +02:00
Helldez
e90149b279 docs(readme): launch-ready front page (result banner, hero GIF, quickstart, gate proof) 2026-07-19 19:14:02 +02:00