Commit graph

129 commits

Author SHA1 Message Date
Helldez
022758e467 docs(android): list Qwen3.6 in the catalog, and separate sweep from default
The app README enumerated a two-model catalog; Qwen3.6-35B-A3B has shipped
in it since 0.11.0. Qwen3-30B is labelled to match what the app displays,
the release-signing prerequisite is stated so a published build is
reproducible, and the 4000 MiB in the expected-numbers section is marked as
a sweep point rather than the shipped default, which is a fixed 2000 MiB.

prefetch.md and the config.h comment both said prefetch needs
--cache-mb > 0; validate() also accepts --cache-mb auto.
2026-07-19 11:27:44 +02:00
Helldez
fa4712e6c4 docs(cache): correct the stale 'tracks free memory' claims about --cache-mb auto
The runtime governor was retired; auto sizes the budget once at load and
holds it. The intro and the Android note still described the old control
loop, contradicting the body of the same document.
2026-07-19 10:50:38 +02:00
Helldez
d5d400a0c9 chore(release): 0.11.0 (versionCode 18)
Add Qwen3.6-35B-A3B (qwen35moe) to the model catalog.
2026-07-19 10:11:07 +02:00
Helldez
56acb3bed7 feat(android): add Qwen3.6-35B-A3B to the model catalog
Qwen3.6-35B-A3B reports arch qwen35moe, already in the registry: a hybrid
attention/SSM MoE (256 experts, top-8, 41 blocks) whose routed experts
stream through the existing qwen3moe path unchanged. Add it to the in-app
one-tap download catalog and document it in the README (Supported models
row + a benchmarks subsection). Validated on-device at ~2x RAM: the mmap
baseline collapses (0.1 tok/s) while streaming runs stably at 5.0 tok/s
lossless (5.8 with turbo top-k).
2026-07-19 10:11:07 +02:00
Helldez
af6c7d8ffe chore(release): 0.10.0 (versionCode 17)
Layer-granularity compute trace (--compute-trace-layers, #65) and the
reviewed defaults (#71): n_predict 128 on every surface, fixed 2000 MiB
expert cache in the app.
2026-07-19 09:34:23 +02:00
Helldez
eb455bc4e6 feat(defaults): align n_predict at 128 and fix the app expert cache at 2000 MiB
The generation default diverged across surfaces (CLI 32, app 48) and both
budgets truncate most answers mid-sentence, which reads as broken rather
than slow on a first run. 128 is the smallest choice that lets a typical
answer finish; CLI and app now share it so there is one documented default.

The app's expert cache moves from Auto (ceil 3000) to a fixed 2000 MiB:
Auto sizes to whatever RAM happens to be free at load, so first
impressions and benchmarks varied with unrelated device state. 2000 sits
above the engine's 1500 MiB floor (no --force-cache) and Auto stays
selectable for devices where a fixed budget is wrong. Existing installs
keep their saved prefs; only fresh installs see the new defaults.

Closes #71.
2026-07-19 09:23:22 +02:00
Helldez
51b5a464ed chore(release): 0.9.2 (versionCode 16)
Ships the reasoning-visibility fix (#70): a thinking model's reasoning is shown
in a collapsible block instead of being stripped from the answer. Validated
on-device.
2026-07-19 08:40:44 +02:00
Helldez
1d4de61420 feat(android): show a reasoning model's thinking in a collapsible block (#70)
Consume the new `reasoning` protocol field so a Thinking-on run is legible
while the model reasons instead of sitting on a blank answer.

  - Telemetry parses `reasoning`; RunBus carries an in-flight `reasoning` span
    and ChatTurn keeps a per-turn reasoning field; RunService streams it live
    and commits it with the finished turn (falling back to the last streamed
    span if BMOE_DONE omits it).
  - The transcript renders reasoning as a dimmed, collapsible "Thinking" block
    above the answer — open while it streams, collapsed once the turn is
    committed. The live turn now shows during the thinking phase (before any
    answer text), so a slow reasoning decode no longer looks stuck.

Also make the Thinking setting description model-agnostic: it no longer names
specific model families, describing the behaviour (reasoning shown in a
collapsible block; off skips thinking; no effect on non-reasoning models)
instead.
2026-07-18 22:43:34 +02:00
Helldez
7d7108de2d build(android): sign release builds with a stable keystore (0.9.1)
Sideload releases were debug-signed. A debug key is generated per build machine, so
every published APK had a different signature and Android refused to update in place
(INSTALL_FAILED_UPDATE_INCOMPATIBLE), forcing an uninstall that wiped the models in
filesDir. Release builds now sign with a stable keystore whose secrets live in
keystore.properties (gitignored, never committed); when it is absent the release
falls back to debug signing so a fresh clone still builds. The distributed artifact
is app-dev-release.apk. Bumps to 0.9.1 (versionCode 15).
2026-07-18 20:25:28 +02:00
Helldez
8b22f1bd9f chore(release): 0.9.0 (versionCode 14)
In-app model downloads to O_DIRECT-capable internal storage (#67).
2026-07-18 20:02:14 +02:00
Helldez
172d81cae1 fix(android): keep the download's MiB counter visible and clear the finished bar
The progress line put the filename and the MiB counter on one ellipsized line, so a
long model name (e.g. a Hugging Face Qwen filename) pushed the numbers off the end as
"…". Put the name on its own line above the counter. Also clear the progress map when
no download is in flight, so the last bar doesn't linger after a cancel or completion.
2026-07-18 20:00:13 +02:00
Helldez
565b854031 fix(android): make the WorkManager download run and resume on-device
Three defects found testing the download on a Pixel-class OxygenOS device:

- WorkManager runs the foreground transfer through its own SystemForegroundService,
  which it declares with no foregroundServiceType. On API 34+ the type passed to
  startForeground (dataSync) must be a subset of the manifest declaration, so the
  first download crashed with "foregroundServiceType 0x1 is not a subset of 0x0".
  Merge dataSync onto that service in the manifest.

- Resume restarted from zero: HttpURLConnection's automatic redirect drops request
  headers across a cross-host 3xx (exactly Hugging Face's resolve -> CDN hop), so the
  Range header was lost, the CDN returned 200, and the .part was truncated. Follow
  redirects manually, re-applying Range and identity encoding on every hop.

- Cancel raced: cancelUniqueWork is async, so the immediate activeDownloads() reseed
  read the work back as still-active and the row stayed "downloading". Block on the
  cancel operation before reseeding, matching enqueue.
2026-07-18 20:00:13 +02:00
Helldez
35c63accfe docs(android): note in-app downloads now use O_DIRECT internal storage
The model-acquisition section and the manifest/ModelManager comments described
downloads landing on the app-specific external dir; they now land in filesDir
where O_DIRECT works, and a download needs no temporary second copy.
2026-07-18 18:21:28 +02:00
Helldez
9dcb250eee feat(android): download models with WorkManager into internal storage
In-app downloads went to the app-specific external files dir (getExternalFilesDir),
a FUSE volume where O_DIRECT is silently broken: pread reports success without
filling the buffer, so the engine's verify catches it and falls back to buffered
I/O — losing the streaming advantage exactly on the >RAM models it exists for.
DownloadManager cannot write to internal storage by design (system process,
external-only), so it is replaced.

DownloadWorker is a foreground CoroutineWorker that downloads over HTTP straight
into filesDir/models (real f2fs/ext4, where O_DIRECT works — the same dir SAF
import uses), with a Range request to resume an interrupted .part instead of
restarting a multi-GB transfer, and rename-on-complete so a partial file is never
listed. This mirrors Google's AI Edge Gallery downloader (HttpURLConnection +
Range + WorkManager foreground worker). It also removes the 2x free-space cost of
a download-then-copy workaround: a model needs free space equal to its own size.

ModelDownloader keeps its State/Progress contract and enqueue/query/cancel shape,
now backed by unique WorkManager work keyed by filename, so the UI change is small
and in-flight transfers stay re-discoverable after process death.

Fixes #67
2026-07-18 18:21:20 +02:00
Helldez
baefcb8da7 chore(app): bump to 0.8.3 (versionCode 13) 2026-07-18 16:53:13 +02:00
Helldez
ee9ea072cf fix(android): dismiss keyboard on send and keep the answer clear of the IME
The prompt field is multiline, so it exposes no Done action, and the activity
declared no windowSoftInputMode with no imePadding in the layout: once the soft
keyboard opened there was no in-app way to close it, and the streaming answer
drew behind it (issue #56).

Declare adjustResize, pad the content Box for the IME (covers the edge-to-edge
path on Android 15+ where the window no longer shrinks), and clear focus on Send
so the keyboard retracts into the space the answer streams into.
2026-07-18 16:35:23 +02:00
Helldez
2c710c8f10 chore(app): bump to 0.8.2 (versionCode 12) 2026-07-18 16:10:31 +02:00
Helldez
a89814dfbc
Merge pull request #61 from Helldez/feat/52-model-delete
feat(app): delete on-device models from the catalog (#52)
2026-07-18 16:08:44 +02:00
Helldez
a337edc25a
Merge pull request #60 from Helldez/fix/50-metrics-glossary
fix(app): correct the metrics glossary and three telemetry miscounts (#50)
2026-07-18 16:08:40 +02:00
Helldez
6e25dec76e feat(app): delete on-device models from the catalog (#52)
The catalog could download a model but never remove one, so a user who pulled
two or three multi-GB ggufs had no in-app way to reclaim the storage. Models
brought in by pasted URL or the file picker had no per-row surface at all.

Add a Delete affordance to every on-device model — the catalog's ON_DEVICE
rows and a new "Imported models" list for non-catalog ggufs (URL / SAF). Both
open one Material3 confirm dialog that:

- lists every on-disk copy the name resolves to across the scanned dirs, with
  its size (the same gguf can sit in two places at once — adb-pushed AND
  downloaded — and deleting one must not silently orphan the other);
- flags copies the app cannot remove: /data/local/tmp/bmoe is shell-owned, so
  those are shown with the adb command to remove them instead;
- blocks deletion of the currently loaded model (its gguf is mmap'd by the live
  engine), telling the user to start a new chat first.

On confirm it deletes the app-deletable copies off the main thread and triggers
the existing rescan, so the catalog status and the model picker both refresh.
ModelManager gains copiesOf() (all copies of a name) and isAppDeletable().
2026-07-18 10:59:58 +02:00
Helldez
12f3b42339 fix(app): correct the metrics glossary and three telemetry miscounts (#50)
The glossary and live panel described numbers the engine does not compute.

Text (MetricFields.kt, AppSettings.kt):
- majflt_mib is majflt x the runtime page size, not a hardcoded 4 KiB (16 KiB
  on large-page devices).
- dense_resident_frac is a 256-page mincore SAMPLE refreshed every few tokens,
  not the exact fraction; -1 means "not sampled" only before the first sample,
  with streaming off, or when /proc is unreadable (not on most rows).
- mgmt_ms also folds in the dense-residency probe.
- compute_ms is wall_ms with streaming off, and is clamped at 0.
- cache_hit_pct includes prefill lookups, which dominate early.
- rss_anon/rss_file: under the default 'anon' dense policy the dense weights are
  anonymous, not file-backed, so "anon = the cache / file = the model" inverts.

Code:
- flash wait (MainActivity): show the MEASURED wall-additive read term —
  stall_ms under overlap, io_ms in serial — and derive compute as the leftover,
  so the clamp lands on compute instead of the old wall-compute-mgmt silently
  dumping a clamped-to-0 compute into "flash wait". Adds io_ms/stall_ms to the
  parsed telemetry.
- CPU occupancy: the CPU numerator is whole-process, so under overlap the I/O
  lanes must be in the denominator too (threads + io lanes), or it reads >100%.
- MetricsCsv.column(): map the -1 "not sampled" sentinel to null for
  cache_hit_pct and dense_resident_frac, so it stops dragging min/mean/max and
  pearson (both index-aligned, so null preserves pairing). Genuine -1s in other
  columns are untouched.
2026-07-18 10:55:26 +02:00
Helldez
b5ff12eb67 fix(app): keep all metric rows reachable in the CSV view (#53)
The CSV view laid the run-config card and one summary card per turn in a
non-scrolling Column above the token LazyColumn. Each turn appended another
summary, so the pinned block grew until it squeezed the LazyColumn and the
last token rows fell off-screen with nothing able to scroll them into view.

Make the whole view one LazyColumn: the config card, the per-turn summaries
and the caption are now lazy items that scroll away with the content, and the
column header pins via stickyHeader so the labels stay visible. The shared
horizontal scroll state moves onto each Row (header and every data row) so the
columns still scroll sideways together.
2026-07-18 10:47:52 +02:00
Helldez
1dcbe94bad feat(android): default dense weights to anon (O_DIRECT)
On a >RAM model the anon policy reads the dense weights via O_DIRECT into private
buffers, so a reclaim swaps them to zram (fast) instead of refaulting from flash,
and an expert cache finally has room to earn its budget — measured 0.998 vs 0.711
tok/s on gpt-oss-120b with a 2000 MiB cache. Warm stays a click away for models
that fit in RAM, where its page-cache prefetch is the better trade.

The fresh-install branch in load() now defers to the field default instead of
hardcoding the policy a second time.

versionCode 11, versionName 0.8.1.
2026-07-17 18:30:44 +02:00
Helldez
ebfeb1c2d8 fix(android): report a download refusal in the row that caused it, in GB
Tapping Download on a model too big for the device quoted the row as ~17.0 GB
and then refused it for '15 GiB' — both true, and together they read as a bug.
Every user-facing size now goes through one formatter, in the decimal GB the
model repositories quote.

The message also appeared at the foot of the card, under the unrelated 'Other
model' heading, a screen away from the button that triggered it. Catalog
failures now render against their own row.
2026-07-17 15:32:08 +02:00
Helldez
3e28c4866b fix(android): prefer the O_DIRECT copy when a model exists twice
Dedup-by-name keeps the first hit, and the scan order put /sdcard/Download ahead
of /data/local/tmp — so on a device holding the same gguf in both places (which
is the normal state after adb-pushing a model and also downloading it) the app
would silently switch to the emulated-storage copy, where O_DIRECT does not work
and the engine falls back to buffered I/O. A benchmark would just get slower with
nothing in the diff to explain it.

Real filesystems now come first in the scan order.
2026-07-17 13:44:37 +02:00
Helldez
43ed4b582b fix(android): stop the catalog rows wrapping, and let the card be closed
The blurb sat in a weighted column beside the action button, so a sentence had
only the leftover width and came out as a stack of one-word lines. Name and
action now share the top line, the description gets the full card width below it,
and the texts are short enough to fit. The gpt-oss install recipe scrolls
sideways instead of wrapping — a wrapped shell line is a broken shell line.

The card also collapses now. It decides its initial state once the first scan
lands rather than before it: open on a first run, out of the way on a device that
already has a model, and the user's choice from then on. While a download runs
the closed header still says so, so nothing disappears silently.
2026-07-17 13:12:56 +02:00
Helldez
88562873dd fix(android): declare INTERNET, without which no download ever ran
DownloadManager needs android.permission.INTERNET both to run a download and to
report on one, and no manifest declared it. The URL downloader has therefore
never worked: enqueuing threw a SecurityException the UI swallowed as a failed
Result, so the feature looked present and silently did nothing. adb push was the
only path that ever delivered a model.

Seeding the catalog's state from DownloadManager on first composition turned that
silent failure into a crash on launch, which is how it surfaced.

Also stop trusting the download provider not to refuse a query: both query() and
activeDownloads() now swallow it. This state is a progress bar; it must never be
able to take the screen down with it.
2026-07-17 12:43:23 +02:00
Helldez
f45552ce3d build(android): bump to 0.8.0 (versionCode 10) 2026-07-17 12:43:23 +02:00
Helldez
fd049973f5 docs(android): document the catalog, the gpt-oss merge and the dir rename 2026-07-17 12:43:23 +02:00
Helldez
2ededf1d7a feat(android): built-in model catalog with one-tap downloads
Getting a first model meant knowing which MoE to look for and pasting a raw
Hugging Face link. The Get-a-model card now offers the models this engine is
measured on — Qwen3-30B-A3B-Q4_K_M and Gemma-4-26B-A4B-it-Q4_K_M — as a single
tap each, with size, free-space check before enqueuing, and per-entry progress.

gpt-oss-120b is listed but not downloadable in-app: Hugging Face ships that
quant as two shards, and expert streaming reads tensors by byte offset from one
file. The entry carries the PC-side merge recipe instead; once the merged file
reaches the device it is recognized like any other.

Downloads are keyed by filename and seeded from DownloadManager, so a multi-GB
transfer started before the app was killed is picked back up rather than left
running unseen. Arbitrary URLs and the file picker stay, under Other model.

The catalog is UI convenience only — nothing about these models reaches the
engine, which still discovers every architecture at runtime.
2026-07-17 12:43:23 +02:00
Helldez
33ac8f97ff refactor(android): rename the dev shared model dir shardllm -> bmoe
The directory carried a name from an earlier project. It is a plain rename: the
constant, the three benchmark script defaults (all already env-overridable), and
the comments. Recorded measurements under docs/bench-data keep the old path —
they say where those runs actually read their model from.

Devices with models already pushed there:
  adb shell mv /data/local/tmp/shardllm /data/local/tmp/bmoe
2026-07-17 12:43:23 +02:00
Helldez
4269155654 refactor(android): split MetricsScreen, sweep the governor's leftovers
MetricsScreen.kt had grown into a god-file: a CSV parser, file IO, a regex, the
Pearson math, a share intent and five views in 755 lines. The data layer moves to
MetricsCsv.kt (Csv, runTitle, fmt, pearson, shareCsv) and the screen keeps the
views — nothing in it touches a file now. 755 -> 655 + 125.

Governor-era leftovers, from a time when a governor moved the cache budget under
memory pressure:
  - compare defaulted to cache_budget_mib, "the column this whole feature was
    built to look at". The budget is fixed for a run now, so that default opened
    the view on a flat line. Defaults to cache_hit_pct.
  - correlate defaulted to {majflt_mib, rss_anon_mib, cache_budget_mib}; the third
    is now dense_resident_frac, which is what tells "the faults are the model"
    apart from "the faults are the cache". (Not cache_resident_mib: that lives in
    the summary trailer, not the per-token columns, so it would have silently
    fallen through to the first-two-columns fallback.)
  - Telemetry.cacheBudgetMib was documented as "possibly auto-adapting"; MetricFields
    already said "fixed for the run". Now they agree.
  - "was warm-dense on in both?" predates the DenseWeights enum: dense=warm.

Dead live-telemetry fields, parsed off the wire and never rendered: ioMs,
avgIoMs, denseResidentFrac. io_ms in particular is not an oversight — the panel
derives flash wait as wall - compute - mgmt on purpose, because under overlap the
read time is not a wall component (MainActivity: "instead of asking the user to
reconcile a compute residual with a parallel flash-I/O figure"). Dropping the
fields matches that decision.

Also: the n_predict default was literal 48 in both AppSettings and the service's
intent fallback (now AppSettings.DEFAULT_N_PREDICT), and the two charts each had
their own copy of the null-gap-skipping path builder (now DrawScope.drawSeries,
with the y-range choice left at the call site — that choice is what makes the two
charts different).

APK compiles (compilePlayDebugKotlin --rerun-tasks clean); UI-only, no engine change.
2026-07-17 10:33:16 +02:00
Helldez
75b5a634fc refactor(moe): retire the adaptive cache governor; fixed LRU + one-shot auto
Measured net loss on >RAM models (cache-off is the ceiling), so the runtime governor, --cache-dynamic/--cache-gov2, and the sense/resize loop go. Kept: the fixed --cache-mb N LRU, --cache-mb auto sizing once at load, cache-off default. Telemetry aligned (dense_resident_frac made live; resident_frac/cache_cuts removed). App: one 3-way dense-weights selector. Host+ctest green.
2026-07-17 08:55:18 +02:00
Helldez
e78fb78194 feat(moe): --dense-weights — dense (non-expert) weight residency policy
Read each dense weight once via O_DIRECT into an anon buffer and rebind (anon), page-cache them at load (warm), or leave them mmap'd (mmap). RouterHook captures the non-expert weight leaves; the streamer reads+rebinds and drops the mmap copy. Byte-identity gates G6/G7.
2026-07-17 08:55:17 +02:00
Helldez
9a1d1f8a1c feat(metrics): on-device memory telemetry, pressure sensing, and an adaptive cache governor
The measure-your-own-memory series: fault/CPU decomposition, the anon/file RSS split, mincore residency sensors, the Android metrics screens, and a --cache-dynamic governor that sizes the expert cache to what the device concedes. (The governor is retired further down this history; the telemetry stays.)
2026-07-17 08:54:06 +02:00
Helldez
4bd9cb24e9 chore(release): 0.7.0
versionCode 9, versionName 0.7.0. The 0.6.1 that sat in build.gradle was never
tagged or published, and what landed since v0.6.0 is not a patch: route traces,
Markdown answers in the app, decode traces, the small cache rungs, and a session
teardown fix.

CHANGELOG records what is new since v0.6.0 under [Unreleased]. It is not
converted to a [0.7.0] section because that bucket has accumulated since [0.1.0]
and still holds work shipped in v0.3.0-v0.6.0 — relabelling it wholesale would
credit this release with other releases' changes.
2026-07-15 21:48:50 +02:00
Helldez
5db3399897 docs(android): drop the reference to the retired dense pins
The teardown comment listed "unlock the dense pins" among the ordered steps.
That step belonged to --lock-dense, which is retired: RLIMIT_MEMLOCK is 65536
bytes hard on the target, so the flag could pin 0.003% of a 2 GB cache and its
premise does not survive measurement. See docs/android-memory.md.

The teardown is unchanged — only the comment described a step that never ran.
2026-07-15 20:19:21 +02:00
Helldez
5a51a175a0 fix(android): free the session's RAM before the next one loads
Two defects in the unload path, both on the model/settings-change route.

The delayed force-kill that backs up `close` was posted as an anonymous
Runnable, so nothing could cancel it. Unloading (or hitting the idle
timeout) and then sending a prompt within 1.5s let that stale kill land
on the freshly spawned process: the model loaded and was SIGKILLed
mid-flight, surfacing as "bmoe-cli exited 137". It is a field now, and
startSession cancels it along with the idle unload.

Superseding a session only called destroy(), which signals and returns,
then spawned the replacement immediately. The two processes overlapped,
so the new one sized its expert cache from a MemAvailable still deflated
by the model the old one had not released yet — quietly starving the
cache the settings change was meant to retune, and risking an OOM kill
on a >RAM model. The old process is now handed to the session thread,
which waits for it to exit before spawning (off the main thread).

Both teardown routes ask the session to close instead of killing it
outright, so it unhooks, joins the IO pool, unlocks the dense pins and
frees the expert cache itself; the kill stays as the deadline.
2026-07-15 20:19:21 +02:00
Helldez
019178bff7 feat(android): 500 and 1000 MiB cache rungs, to find where the cache stops paying
The Settings only offered 0 or >= 2000 because the engine rejects anything under its
1500 MiB floor. That floor says a cache smaller than one token's routed working set
can only thrash — sound, but measured on models whose cache pays for itself. On a
>RAM model the question is live: gpt-oss-120b at top-2 routes ~886 MB per token and
returns an 8-13% hit from a 2000-3000 MiB budget, because that budget covers ~5% of a
56.8 GB expert bank. It may already be below the floor's intent while sitting well
above its number.

So the rungs need the floor's own escape hatch: a fixed budget under 1500 now sends
--force-cache, which exists for exactly this ("tests/experiments"). 1000 sits just
above one token's working set, 500 below it — the two sides of the floor's claim,
which is the point of having both.

The help text stops asserting ">= 2000 MiB" (untrue now) and says what the small
rungs are for, since a knob that thrashes by design should say so where it is turned.
2026-07-15 20:10:12 +02:00
Helldez
c5c266a351 fix(android): clear the I/O mode when the session is replaced
The engine honoured --no-odirect and reported o_direct=0, but the app kept
showing "direct (O_DIRECT)": startSession reset the rest of the UI state and
left ioMode alone, so the stderr sniffer hit its first-writer-wins guard and
kept the value the previous session had published. Turning O_DIRECT off in
Settings therefore looked like it did nothing. The mmap baseline had it
worse — it prints no "expert streaming ON" line at all, so it inherited the
previous streaming session's I/O mode outright.

The guard itself is right: it protects the fallback notice ("O_DIRECT
returns wrong data on this storage"), which the engine prints ahead of the
streaming banner, from being overwritten by the plain "buffered" that
follows. It just needs a null to start each session from.
2026-07-15 17:56:21 +02:00
Helldez
3b64747f1d feat(android): render answers as Markdown and stop the scroll from fighting the user (#24)
The demo showed the model's answer as one flat string, so every heading, list and
code fence arrived as literal syntax. A hand-rolled renderer covers the subset a
chat model emits; a library would be a heavier dependency than the feature, and a
parser we own can leave the half-written markup of a streaming answer as plain
text instead of swallowing it until the closing token lands.

The transcript also scrolled to the tail on every token, which made a long answer
unreadable: scrolling back was undone by the next token, milliseconds later. The
follow now detaches when the user drags and re-arms only where a scroll comes to
rest at the bottom, with a Jump to latest button as the way back. Scrolling pins
to the end of the last item rather than its top, so the newest text is what stays
on screen once an answer grows past the viewport.
2026-07-15 11:03:59 +02:00
Helldez
7802b23f32 feat(android): wall-additive telemetry panel with CPU occupancy and faults
The decode meters showed compute vs. flash I/O, but those do not sum to the token
time — under overlap io_ms is a per-lane busy sum that runs parallel to compute —
so tok/s could not be read off the panel. Replace them with the three terms that
sum to the wall time by construction (compute + flash-wait + cache-mgmt), under a
"<ms>/token -> <tok/s>" headline, so the rate is the direct inverse of the total.
Flash-wait is derived as wall - compute - mgmt, which equals the measured stall.

Add a diagnostic line — CPU occupancy (cpu_s / (wall x threads)), major
faults/token and cache hit — so the compute bar is qualified rather than trusted
blindly: low occupancy or high faults means part of "compute" is a throttled core
or a flash fault, not matmul. Parse the new mgmt_ms / majflt / cpu_ms per-token
fields and mgmt_s_tok / majflt_tok / cpu_s_tok run averages from the session
lines. Fix the summary token count to show tokens actually generated (t.step)
instead of the n_predict target.
2026-07-15 07:03:17 +02:00
Helldez
6439a06a07 feat(android): expose the dense warm-up toggle; bump to 0.6.0
Add a "Dense weight warm-up" switch to the streaming settings, wired through
AppSettings.toArgv as --no-warm-dense when disabled (engine default is on). It
joins the session signature so toggling reopens the session, and persists like
the other run options. versionCode 6 -> 7, versionName 0.5.0 -> 0.6.0.
2026-07-14 18:44:15 +02:00
Helldez
6e773f553d feat(android): show CPU temperature instead of battery temperature
The live temperature figure shown during generation now reports the SoC/CPU
temperature read from the kernel thermal zones (/sys/class/thermal), which
track compute load directly, rather than the battery pack temperature, which
lags behind and reflects charging as much as inference heat.

The CPU zone is discovered once by matching its `type` (e.g. cpu-0-0-usr,
cpu_thermal, mtktscpu) and cached; readings are normalised by magnitude
(millidegrees / tenths / degrees). If no CPU zone is readable, it falls back to
the battery temperature so the figure never goes blank.
2026-07-14 16:16:40 +02:00
Helldez
51933985d7 feat(android): richer telemetry — prefill, TTFT, streamed MB, cache, temperature
The live panel showed only decode tok/s, the compute/flash split and cache
hit rate. Surface the rest of what the engine already reports, plus a device
temperature reading:

- prefill rate (tok/s) and time-to-first-token (model load + prompt prefill)
- flash streamed this turn (MB) and expert-cache footprint (resident/budget MiB)
- live battery temperature (BatteryManager, no permission) as a thermal-headroom
  proxy — read on the Android side, it does not travel through the engine

The first four were already computed; only the streamed total needed a new
read_mib field on the BMOE_DONE session line (docs/telemetry.md updated). The
prefill rate and TTFT are also folded into the per-turn transcript line and the
summary. Bumps the app to 0.5.0 (versionCode 6).
2026-07-14 11:01:28 +02:00
Helldez
9eea9743b6 refactor(moe): remove speculative gating to restore the modular seam
Speculative gating was the only feature that broke the ports-and-adapters
seam: it made router_hook reach into architecture-specific router math
(RouterPre), spawned a second thread inside the eval-callback bridge, and
inlined predictor logic into the streaming hot path. It was experimental and
default-off, and never paid its way in steady-state decode on device.

Removing it collapses MoeRecipe back to {arch, expert suffixes} — the header's
stated design intent — and router_hook back to capture -> gather -> load_layer
plus temporal prefetch. The shared speculative-prefetch queue in
expert_stream_source (used by --prefetch) is untouched.

- delete core/src/moe/spec_dot.{h,cpp} and docs/spec-gating.md
- strip RouterPre + router-node fields from recipe.h and every registry row
- drop spec_gate / spec_recall_* config, the run() wiring, and the
  moe_spec_recall_pct / moe_spec_auto_off summary fields (+ CLI flags,
  BMOE_SPEC_GATE env, BMOE_DONE + CSV columns, moe-spec-gate print)
- remove the Android "Speculative gating" toggle and specGate setting
- drop gates G6a-d; G1-G5 and S1-S3 still prove streamed == resident
- clean the reusable bench scripts of --spec-gate; keep docs/bench-data as an
  archive of the historical measurements

Host byte-identity gates pass for qwen3moe and gemma4.
2026-07-14 10:41:27 +02:00
Helldez
e732d22aa3 feat(android): add 3 and 2 to the active-experts (top-k) dropdown
The gpt-oss family routes top-4 by default, so 4/3/2 give a useful
speed/quality sweep on it (and 2/3 on the wider Qwen/Gemma widths).
Adds 3 and 2 to N_EXPERT_CHOICES; the override flows through the existing
--n-expert-used kv_override, valid in both streaming and mmap mode.
2026-07-13 20:13:32 +02:00
Helldez
c29ec79446 build(android): bump app version to 0.4.0
Cuts a release from main since v0.3.0: Auto expert-cache default (ceiling
3000), the hybrid attention/SSM MoE recipe, and the armv8.2 APK baseline that
drops i8mm so a single binary runs across the full device range (fixes the
SIGILL/exit-132 prefill trap on pre-armv8.6 SoCs such as the Snapdragon 865).
2026-07-13 18:37:52 +02:00
Helldez
36d49a0b00 feat(android): default to Auto cache (ceiling 3000), add a 2000 ceiling option, drop the load-all debug toggle 2026-07-13 16:19:57 +02:00
Helldez
a13fc99c9b fix(android): session-reload race + device-agnostic defaults and telemetry
Changing the model or any streaming setting restarts the engine session, but the torn-down
session's thread — unblocked the moment its process is destroyed — ran its finally/waitFor with
shuttingDown=false and reset the UI to IDLE (or ERROR "bmoe-cli exited") and nulled the process
handles, clobbering the fresh session that was already loading. Each session now carries an epoch;
a superseded thread no longer touches the shared process, UI state, or foreground service.

Also, per device-agnostic feedback:
- Default expert cache is a fixed 3000 MiB (was auto-capped); no benchmark- or device-specific
  tuning in the defaults.
- Settings help text is neutral — describes what each knob does, with no measured numbers or
  device/storage claims. Experimental knobs (prefetch, spec-gate) still marked experimental.
- The prefill phase after load is now signalled in the UI (a slow prefill no longer looks stuck).
- At the end of a run the compute and flash-I/O meters show the per-token AVERAGE, not the last
  token, alongside the average tok/s. BMOE_DONE gains prefill_tps, compute_s_tok, io_s_tok.
2026-07-13 13:56:50 +02:00