Commit graph

20 commits

Author SHA1 Message Date
Vincent Koc
cf41c606df
docs: close remaining one-way link findings in cli, plugins, tools, providers (#143855)
* docs: close remaining one-way link findings in cli, plugins, tools, providers

Adds the back-links and anchors that PR #143157 did not cover, and gives two
"see below" tables real headings to link to.

- Related back-links: cli/tui -> resume, cli/doctor -> status,
  configuration-reference -> configure, voice-call -> voicecall CLI,
  onepassword -> secrets CLI, acp-agents-setup -> acpx reference,
  llama-cpp -> llama-cpp reference, cli/policy -> policy reference,
  ollama -> LM Studio and Memory LanceDB, image/video generation -> OpenRouter,
  google-meet -> ElevenLabs, media-understanding -> Mistral,
  tts service links -> Fish Audio, tools/secrets -> ask_user.
- manifest/config-and-secrets: H3 headings for the dangerousFlags and
  secretInputs detail tables; the two "See below" cells now link to them.
- sdk-overview/capabilities: new "Worker providers" heading; the manifest
  worker-provider contract link now lands on it instead of 36 lines above.
- providers/openrouter: the model-list Note pointed at /concepts/model-providers,
  which carries no OpenRouter catalog; it now points at OpenRouter's own catalog
  and keeps a separate pointer to OpenClaw model selection.
- glossary.zh-CN: 7 sources for the new list-item link labels.

* docs(google): link the Gemini CLI runtime tab to the CLI backends page

r3-2107. `google-gemini-cli` is the CLI backend id the bundled Google plugin
registers, so the tab that configures it should point at the page documenting
its argv, JSONL dialect, and session behavior. The Related card alone left the
tab itself unlinked.

---------

Co-authored-by: Vincent Koc <vincent@openclaw.org>
2026-09-10 16:15:08 +08:00
Vincent Koc
f72699f858
docs: STE pass on terminology consistency and run-on sentences (#143770)
* docs: STE pass on terminology consistency and run-on sentences

Closes a batch of ASD-STE100 audit findings grouped by the two most
objectively checkable rules in .audit/ste-policy.md: one word / one
meaning (scan-checklist item 1) and run-on sentences joined by
semicolons (item 5, Rule 8.1).

Terminology:
- Normalize the component name to "Gateway" in prose on
  announcements/bluebubbles-imessage, channels/line, gateway/discovery,
  install/docker, install/digitalocean, install/northflank,
  platforms/android, platforms/ios, platforms/mac/remote.
  Hyphenated attributive compounds (per-gateway, gateway-host,
  gateway-side) stay lowercase, matching the tree and the
  user@gateway-host placeholder mandated by docs/AGENTS.md.
- gateway/discovery: "Node Gateway" -> "Gateway" (used once).
- platforms/mac/remote: "Web Chat" -> "WebChat" (157:7 in the tree).
  The "## Web Chat" heading becomes "## WebChat" with an
  <a id="web-chat" /> stub so the published id survives.
- platforms/android, platforms/ios: one arrow form for menu paths (->
  and > become the dominant ->).
- providers/moonshot: "Kimi Code K3" -> "Kimi Coding K3" in prose;
  vendor product names in link labels are unchanged.
- tools/diffs: one plugin name in the page ("Diff Viewer Language
  Pack plugin", the label in official-external-plugin-catalog.json).

Run-on sentences (Rule 8.1) split into separate sentences on
cli/hooks, cli/webhooks, channels/matrix-push-rules, channels/a2a,
plugins/plugin-permission-requests, plugins/llama-cpp,
platforms/mac/canvas, platforms/mac/health, plugins/teams-meetings,
plugins/zoom-meetings, providers/mistral, tools/diffs,
gateway/secrets-plan-contract, concepts/session-pruning.
The two scrub passes in gateway/secrets-plan-contract and the
changed-files summary card in tools/diffs become lists.

Concrete defects:
- gateway/telemetry: drop the dangling "and applies".
- plugins/zalouser: the opening sentence fragment gets a subject and
  a verb.
- tools/exa-search: the intro named "keyword" and "hybrid" modes that
  the mode table does not list; it now names the table's own ids.
- concepts/session-pruning: state the real ttl default (5 minutes when
  cache-ttl mode is on with no ttl set, per
  resolveCacheTtlPruningSettings) instead of "default 5 minutes when
  set manually".

No fact, hedge, scope qualifier or number changed. No anchor id was
dropped: ids enumerated with parseDocsDocument before and after are
identical across all 28 files except remote.md, which gains "webchat"
and keeps "web-chat".

* docs: turn the diffs baseUrl rules into a list

Three validation rules were joined by two semicolons on one line
(Rule 8.1, and the policy's 'lists for 3 or more conditions'). The
rules and their order are unchanged.

* docs: drop the secrets-plan-contract edit from this PR

.github/CODEOWNERS routes docs/gateway/secrets-plan-contract.md to
@openclaw/openclaw-secops, and this PR has no secops involvement to
point at. ClawSweeper blocked on that and offered omitting the
restricted-path edit as the alternative. The file is back to its
merge-base blob; audit finding r5-0018 is deferred to a secops-routed
PR.

---------

Co-authored-by: Vincent Koc <vincent@openclaw.org>
2026-09-10 15:51:49 +09:00
Ayaan Zaidi
0ed5f76cbe
fix(llama-cpp): report external server model discovery failures (#139846)
## What Problem This Solves

Fixes an issue where refreshing models from a configured external llama-server silently returned an empty or cached model list when the server rejected authentication or discovery failed.

Related: #136257. This extracts the external llama-server producer repair; the source draft stays open and unmerged.

## Why This Change Was Made

The external producer converted structured discovery failures into successful provider results. It now uses the existing live-catalog outcome handler and bypasses the advisory discovery cache on explicit refresh. The shared publication owner can retain compatible inventory, expose the failure, and replace it on recovery.

The provider preserves explicit Authorization precedence and attributes an outcome only to the profile whose actual credential supplied discovery. Synthetic local markers do not turn unconfigured probes into failures. The canonical provider builder already normalizes the endpoint, so redundant reconstruction and the success-shaped failure branch are removed.

## User Impact

- HTTP 401/403 reports authentication rejection; service, transport and malformed-response failures report unavailability.
- Compatible inventory survives temporary failure; replacing the endpoint or credentials cannot retain the prior scope's discovered rows.
- A successful empty catalog removes discovered rows while preserving explicit models.
- Managed llama.cpp, dynamic model preparation, and the published advisory helper defaults stay unchanged.

## Evidence

- Built baseline `ec29145c0e`: Gateway `models.list` returned `models: []` with no outcome during a captured `/models` HTTP 503. A healthy control returned `catalog-one`; the following 401 refresh still had no failure outcome.
- Built candidate `7d76f5eccc52801ebcd61cbc729ab2649517c39a`: the same cold-failure flow returns `providerOutcomes: [{ provider: "llama-cpp", status: "unavailable" }]`. The existing shared static starter remains distinguishable from live discovery by this provider outcome.
- Seven changed regressions fail on the baseline. Candidate producer tests pass 15/15; retained sibling coverage passes 87/87 and doctor contracts pass 5/5. Syntax lint checks both changed TypeScript files with no diagnostics; format and assertion/line ratchets pass.
- Independent code review is clean through P2. Exact builds use isolated Node 24 containers and synthetic HTTP faults through the normal Gateway/CLI; no operator state or live inference account is used.
- Independent runtime acceptance covers immediate refresh, six failure inputs, retained inventory and recovery, endpoint/key/header/profile-secret replacement, authoritative empty/manual rows, actual profile attribution, header precedence, quiet unconfigured reads, and the public advisory helper. Direct curl distinguishes HTTP 503 from an empty-response transport failure; the Gateway reports `unavailable` for both. The public outcome schema intentionally omits the internal rejection scope, and catalog-only rejection leaves explicit execution routes available.
- Stable `@openclaw/llama-cpp-provider@2026.9.2` exposes the plugin entry and embedding policy helper, not this private producer. The published `provider-setup` advisory discovery default and its existing `rawResult` selector are unchanged. The owning manifest already declares refreshable discovery.
- The outcome helper comes from the merged shared catalog prerequisite. Normal release synchronization advances plugin API compatibility and the published host peer range; this PR does not publish a standalone plugin for the old host.
- Scope: production +10 lines, tests/support +37, docs +6; no generated, dependency, configuration or schema changes. The small production increase supplies typed outcomes and credential attribution after removing duplicated provider construction; broader transport or managed-path cleanup would be unrelated.

Known pre-existing follow-up: shared discovery selects synthetic no-auth when only a stored llama.cpp profile is configured and the provider key is unset. This PR leaves shared credential selection unchanged. Positive profile attribution and replacement were exercised through the supported `LLAMA_SERVER_API_KEY` placeholder route with no environment value. Managed/native execution is covered by unchanged owner paths and tests, not live inference.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-09-06 12:08:42 +05:30
Peter Steinberger
408301d36d
fix(llama-cpp): preserve configured model preset settings (#139646)
Keep native preset globals, comments and retained model options while preserving configured inventory pruning and serialized reload ownership. Replace destructive section rendering with one pure preset-format owner.
2026-09-05 22:20:51 -07:00
Peter Steinberger
b3dcb2d52b
fix: preserve local-model capabilities and conversation history (#138850)
* fix: retain local-model tools without changing sibling agents

Infer structured Tool Search from the final resolved provider route instead
of writing global lean mode during model setup. Preserve explicit user
choices and migrate only the lean flag owned by the retired setup marker.

Align CLI and native Ollama context caps and scale the built-in compaction
reserve to small context windows. Preserve ordinary forced message delivery
through the shared tool catalog while keeping private replies restricted.

Full build, focused owner tests, config baseline, and independent review
passed. Fresh CLI/native Ollama evaluation is tracked before landing.

Fixes #138753.

* test: cover private delivery and isolate local-model regressions

* fix: clarify deferred tool calls and trim repeated metadata

Keep full structured call details while presenting only exact tool identity
and the unchanged target result to the model. Explain tools-mode wrapping
and compact automation summaries without changing execution permissions.

Validation: 468 focused tests, failing-before compact-output regression,
and independent Codex review. Fresh real-model setup and recall proof follows.

* test: refresh automation prompt snapshots

Regenerate the canonical prompt fixtures for compact list summaries and
full get details. Snapshot drift check and all16 reconstruction tests pass.
Production remains byte-identical to8fa7d1488c3.

* fix: retain exact timing in compact automation lists

Include existing at/every/cron schedules after the normal visibility filter
so disabled jobs and recurring intervals can be understood in one tool call.
Keep event commands, working directories, payloads, and delivery definitions
behind the existing full-job read. Preserve scheduleKind for API compatibility.

The real local-model trial found the job but invented its schedule because
only the kind was exposed. Owner regressions reproduce that missing state.

* test: dispose assistant panel session capabilities

The panel fixture owns its application-level session capability. Removing
DOM alone leaves subscription retries alive and contaminates later fake
clocks. Register disposal at creation without changing runtime behavior or
weakening the Sessions typing timer assertion.

The original 11-file worker order reproduces the failure before cleanup and
passes all 233 tests afterward. Independent review is clean.

* fix: expose readable run times in compact automation results

Add exact ISO next/last run dates using the canonical timestamp formatter.
Keep the shipped millisecond fields for programmatic clients and preserve
null for absent run times. Models can now report dates without calendar
arithmetic on epoch values.

All 253 Gateway tests pass; five old-producer failures establish timestamp
and absence coverage. Independent review is clean.

* fix: preserve current Ollama context caps during doctor

Do not turn catalog windows or provider output budgets into stronger
num_ctx pins when a current contextTokens cap is already present. Avoid
provider-wide synthesis for mixed current/legacy models; migrate uncapped
native siblings individually and preserve existing explicit overrides.

Retain the shipped native legacy budget precedence and the compatibility
adapter's own API-specific behavior. Simplify the private API classifiers
and budget resolver while fixing the migration owner.

Live reproduction: Doctor tried to pin 262144 over the onboarded 32768 cap.
Eight pre-fix regressions fail; all 98 migration tests pass after the fix.
Independent review is clean. Existing oversized pins remain operator-owned.

* fix(ollama): preserve prepared tool text and continuation hints

Keep caller-budgeted explicit tool text intact when adding structured
fallback content. Bound structured data separately so large file pages
and Tool Search results retain their complete instructions.

Live native wire reproduction showed a 15,976-character tool result
truncated to 8,014 characters, deleting its continuation offset.
Regression coverage preserves long text, Unicode and whitespace while
retaining structured-data bounds and redaction.

* fix(agents): share the command budget with post-turn maintenance

Use the foreground run's remaining allowance for memory flushing and
compaction instead of restarting a full timeout for each phase. Cancel
and settle maintenance before returning the completed reply, while
preserving caller cancellation, restart fencing, and accepted compaction
session/count facts.

A live local-model turn completed in 462 seconds, but its CLI waiter
expired at 630 seconds while a separate memory flush was still running.
The existing model and Gateway deadlines remain unchanged.

Regression coverage includes exhausted and unlimited budgets, shared
phase expiry, committed successors, CLI siblings, and caller abort.

* fix(ollama): preserve native chat history on overflow

Default local native requests to truncate:false and shift:false so context
pressure reaches OpenClaw's compaction/recovery path instead of silently
removing conversation history. Preserve explicit model parameters and keep
hosted routes unchanged. Document the verified runtime and partial-output
behavior.

* fix(ollama): preserve hosted defaults through custom proxies

Exclude explicit Ollama Cloud provider routes from local history-preserving
request defaults even when their model and proxy URL lack cloud markers.
Reuse the existing cloud-origin check and cover the real native request.

* test(ollama): expect history preservation in timeout requests

Keep the complete request expectation aligned with the intentional native
truncate:false and shift:false defaults. Preserve the timeout and every
other existing assertion. No production behavior changes.

* fix(agents): use provider compaction effort and one summary format

Native Ollama summaries default to thinking off through prepared provider
metadata; explicit compaction settings and hosted routes retain precedence.
Use one primary summary format across history, prefix, stage and fallback
requests, preserving focus and existing output budgets.

Three live 4B compactions hit the existing 180-second watchdog with low
effort. A controlled thinking-off request completed in 39 seconds but
exposed conflicting formats. Owner regressions cover both failures.
No new operator configuration, schema, or timeout increases.
2026-09-05 05:32:07 -07:00
Peter Steinberger
4f695ddcef
fix: local model setup leaves CPU chats waiting (#138717)
* fix: keep managed local model setup usable on CPU hosts

* test: complete setup run timing fixture

* test: provide complete managed provider fixtures
2026-09-04 17:59:00 -07:00
Peter Steinberger
9b95e87c4b
feat: choose and verify local models for Gateway hardware (#138464)
* feat: choose and verify local models for Gateway hardware

Recommend pinned Qwen, Gemma, and Muse recipes from available host resources. Install a verified runtime and require a real tool roundtrip before activation while preserving the prior configuration on failure.

Closes #138426

* test: isolate local setup from runner hardware capacity
2026-09-04 11:53:06 -07:00
Ayaan Zaidi
06038f9df8
fix(llama): support embedding-only managed setup (#130883)
Allow low-memory local-memory setups to install only the verified managed server and embedding model while preserving every configured chat route.

Co-authored-by: Ayaan Zaidi <hi@obviy.us>
2026-08-27 18:09:53 +05:30
Peter Steinberger
4118f31d89
fix(llama-cpp): make endpoint auth transitions reproducible (#126498) 2026-08-19 18:37:28 -07:00
Peter Steinberger
0135046830
refactor(llama-cpp): use one provider for managed and existing servers (#126434)
* refactor(llama-cpp): unify server ownership modes

* test(llama-cpp): preserve shared discovery limits

* fix(plugin-sdk): retain provider auth removal export
2026-08-19 13:57:33 -07:00
Onur Solmaz
c2de3206d4
feat(llama-cpp): support external llama-server
* feat(llama-cpp): add external server provider

* feat(llama-cpp): document external server setup

* refactor(llama-cpp): harden external provider boundaries

* fix(llama-cpp): support external structured output

* fix(llama-cpp): isolate replacement endpoint credentials

* test(llama-cpp): register external live shard

* fix(llama-cpp): preserve explicit endpoint authorization

* fix(llama-cpp): clear disabled inline credentials

* fix(llama-cpp): preserve external local service configs

* test(llama-cpp): cover retained external configs

* test(llama-cpp): cover authorization precedence
2026-08-19 17:32:00 +03:00
Peter Steinberger
f65a6f81de
feat(llama-cpp): raise default context size to 64K (#123701)
The managed llama-server default ctx-size was 8192, but the full OpenClaw
agent system prompt alone is ~31K tokens, so the first agent turn overflowed
the context window and forced immediate compaction (observed live on the Mac
app local-model onboarding). Raise the default to 65536 so a fresh local-model
install can run a real agent turn out of the box.

The default-download 16 GiB RAM floor already bounds weaker machines, and
Gemma 4 supports far more than 64K, so this only changes headroom, not the
offer gate. Docs updated to match.
2026-08-14 08:26:41 -07:00
Peter Steinberger
1348387076
refactor(plugins): replace node-llama-cpp with managed llama-server (#123105)
Move llama.cpp chat and local embeddings onto a verified externally managed llama-server runtime. Remove the in-process native runtime, forked embedding workers, and node-llama-cpp dependency while preserving guided setup, local GGUF models, tool-capable agent runs, diagnostics, and operator docs.
2026-08-13 16:58:20 -07:00
Peter Steinberger
914f73ac99
docs: replace retired config keys with canonical schema keys (#121330) 2026-08-09 20:30:43 -07:00
Vincent Koc
e96b9d2cd0
improve(ui): verify llama.cpp model setup 2026-07-30 23:41:58 +08:00
Peter Steinberger
edecdbd05e
refactor(config): config-surface reduction tranche 3 — product consolidations (review request) (#111527)
* refactor(config): consolidate media model lists

* refactor(config): unify memory configuration

* refactor(config): consolidate TTS ownership

* refactor(config): move typing policy to agents

* refactor(config): retire product-level config surfaces

* refactor(config): share scoped tool policy type

* chore(config): refresh generated baselines

* fix(config): honor agent typing overrides

* fix(config): migrate sibling config consumers

* refactor(infra): keep base64url decoder private

* fix(config): strip invalid legacy TTS values

* chore(config): refresh rebased baseline hash

* fix(doctor): route legacy messages.tts.realtime voice to talk during tts move

* refactor(config): polish final layout names

* refactor(config): freeze retired tuning defaults

* feat(config): add fast mode default symmetry

* refactor(config): key agent entries by id

* docs(config): update final layout reference

* test(config): cover final layout migrations

* chore(config): refresh final layout baselines

* fix(config): align final layout runtime readers

* fix(config): align remaining readers

* fix(config): stabilize final layout migrations

* fix(config): finalize config projection proof

* fix(config): address final layout review

* docs(release): preserve historical config names

* fix(config): complete keyed agent migration

* fix(config): close final migration gaps

* fix(config): finish full-branch review

* fix(config): complete runtime secret detection

* fix(config): close final review findings

* fix(config): finish canonical docs and heartbeat migration

* fix(config): integrate latest main after rebase

* refactor(env): isolate test-only controls

* refactor(env): isolate build and development controls

* refactor(env): collapse process identity indirection

* refactor(env): remove duplicate config and temp aliases

* docs(env): define the operator-facing allowlist

* ci(env): ratchet production variable count

* fix(env): remove stale provider helper import

* fix(env): make ratchet sorting explicit

* test(env): keep test seam in dead-code audit

* test(env): cover ratchet growth and boundary; document surface budgets

* docs(config): document tier-eval consolidations

* docs(config): clarify speech preference ownership

* test(memory): align retired tuning fixtures

* refactor(memory): freeze engine heuristics

* refactor(config): apply tier-eval tranche

* refactor(tts): move persona shaping to providers

* refactor(compaction): move prompt policy to providers

* test(config): align hookified prompt fixtures

* chore(deadcode): classify test-only exports

* chore(github): remove unused spawn helper

* chore(deadcode): classify queue diagnostics

* chore(deadcode): remove unused lane snapshot export

* chore(plugin-sdk): ratchet consolidated surface

* fix(config): integrate latest main after rebase
2026-07-21 20:28:43 -07:00
Peter Steinberger
658b601ee5
feat(llama-cpp): in-process local GGUF text inference provider (#109444)
* feat(llama-cpp): add in-process text inference

* test(llama-cpp): narrow setup provider fixture

* fix(llama-cpp): trim public surface and refresh docs map

* fix(llama-cpp): import Context type in inference test
2026-07-16 18:53:55 -07:00
Vincent Koc
3f55aa837d docs(memory): explain llama.cpp runtime diagnostics 2026-07-11 16:40:14 +08:00
Peter Steinberger
f7d7148cf0
docs: rewrite published docs grounded in current source (#100142)
Source-grounded rewrite of 529 published docs pages with per-unit information-loss verification: 1,713 factual corrections cited to src/**, generated surfaces regenerated, frontmatter titles preserved for i18n, release notes pages untouched. All docs gates green.

Closes #100141
2026-07-05 00:32:47 -04:00
Onur Solmaz
3137110167
fix(memory): move local llama.cpp runtime to provider plugin
* fix(memory): move local llama.cpp runtime to provider plugin

* chore: ignore llama cpp dynamic dependency

* test: remove invalid local provider alias fixture

* chore: refresh llama cpp shrinkwrap

* chore: drop stale memory embedding defaults facade
2026-06-09 14:30:35 +08:00