* docs: close remaining one-way link findings in cli, plugins, tools, providers
Adds the back-links and anchors that PR #143157 did not cover, and gives two
"see below" tables real headings to link to.
- Related back-links: cli/tui -> resume, cli/doctor -> status,
configuration-reference -> configure, voice-call -> voicecall CLI,
onepassword -> secrets CLI, acp-agents-setup -> acpx reference,
llama-cpp -> llama-cpp reference, cli/policy -> policy reference,
ollama -> LM Studio and Memory LanceDB, image/video generation -> OpenRouter,
google-meet -> ElevenLabs, media-understanding -> Mistral,
tts service links -> Fish Audio, tools/secrets -> ask_user.
- manifest/config-and-secrets: H3 headings for the dangerousFlags and
secretInputs detail tables; the two "See below" cells now link to them.
- sdk-overview/capabilities: new "Worker providers" heading; the manifest
worker-provider contract link now lands on it instead of 36 lines above.
- providers/openrouter: the model-list Note pointed at /concepts/model-providers,
which carries no OpenRouter catalog; it now points at OpenRouter's own catalog
and keeps a separate pointer to OpenClaw model selection.
- glossary.zh-CN: 7 sources for the new list-item link labels.
* docs(google): link the Gemini CLI runtime tab to the CLI backends page
r3-2107. `google-gemini-cli` is the CLI backend id the bundled Google plugin
registers, so the tab that configures it should point at the page documenting
its argv, JSONL dialect, and session behavior. The Related card alone left the
tab itself unlinked.
---------
Co-authored-by: Vincent Koc <vincent@openclaw.org>
* docs: STE pass on terminology consistency and run-on sentences
Closes a batch of ASD-STE100 audit findings grouped by the two most
objectively checkable rules in .audit/ste-policy.md: one word / one
meaning (scan-checklist item 1) and run-on sentences joined by
semicolons (item 5, Rule 8.1).
Terminology:
- Normalize the component name to "Gateway" in prose on
announcements/bluebubbles-imessage, channels/line, gateway/discovery,
install/docker, install/digitalocean, install/northflank,
platforms/android, platforms/ios, platforms/mac/remote.
Hyphenated attributive compounds (per-gateway, gateway-host,
gateway-side) stay lowercase, matching the tree and the
user@gateway-host placeholder mandated by docs/AGENTS.md.
- gateway/discovery: "Node Gateway" -> "Gateway" (used once).
- platforms/mac/remote: "Web Chat" -> "WebChat" (157:7 in the tree).
The "## Web Chat" heading becomes "## WebChat" with an
<a id="web-chat" /> stub so the published id survives.
- platforms/android, platforms/ios: one arrow form for menu paths (->
and > become the dominant ->).
- providers/moonshot: "Kimi Code K3" -> "Kimi Coding K3" in prose;
vendor product names in link labels are unchanged.
- tools/diffs: one plugin name in the page ("Diff Viewer Language
Pack plugin", the label in official-external-plugin-catalog.json).
Run-on sentences (Rule 8.1) split into separate sentences on
cli/hooks, cli/webhooks, channels/matrix-push-rules, channels/a2a,
plugins/plugin-permission-requests, plugins/llama-cpp,
platforms/mac/canvas, platforms/mac/health, plugins/teams-meetings,
plugins/zoom-meetings, providers/mistral, tools/diffs,
gateway/secrets-plan-contract, concepts/session-pruning.
The two scrub passes in gateway/secrets-plan-contract and the
changed-files summary card in tools/diffs become lists.
Concrete defects:
- gateway/telemetry: drop the dangling "and applies".
- plugins/zalouser: the opening sentence fragment gets a subject and
a verb.
- tools/exa-search: the intro named "keyword" and "hybrid" modes that
the mode table does not list; it now names the table's own ids.
- concepts/session-pruning: state the real ttl default (5 minutes when
cache-ttl mode is on with no ttl set, per
resolveCacheTtlPruningSettings) instead of "default 5 minutes when
set manually".
No fact, hedge, scope qualifier or number changed. No anchor id was
dropped: ids enumerated with parseDocsDocument before and after are
identical across all 28 files except remote.md, which gains "webchat"
and keeps "web-chat".
* docs: turn the diffs baseUrl rules into a list
Three validation rules were joined by two semicolons on one line
(Rule 8.1, and the policy's 'lists for 3 or more conditions'). The
rules and their order are unchanged.
* docs: drop the secrets-plan-contract edit from this PR
.github/CODEOWNERS routes docs/gateway/secrets-plan-contract.md to
@openclaw/openclaw-secops, and this PR has no secops involvement to
point at. ClawSweeper blocked on that and offered omitting the
restricted-path edit as the alternative. The file is back to its
merge-base blob; audit finding r5-0018 is deferred to a secops-routed
PR.
---------
Co-authored-by: Vincent Koc <vincent@openclaw.org>
## What Problem This Solves
Fixes an issue where refreshing models from a configured external llama-server silently returned an empty or cached model list when the server rejected authentication or discovery failed.
Related: #136257. This extracts the external llama-server producer repair; the source draft stays open and unmerged.
## Why This Change Was Made
The external producer converted structured discovery failures into successful provider results. It now uses the existing live-catalog outcome handler and bypasses the advisory discovery cache on explicit refresh. The shared publication owner can retain compatible inventory, expose the failure, and replace it on recovery.
The provider preserves explicit Authorization precedence and attributes an outcome only to the profile whose actual credential supplied discovery. Synthetic local markers do not turn unconfigured probes into failures. The canonical provider builder already normalizes the endpoint, so redundant reconstruction and the success-shaped failure branch are removed.
## User Impact
- HTTP 401/403 reports authentication rejection; service, transport and malformed-response failures report unavailability.
- Compatible inventory survives temporary failure; replacing the endpoint or credentials cannot retain the prior scope's discovered rows.
- A successful empty catalog removes discovered rows while preserving explicit models.
- Managed llama.cpp, dynamic model preparation, and the published advisory helper defaults stay unchanged.
## Evidence
- Built baseline `ec29145c0e`: Gateway `models.list` returned `models: []` with no outcome during a captured `/models` HTTP 503. A healthy control returned `catalog-one`; the following 401 refresh still had no failure outcome.
- Built candidate `7d76f5eccc52801ebcd61cbc729ab2649517c39a`: the same cold-failure flow returns `providerOutcomes: [{ provider: "llama-cpp", status: "unavailable" }]`. The existing shared static starter remains distinguishable from live discovery by this provider outcome.
- Seven changed regressions fail on the baseline. Candidate producer tests pass 15/15; retained sibling coverage passes 87/87 and doctor contracts pass 5/5. Syntax lint checks both changed TypeScript files with no diagnostics; format and assertion/line ratchets pass.
- Independent code review is clean through P2. Exact builds use isolated Node 24 containers and synthetic HTTP faults through the normal Gateway/CLI; no operator state or live inference account is used.
- Independent runtime acceptance covers immediate refresh, six failure inputs, retained inventory and recovery, endpoint/key/header/profile-secret replacement, authoritative empty/manual rows, actual profile attribution, header precedence, quiet unconfigured reads, and the public advisory helper. Direct curl distinguishes HTTP 503 from an empty-response transport failure; the Gateway reports `unavailable` for both. The public outcome schema intentionally omits the internal rejection scope, and catalog-only rejection leaves explicit execution routes available.
- Stable `@openclaw/llama-cpp-provider@2026.9.2` exposes the plugin entry and embedding policy helper, not this private producer. The published `provider-setup` advisory discovery default and its existing `rawResult` selector are unchanged. The owning manifest already declares refreshable discovery.
- The outcome helper comes from the merged shared catalog prerequisite. Normal release synchronization advances plugin API compatibility and the published host peer range; this PR does not publish a standalone plugin for the old host.
- Scope: production +10 lines, tests/support +37, docs +6; no generated, dependency, configuration or schema changes. The small production increase supplies typed outcomes and credential attribution after removing duplicated provider construction; broader transport or managed-path cleanup would be unrelated.
Known pre-existing follow-up: shared discovery selects synthetic no-auth when only a stored llama.cpp profile is configured and the provider key is unset. This PR leaves shared credential selection unchanged. Positive profile attribution and replacement were exercised through the supported `LLAMA_SERVER_API_KEY` placeholder route with no environment value. Managed/native execution is covered by unchanged owner paths and tests, not live inference.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
Keep native preset globals, comments and retained model options while preserving configured inventory pruning and serialized reload ownership. Replace destructive section rendering with one pure preset-format owner.
* fix: retain local-model tools without changing sibling agents
Infer structured Tool Search from the final resolved provider route instead
of writing global lean mode during model setup. Preserve explicit user
choices and migrate only the lean flag owned by the retired setup marker.
Align CLI and native Ollama context caps and scale the built-in compaction
reserve to small context windows. Preserve ordinary forced message delivery
through the shared tool catalog while keeping private replies restricted.
Full build, focused owner tests, config baseline, and independent review
passed. Fresh CLI/native Ollama evaluation is tracked before landing.
Fixes#138753.
* test: cover private delivery and isolate local-model regressions
* fix: clarify deferred tool calls and trim repeated metadata
Keep full structured call details while presenting only exact tool identity
and the unchanged target result to the model. Explain tools-mode wrapping
and compact automation summaries without changing execution permissions.
Validation: 468 focused tests, failing-before compact-output regression,
and independent Codex review. Fresh real-model setup and recall proof follows.
* test: refresh automation prompt snapshots
Regenerate the canonical prompt fixtures for compact list summaries and
full get details. Snapshot drift check and all16 reconstruction tests pass.
Production remains byte-identical to8fa7d1488c3.
* fix: retain exact timing in compact automation lists
Include existing at/every/cron schedules after the normal visibility filter
so disabled jobs and recurring intervals can be understood in one tool call.
Keep event commands, working directories, payloads, and delivery definitions
behind the existing full-job read. Preserve scheduleKind for API compatibility.
The real local-model trial found the job but invented its schedule because
only the kind was exposed. Owner regressions reproduce that missing state.
* test: dispose assistant panel session capabilities
The panel fixture owns its application-level session capability. Removing
DOM alone leaves subscription retries alive and contaminates later fake
clocks. Register disposal at creation without changing runtime behavior or
weakening the Sessions typing timer assertion.
The original 11-file worker order reproduces the failure before cleanup and
passes all 233 tests afterward. Independent review is clean.
* fix: expose readable run times in compact automation results
Add exact ISO next/last run dates using the canonical timestamp formatter.
Keep the shipped millisecond fields for programmatic clients and preserve
null for absent run times. Models can now report dates without calendar
arithmetic on epoch values.
All 253 Gateway tests pass; five old-producer failures establish timestamp
and absence coverage. Independent review is clean.
* fix: preserve current Ollama context caps during doctor
Do not turn catalog windows or provider output budgets into stronger
num_ctx pins when a current contextTokens cap is already present. Avoid
provider-wide synthesis for mixed current/legacy models; migrate uncapped
native siblings individually and preserve existing explicit overrides.
Retain the shipped native legacy budget precedence and the compatibility
adapter's own API-specific behavior. Simplify the private API classifiers
and budget resolver while fixing the migration owner.
Live reproduction: Doctor tried to pin 262144 over the onboarded 32768 cap.
Eight pre-fix regressions fail; all 98 migration tests pass after the fix.
Independent review is clean. Existing oversized pins remain operator-owned.
* fix(ollama): preserve prepared tool text and continuation hints
Keep caller-budgeted explicit tool text intact when adding structured
fallback content. Bound structured data separately so large file pages
and Tool Search results retain their complete instructions.
Live native wire reproduction showed a 15,976-character tool result
truncated to 8,014 characters, deleting its continuation offset.
Regression coverage preserves long text, Unicode and whitespace while
retaining structured-data bounds and redaction.
* fix(agents): share the command budget with post-turn maintenance
Use the foreground run's remaining allowance for memory flushing and
compaction instead of restarting a full timeout for each phase. Cancel
and settle maintenance before returning the completed reply, while
preserving caller cancellation, restart fencing, and accepted compaction
session/count facts.
A live local-model turn completed in 462 seconds, but its CLI waiter
expired at 630 seconds while a separate memory flush was still running.
The existing model and Gateway deadlines remain unchanged.
Regression coverage includes exhausted and unlimited budgets, shared
phase expiry, committed successors, CLI siblings, and caller abort.
* fix(ollama): preserve native chat history on overflow
Default local native requests to truncate:false and shift:false so context
pressure reaches OpenClaw's compaction/recovery path instead of silently
removing conversation history. Preserve explicit model parameters and keep
hosted routes unchanged. Document the verified runtime and partial-output
behavior.
* fix(ollama): preserve hosted defaults through custom proxies
Exclude explicit Ollama Cloud provider routes from local history-preserving
request defaults even when their model and proxy URL lack cloud markers.
Reuse the existing cloud-origin check and cover the real native request.
* test(ollama): expect history preservation in timeout requests
Keep the complete request expectation aligned with the intentional native
truncate:false and shift:false defaults. Preserve the timeout and every
other existing assertion. No production behavior changes.
* fix(agents): use provider compaction effort and one summary format
Native Ollama summaries default to thinking off through prepared provider
metadata; explicit compaction settings and hosted routes retain precedence.
Use one primary summary format across history, prefix, stage and fallback
requests, preserving focus and existing output budgets.
Three live 4B compactions hit the existing 180-second watchdog with low
effort. A controlled thinking-off request completed in 39 seconds but
exposed conflicting formats. Owner regressions cover both failures.
No new operator configuration, schema, or timeout increases.
* feat: choose and verify local models for Gateway hardware
Recommend pinned Qwen, Gemma, and Muse recipes from available host resources. Install a verified runtime and require a real tool roundtrip before activation while preserving the prior configuration on failure.
Closes#138426
* test: isolate local setup from runner hardware capacity
Allow low-memory local-memory setups to install only the verified managed server and embedding model while preserving every configured chat route.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
The managed llama-server default ctx-size was 8192, but the full OpenClaw
agent system prompt alone is ~31K tokens, so the first agent turn overflowed
the context window and forced immediate compaction (observed live on the Mac
app local-model onboarding). Raise the default to 65536 so a fresh local-model
install can run a real agent turn out of the box.
The default-download 16 GiB RAM floor already bounds weaker machines, and
Gemma 4 supports far more than 64K, so this only changes headroom, not the
offer gate. Docs updated to match.
Move llama.cpp chat and local embeddings onto a verified externally managed llama-server runtime. Remove the in-process native runtime, forked embedding workers, and node-llama-cpp dependency while preserving guided setup, local GGUF models, tool-capable agent runs, diagnostics, and operator docs.
* feat(llama-cpp): add in-process text inference
* test(llama-cpp): narrow setup provider fixture
* fix(llama-cpp): trim public surface and refresh docs map
* fix(llama-cpp): import Context type in inference test