mirror of
https://github.com/openclaw/openclaw.git
synced 2026-10-03 17:53:39 +00:00
docs: document image replay cost for vision models (#159205)
* docs: document image replay cost for vision models Native-vision sessions skip image understanding by default, so raw images re-attach on later turns inside the replay window at full image-token cost. Document that default, its cost consequence, the explicit tools.media.models[] entry that opts into description-based replay suppression, and the preferredModel/enabled boundaries that do not trigger it. * docs: qualify image description and replay boundaries --------- Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
This commit is contained in:
parent
94e17d521e
commit
32cc3f063d
1 changed files with 25 additions and 1 deletions
|
|
@ -7,7 +7,7 @@ title: "Media understanding"
|
|||
sidebarTitle: "Media understanding"
|
||||
---
|
||||
|
||||
OpenClaw can summarize inbound media (image/audio/video) before the reply pipeline runs, so command parsing and routing work off short text instead of raw bytes. Understanding auto-detects local tools or provider keys, or you can configure explicit models. Original media is always delivered to the model as usual; when understanding fails or is disabled, the reply flow continues unchanged.
|
||||
OpenClaw can summarize inbound media (image/audio/video) before the reply pipeline runs, so command parsing and routing work off short text instead of raw bytes. Understanding auto-detects local tools or provider keys, or you can configure explicit models. In the OpenClaw-managed reply path, successfully described images are represented by text instead of being attached again as native images. Failed, disabled, or unselected images remain eligible for normal native image delivery; understanding errors do not block the reply.
|
||||
|
||||
Vendor plugins register capability metadata (which provider supports which media type, default model, priority). OpenClaw core owns the shared `tools.media` config, fallback order, and reply-pipeline integration.
|
||||
|
||||
|
|
@ -278,6 +278,30 @@ Local attachments stay within the session's allowed media roots. Directory alias
|
|||
- Image, audio, and video decisions record one closed disposition for every attachment candidate: handled, handed to native vision, not selected after the attachment limit, disabled, missing a model, denied by chat scope, or failed.
|
||||
- Unhandled media gets a bounded model-visible marker. Images handed to native vision do not add markers. When a native harness owns the turn and OpenClaw runs only audio preprocessing, failed or skipped audio still gets a marker; image, video, and document inputs remain owned by the harness. Too-small audio keeps its placeholder transcript without a duplicate marker.
|
||||
|
||||
## Vision models and image replay cost
|
||||
|
||||
In the OpenClaw-managed reply path, a vision-capable reply model normally receives native images without a separate description pass. Images retained in later requests can contribute to input cost; the amount depends on the provider and the active runtime's context handling. See [Session pruning](/concepts/session-pruning#legacy-image-cleanup) for the applicable cleanup behavior.
|
||||
|
||||
For identification or extraction tasks that do not need continued pixel access, an image-capable entry in `tools.media.models[]` opts into image understanding even when the reply model supports vision. For each selected image that is successfully described, the reply receives description text in place of that native image. This is not a guarantee that every attached image is converted or that later requests contain no images.
|
||||
|
||||
```json5
|
||||
{
|
||||
tools: {
|
||||
media: {
|
||||
models: [{ provider: "google", model: "gemini-3-flash-preview", capabilities: ["image"] }],
|
||||
},
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
Tradeoffs and boundaries:
|
||||
|
||||
- Selection defaults to the first matching image only. Use the existing `tools.media.image.attachments` policy to select more, with an explicit `maxAttachments` cap. Failed or unselected images can still be delivered natively.
|
||||
- A successful description replaces that image's pixels on the current reply turn. Avoid this when the reply model needs to inspect pixels directly, such as screenshots, UI review, or visual debugging.
|
||||
- An explicit image-capable `tools.media.models[]` entry overrides the native-vision skip, subject to the existing enablement, scope, size, and attachment limits. `tools.media.image.preferredModel` alone does not override that skip.
|
||||
- `tools.media.image.enabled: false` disables image understanding; it does not disable native image delivery.
|
||||
- This describes OpenClaw-managed preprocessing. A native harness that owns image handling has its own delivery and context rules; do not assume this setting controls its image replay or billing.
|
||||
|
||||
## Config examples
|
||||
|
||||
<Tabs>
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue