* docs: fix concrete gateway/provider setup defects from the ux audit
Closes 20 `ux`-kind audit findings in docs/gateway/ and docs/providers/
that named an objectively checkable defect: a missing prerequisite, a step
with no command, an unexplained placeholder, or a silent flag divergence.
- gateway/1password: the token-file block reads and then unsets
$OP_SERVICE_ACCOUNT_TOKEN, but nothing told the reader to set it, so a
pasted block wrote an empty token file (r3-1654).
- gateway/prometheus: "Restart the Gateway" had no restart command (r3-1618).
- gateway/doctor/config-migrations: "fix those keys by hand against the
current config reference" had no link (r3-1530).
- gateway/bonjour: dropped an undefined acronym ("OCM"), defined nowhere in
docs/ (r3-1641).
- providers/litellm: `litellm --model claude-opus-4-6` needs the upstream
provider key, named only in a collapsed accordion 110 lines below (r3-2201).
- providers/vllm: step 1 said to start a vLLM server but gave no command
(r3-2132).
- providers/vercel-ai-gateway: install step omitted `openclaw gateway
restart`, which the next step needs (r3-2190).
- providers/synthetic: "Verify the default model" had no verify command
(r3-2205).
- providers/gmi: the non-interactive hint named one flag, no command
(r5-0059).
- providers/stepfun: both steps titled "Non-interactive alternative" omitted
`--non-interactive` and `--accept-risk` (r3-2167).
- providers/cerebras, fireworks, meta, tencent: two non-interactive commands
per page differed only by `--mode local`, unexplained; `--mode` defaults to
`local` (docs/cli/onboard.md) (r3-2153, r3-2151, r3-2155, r3-2188).
- providers/xiaomi: TTS example used `apiKey: "xiaomi_api_key"`, drifting from
the canonical docs/tools/tts/configuration.md (r3-2136).
- providers/nvidia: the custom provider id read as a reserved value (r3-2145).
- providers/baseten: the daemon-env note stated the problem without the fix
the other eight copies give (r3-2150).
- providers/groq: a Step code fence was unindented out of its container
(r3-2194).
- providers/fal, zai: neither page said where the API key comes from
(r3-2142, r3-2139).
* docs: keep provider onboarding on the Gateway host (ClawSweeper P2)
Remote-client onboarding writes gateway.remote connection settings only:
runNonInteractiveRemoteSetup returns before local provider setup and never
consumes the provider credential (src/commands/onboard-non-interactive/
remote.ts), and docs/start/wizard-cli-reference.md:434 already says so.
The --mode remote alternative I added to cerebras, fireworks, meta and
tencent would therefore have left the provider unconfigured. Replaced it
with the Gateway-host instruction on all four pages; the --mode local
equivalence statement, which was the point of the finding, stays.
12 KiB
| summary | title | read_when | |||
|---|---|---|---|---|---|
| fal image, video, and music generation setup in OpenClaw | Fal |
|
OpenClaw ships a bundled fal provider for hosted image, video, and music
generation.
| Property | Value |
|---|---|
| Provider | fal |
| Auth | FAL_KEY (canonical; FAL_API_KEY also works as a fallback) |
| API | fal model endpoints (https://fal.run; video jobs use https://queue.fal.run) |
| Base URL | Override with models.providers.fal.baseUrl |
Getting started
Create a key at [fal.ai/dashboard/keys](https://fal.ai/dashboard/keys). ```bash openclaw onboard --auth-choice fal-api-key ```Non-interactive setups can pass `--fal-api-key <key>` or export `FAL_KEY`.
Onboarding also sets `fal/fal-ai/flux/dev` as the default image model when
none is configured.
```json5
{
agents: {
defaults: {
mediaModels: {
image: {
primary: "fal/fal-ai/flux/dev",
},
},
},
},
}
```
Image generation
The bundled fal image-generation provider defaults to
fal/fal-ai/flux/dev.
| Capability | Value |
|---|---|
| Max images | 4 per request; Krea 2: 1 per request |
| Size overrides | 1024x1024, 1024x1536, 1536x1024, 1024x1792, 1792x1024 |
| Aspect ratio | Supported everywhere except Flux image-to-image |
| Resolution | 1K, 2K, 4K (per-model limits below) |
| Output format | png (default) or jpeg; GPT Image 2.5 also supports webp; Krea 2 rejects overrides |
Edit requests (reference images via the shared image / images parameters)
route to a per-model edit endpoint with per-model reference limits:
| Model family | Model ref after fal/ |
Edit endpoint | Max reference images |
|---|---|---|---|
| Flux and other fal models | fal-ai/flux/dev (default) |
/image-to-image |
1 |
| GPT Image 2.5 | openai/gpt-image-2.5/{flare,sunburst}/text-to-image |
sibling /edit |
16 |
| Older GPT Image | openai/gpt-image-* |
/edit |
10 |
| Grok Imagine | xai/grok-imagine-image |
/edit |
3 |
| Nano Banana (legacy) | fal-ai/nano-banana |
/edit |
3 |
| Nano Banana 2 | fal-ai/nano-banana-* |
/edit |
14 |
| Nano Banana 2 Lite | google/nano-banana-2-lite |
/edit |
14 |
| Krea 2 | krea/v2/{medium,large}/text-to-image |
none (style refs) | 10 style references |
GPT Image 2.5
Select either variant:
fal/openai/gpt-image-2.5/flare/text-to-imagefal/openai/gpt-image-2.5/sunburst/text-to-image
References select the sibling /edit endpoint. You can also select
fal/openai/gpt-image-2.5/flare/edit or fal/openai/gpt-image-2.5/sunburst/edit
explicitly.
Both variants support quality: "low", "medium", "high", "xhigh",
"max", or "auto". The fal default is high.
They accept background: "transparent", "opaque", or "auto".
For transparency, use outputFormat: "png" or "webp".
These controls do not change older fal models.
Use size: "auto" or explicit dimensions such as 1536x864.
Dimensions must be divisible by 16, with no edge above 3840 pixels.
Total pixels must be 655,360-8,294,400, with an aspect ratio from 1:3 to 3:1.
OpenClaw converts aspect-ratio hints to valid dimensions.
For example, aspectRatio: "3:2" produces 1536x1024.
Use size to choose exact dimensions. OpenClaw rejects invalid explicit sizes.
These models reject resolution overrides. Edits without geometry hints keep
fal's automatic size selection.
openclaw infer image generate \
--model fal/openai/gpt-image-2.5/flare/text-to-image \
--prompt "A simple red circle sticker" \
--quality low --size 1024x1024 --json
openclaw infer image edit \
--model fal/openai/gpt-image-2.5/sunburst/edit \
--file /path/to/reference.png \
--prompt "Keep the shape and change the color to blue" \
--quality low --size auto --json
Krea 2
Krea 2 models use fal's native Krea payload schema. OpenClaw sends
aspect_ratio, creativity, and image_style_references instead of the
generic image_size / edit-endpoint payload used by Flux. The model refs are:
fal/krea/v2/medium/text-to-imagefal/krea/v2/large/text-to-image
Use Medium for faster expressive illustration, anime, painting, and artistic
styles. Use Large for slower photoreal, raw texture, film grain, and detailed
looks. Krea defaults to fal.creativity: "medium"; supported values are
raw, low, medium, and high.
Krea 2 exposes aspect ratio, not image_size, in fal's request schema. Prefer
aspectRatio; OpenClaw maps size to the closest supported Krea aspect ratio
and rejects resolution for Krea rather than dropping it.
Use outputFormat: "png" when you want PNG output from fal models that expose
output_format. Outside GPT Image 2.5, fal models do not declare a
transparent-background control in OpenClaw. They report background as an
ignored override.
Krea 2 endpoints do not expose an output_format request field through fal, so
OpenClaw rejects outputFormat overrides for Krea requests.
To use Krea 2 Medium:
{
agents: {
defaults: {
mediaModels: {
image: {
primary: "fal/krea/v2/medium/text-to-image",
},
},
},
},
}
Video generation
The bundled fal video-generation provider defaults to
fal/fal-ai/minimax/video-01-live.
| Capability | Value |
|---|---|
| Modes | Text-to-video, single-image reference, Seedance reference-to-video |
| Runtime | Queue-backed submit/status/result flow for long-running jobs |
| Timeout | 20 minutes per job by default; status polled every 5 seconds |
- `fal/fal-ai/minimax/video-01-live`
**HeyGen video-agent:**
- `fal/fal-ai/heygen/v2/video-agent`
**Kling and Wan:**
- `fal/fal-ai/kling-video/v2.1/master/text-to-video`
- `fal/fal-ai/wan/v2.2-a14b/text-to-video`
- `fal/fal-ai/wan/v2.2-a14b/image-to-video`
**Seedance 2.0:**
- `fal/bytedance/seedance-2.0/fast/text-to-video`
- `fal/bytedance/seedance-2.0/fast/image-to-video`
- `fal/bytedance/seedance-2.0/fast/reference-to-video`
- `fal/bytedance/seedance-2.0/text-to-video`
- `fal/bytedance/seedance-2.0/image-to-video`
- `fal/bytedance/seedance-2.0/reference-to-video`
MiniMax Live and HeyGen requests send only the prompt plus an optional
single reference image; other overrides are not forwarded. Seedance models
accept `aspectRatio`, `size`, `resolution`, durations of 4-15 seconds, and
an audio toggle.
```json5
{
agents: {
defaults: {
mediaModels: {
video: {
primary: "fal/bytedance/seedance-2.0/fast/text-to-video",
},
},
},
},
}
```
```json5
{
agents: {
defaults: {
mediaModels: {
video: {
primary: "fal/bytedance/seedance-2.0/fast/reference-to-video",
},
},
},
},
}
```
Reference-to-video accepts up to 9 images, 3 videos, and 3 audio references
through the shared `video_generate` `images`, `videos`, and `audioRefs`
parameters, with at most 12 total reference files. Audio references require
at least one image or video reference in the same request.
```json5
{
agents: {
defaults: {
mediaModels: {
video: {
primary: "fal/fal-ai/heygen/v2/video-agent",
},
},
},
},
}
```
Music generation
The bundled fal plugin also registers a music-generation provider for the
shared music_generate tool.
| Capability | Value |
|---|---|
| Default model | fal/fal-ai/minimax-music/v2.6 |
| Models | fal-ai/minimax-music/v2.6 (mp3), fal-ai/ace-step/prompt-to-audio (wav), fal-ai/stable-audio-25/text-to-audio (wav) |
| Max duration | 240 seconds |
| Runtime | Synchronous request plus generated audio download |
Use fal as the default music provider:
{
agents: {
defaults: {
mediaModels: {
music: {
primary: "fal/fal-ai/minimax-music/v2.6",
},
},
},
},
}
fal-ai/minimax-music/v2.6 supports explicit lyrics and instrumental mode,
but not both in the same request. ACE-Step and Stable Audio are
prompt-to-audio endpoints; choose them with the model override when you want
those model families. ACE-Step rejects explicit lyrics; Stable Audio rejects
both lyrics and instrumental mode.