mirror of
https://github.com/openclaw/openclaw.git
synced 2026-10-03 01:29:56 +00:00
Related: #156507
## What Problem This Solves
Fixes wrong cost estimates when OpenRouter runs a promotion: OpenRouter's price list was treated as every vendor's price list, so OpenAI direct and the Vercel, Kilo and Cloudflare gateways all showed OpenRouter's discounted rate. It also moves clients to the smaller v2 hosted catalog.
## User Impact
Cost estimates now use the price of whoever bills the request. During OpenRouter's current 50% promotion on `openai/gpt-5.6-sol`, OpenAI direct, Vercel, Kilo and Cloudflare show $4/$20 per million tokens instead of $2/$10, while OpenRouter keeps $2/$10.
Clients download `/models/v2/catalog.json` by default: 1.6 MB instead of v1's 3.0 MB. Prices for models outside the catalog, such as gateway routes and older model IDs, still resolve. Released clients keep reading v1, which the same publisher run keeps in step.
Existing installs keep their downloaded catalog across the upgrade, including offline default installs whose stored catalog came from the v1 default URL. Subscription routes such as Kimi Code keep showing pay-as-you-go estimates.
After merge, the catalog repository's publish run regenerates both feeds from `main`, so released clients also receive the corrected prices through v1. 762 keys stop resolving; the route audit below shows none is a supported route with a vendor or gateway price.
Rollout (maintainer decision: land as one PR): the hosted v2 file gains `upstreamPricing`/`providerPricing` on the first publish after merge (every 4 hours). Released clients read v1 and are unaffected. Only `main`/nightly builds that download v2 before that publish briefly lack estimates for routes outside catalog rows; the next publish restores them.
## Why This Change Was Made
- **v2 wire format:** each model carries its price. Two optional sections cover models without a catalog row:
- `upstreamPricing` holds each vendor's rate once, with its source, instead of v1's per-gateway copies.
- `providerPricing` holds a provider's own rate.
- **Price sources follow who bills the request:**
- A new `modelsDev` source prices each provider from its own models.dev entry, reusing the models.dev download the publisher already makes. The entry is chosen by the plugin's existing `modelCatalog.modelsDev` mapping or its `modelsDev.provider` policy.
- OpenRouter's feed prices only `openrouter/*` keys.
- Gateways with their own list (Kilo, Vercel) use it first. Gateways that bill vendor rates (Cloudflare, per its Unified Billing docs) use the vendor's own catalog row.
- Plugin manifest policies are updated to match.
- **Storage:** v2 is stored in its own slot so older clients never read it. Until the new client writes that slot, it reads the older client's catalog. With the default config, a catalog downloaded from the retired v1 default URL stays active until the v2 download lands; configured mirrors must match exactly. A `304` adopts the older client's row only when that row, read inside the same write transaction, matches the revalidated one; otherwise the v2 slot stays empty. The older slot is never written.
## Evidence
- **Pricing, through `resolveModelCostConfig`:**
- For `openai/gpt-5.6-sol`, OpenAI, Vercel, Kilo and Cloudflare now resolve $4/$20, and OpenRouter keeps $2/$10. The references are OpenRouter's endpoint list (`discount: 0.5`), models.dev, Vercel's and Kilo's live model lists, and Cloudflare's Unified Billing docs.
- Across 13,198 keys, compared with the head before the billing change: 11,842 unchanged, 269 changed (for example Kilo `google/gemini-3.5-flash` now matches Kilo's live $0.75/$4.5, and Qwen routes use Alibaba's pay-as-you-go rate), 120 newly priced, 762 no longer priced.
- **Route audit of the 762 (maintainer decision: accepted):**
- 294 are OpenRouter batch-mode variants (`:batch`), which no OpenClaw route sends.
- 423 are Cloudflare/Vercel keys under OpenRouter-only names: `~vendor/…-latest` aliases, dotted Anthropic IDs (`claude-opus-4.5`; Anthropic's API uses dashes), and OpenRouter variants such as `openai/gpt-5.6-sol-pro` that neither the vendor's own list nor the catalog contains (OpenAI direct already had no price for them). None is a vendor catalog row. One (`inclusionai/ling-3.0-flash-fin`) is in Vercel's live model list and is priced by Vercel's runtime discovery.
- 45 are direct-provider keys: retired models the vendor no longer lists (`kimi-k2*`, `claude-opus-4-1`, `claude-sonnet-4`, Gemma, image variants), OpenRouter-style IDs, `zai/glm-4.7-flash` (free on Z.ai; the old number was OpenRouter's), and `qwen/qwen3-coder-next` (only Coding Plan at $0 and resellers list it).
- The audit found and fixed two real gaps: Qwen catalog rows (and gateways passing them through) now use Alibaba's pay-as-you-go list, like Kimi Code's subscription routes; vendors whose models.dev slug differs from their provider ID (`moonshotai`) are also keyed by that slug, so gateway pass-through finds them.
- Before the billing change, v2 with the standalone sections matched v1 on all 13,198 keys, with no lost or changed prices.
- **Real clients against a local mirror of the new feeds:**
- This PR's CLI downloaded v1 and v2 (44 providers, 1,071 models) and revalidated with `304`, and its built Gateway started with v2 stored.
- Released 2026.9.4 and 2026.8.1 accepted the new v1 and rejected v2 cleanly (`schemaVersion: expected 1`, exit 1). A rejected v2 fetch left their stored v1 intact.
- 2026.7.1 predates hosted catalog refresh.
- **Upgrade:**
- A real 2026.9.4 state directory (schema 17) upgraded by this PR's CLI revalidated its existing v1 mirror catalog with `304`, with no full download, and copied it into the v2 slot. The older client's slot was unchanged.
- Default install: 2026.9.4 downloaded the live `catalog.openclaw.ai/models/v1` catalog into fresh state. This PR's code then served that row under the default config, refused it for a configured mirror, and on a mismatched check left the v2 slot empty and the v1 slot byte-identical.
- Published driver × candidate: npm-global 2026.9.4 holding that default v1 catalog ran its own `openclaw update --tag <candidate tarball> --yes --no-restart` onto the package built from `2582b9dd`; the final head adds only a price-projection fix in `remote-bundle.ts` and its test, with no update, Doctor or storage changes. All 11 steps exited 0: global update, candidate migration rehearsal, Doctor lint, config validation, plugin resolution, migration continuation, Gateway canary, install swap, `openclaw doctor`, service reconciliation, and post-plugin Doctor lint. The candidate Gateway then stored the live v2 feed (schema 2, same generation) in its own slot. The v1 slot's bytes were identical (same SHA3) before the update, after it, and after the Gateway ran.
- **Tests:** regressions cover:
- OpenRouter's promotion
- gateway own-list and vendor-row precedence
- pass-through source policy
- standalone parity
- a mirror's standalone price colliding with an unknown catalog row (the row wins)
- the upgrade fallback, including offline default installs and a legacy row changed by an older client before a `304` (both fail on the previous head)
Publisher, catalog, pricing, session-cost, usage-format, Gateway startup and state-database suites pass.
Cost, `pnpm test <file> --maxWorkers=1` (new or changed files; mostly transform time on a cold cache):
| File | Tests | Duration |
|---|---|---|
| `packages/model-catalog-core/src/remote-catalog-bundle.test.ts` | 16 | 0.6 s |
| `test/scripts/publish-model-catalog-v2.test.ts` | 4 | 1.0 s |
| `src/model-catalog/pricing.routing.test.ts` | 4 | 12.2 s |
| `src/model-catalog/pricing.v2.test.ts` | 9 | 13.1 s |
| `src/model-catalog/remote-overlay.test.ts` | 21 | 13.7 s |
| `src/model-catalog/pricing.v2-standalone.test.ts` | 12 | 13.5 s |
| `src/model-catalog/remote-store.test.ts` | 3 | 14.8 s |
| `src/model-catalog/remote-refresh.test.ts` | 10 | 17.6 s |
| `src/infra/session-cost-usage.test.ts` | 62 | 22.1 s |
| `test/scripts/publish-model-catalog.test.ts` | 73 | 23.3 s |
| `src/model-catalog/pricing.test.ts` | 59 | 28.4 s |
| `src/gateway/server.catalog-startup.test.ts` | 3 | 32.5 s |
`server.catalog-startup.test.ts` is over 30 s because it boots a real Gateway, which it already did before this PR; this PR only updates its fixture URL and expected prices.
- **Checks:**
- Hosted CI is running on this head.
- Changed-file typecheck on every lane, core and scripts lint, and dead-code scans pass.
- The extension lint lane runs in hosted CI, because the local worktree is nested inside another install. The extension changes are manifest JSON only.
- **Review:** an independent pricing audit and code reviews. Their findings were fixed, except the follow-ups below.
Follow-ups, not in this PR:
- LiteLLM long-prompt tiers (`_above_Nk_tokens`) are dropped.
- A missing cache-read rate publishes as $0.
- OpenAI and Copilot built-in seeds for gpt-5.6 are stale; the hosted catalog corrects them after the first refresh.
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
250 lines
5.7 KiB
JSON
250 lines
5.7 KiB
JSON
{
|
|
"id": "moonshot",
|
|
"categories": ["models"],
|
|
"activation": {
|
|
"onStartup": false
|
|
},
|
|
"enabledByDefault": true,
|
|
"providerCatalogEntry": "./provider-discovery.ts",
|
|
"providers": [
|
|
"moonshot"
|
|
],
|
|
"providerAuthAliases": {
|
|
"moonshotai": "moonshot",
|
|
"moonshot-ai": "moonshot"
|
|
},
|
|
"providerEndpoints": [
|
|
{
|
|
"endpointClass": "moonshot-native",
|
|
"baseUrls": [
|
|
"https://api.moonshot.ai/v1",
|
|
"https://api.moonshot.cn/v1"
|
|
]
|
|
}
|
|
],
|
|
"providerRequest": {
|
|
"providers": {
|
|
"moonshot": {
|
|
"family": "moonshot",
|
|
"compatibilityFamily": "moonshot"
|
|
}
|
|
}
|
|
},
|
|
"modelPricing": {
|
|
"providers": {
|
|
"moonshot": {
|
|
"modelsDev": {
|
|
"provider": "moonshotai"
|
|
},
|
|
"liteLLM": {
|
|
"provider": "moonshot"
|
|
}
|
|
}
|
|
}
|
|
},
|
|
"modelCatalog": {
|
|
"modelsDev": {
|
|
"moonshot": "moonshotai"
|
|
},
|
|
"aliases": {
|
|
"moonshotai": {
|
|
"provider": "moonshot"
|
|
},
|
|
"moonshot-ai": {
|
|
"provider": "moonshot"
|
|
}
|
|
},
|
|
"providers": {
|
|
"moonshot": {
|
|
"baseUrl": "https://api.moonshot.ai/v1",
|
|
"api": "openai-completions",
|
|
"defaultModel": "kimi-k3",
|
|
"models": [
|
|
{
|
|
"id": "kimi-k3",
|
|
"name": "Kimi K3",
|
|
"reasoning": true,
|
|
"thinkingLevelMap": {
|
|
"off": null,
|
|
"minimal": null,
|
|
"low": "low",
|
|
"medium": null,
|
|
"high": "high",
|
|
"xhigh": "max",
|
|
"max": "max"
|
|
},
|
|
"input": [
|
|
"text",
|
|
"image"
|
|
],
|
|
"contextWindow": 1048576,
|
|
"maxTokens": 1048576,
|
|
"cost": {
|
|
"input": 3,
|
|
"output": 15,
|
|
"cacheRead": 0.3,
|
|
"cacheWrite": 0
|
|
},
|
|
"compat": {
|
|
"supportsReasoningEffort": true,
|
|
"supportedReasoningEfforts": [
|
|
"low",
|
|
"high",
|
|
"max"
|
|
],
|
|
"codeMode": "preferred"
|
|
}
|
|
},
|
|
{
|
|
"id": "kimi-k2.7-code",
|
|
"name": "Kimi K2.7 Code",
|
|
"reasoning": true,
|
|
"input": [
|
|
"text",
|
|
"image"
|
|
],
|
|
"contextWindow": 262144,
|
|
"maxTokens": 262144,
|
|
"cost": {
|
|
"input": 0.95,
|
|
"output": 4,
|
|
"cacheRead": 0.19,
|
|
"cacheWrite": 0
|
|
}
|
|
},
|
|
{
|
|
"id": "kimi-k2.7-code-highspeed",
|
|
"name": "Kimi K2.7 Code HighSpeed",
|
|
"reasoning": true,
|
|
"input": [
|
|
"text",
|
|
"image"
|
|
],
|
|
"contextWindow": 262144,
|
|
"maxTokens": 262144,
|
|
"cost": {
|
|
"input": 1.9,
|
|
"output": 8,
|
|
"cacheRead": 0.38,
|
|
"cacheWrite": 0
|
|
}
|
|
}
|
|
]
|
|
}
|
|
},
|
|
"discovery": {
|
|
"moonshot": "refreshable"
|
|
}
|
|
},
|
|
"setup": {
|
|
"providers": [
|
|
{
|
|
"id": "moonshot",
|
|
"envVars": [
|
|
"MOONSHOT_API_KEY",
|
|
"KIMI_API_KEY"
|
|
]
|
|
}
|
|
]
|
|
},
|
|
"providerAuthChoices": [
|
|
{
|
|
"provider": "moonshot",
|
|
"method": "api-key",
|
|
"choiceId": "moonshot-api-key",
|
|
"appGuidedSecret": true,
|
|
"choiceLabel": "Moonshot API key (.ai)",
|
|
"groupId": "moonshot",
|
|
"groupLabel": "Moonshot AI (Kimi)",
|
|
"groupHint": "Kimi API models \u00b7 https://platform.kimi.ai/docs/pricing/chat",
|
|
"optionKey": "moonshotApiKey",
|
|
"cliFlag": "--moonshot-api-key",
|
|
"cliOption": "--moonshot-api-key <key>",
|
|
"cliDescription": "Moonshot API key"
|
|
},
|
|
{
|
|
"provider": "moonshot",
|
|
"method": "api-key-cn",
|
|
"choiceId": "moonshot-api-key-cn",
|
|
"appGuidedSecret": true,
|
|
"choiceLabel": "Moonshot API key (.cn)",
|
|
"groupId": "moonshot",
|
|
"groupLabel": "Moonshot AI (Kimi)",
|
|
"groupHint": "Kimi API models \u00b7 https://platform.kimi.ai/docs/pricing/chat",
|
|
"optionKey": "moonshotApiKey",
|
|
"cliFlag": "--moonshot-api-key",
|
|
"cliOption": "--moonshot-api-key <key>",
|
|
"cliDescription": "Moonshot API key"
|
|
}
|
|
],
|
|
"uiHints": {
|
|
"webSearch.apiKey": {
|
|
"label": "Kimi Search API Key",
|
|
"help": "Moonshot/Kimi API key (fallback: KIMI_API_KEY or MOONSHOT_API_KEY env var).",
|
|
"sensitive": true
|
|
},
|
|
"webSearch.baseUrl": {
|
|
"label": "Kimi Search Base URL",
|
|
"help": "Kimi base URL override."
|
|
},
|
|
"webSearch.model": {
|
|
"label": "Kimi Search Model",
|
|
"help": "Kimi model override."
|
|
}
|
|
},
|
|
"contracts": {
|
|
"mediaUnderstandingProviders": [
|
|
"moonshot"
|
|
],
|
|
"webSearchProviders": [
|
|
"kimi"
|
|
]
|
|
},
|
|
"mediaUnderstandingProviderMetadata": {
|
|
"moonshot": {
|
|
"capabilities": [
|
|
"image",
|
|
"video"
|
|
],
|
|
"defaultModels": {
|
|
"image": "kimi-k2.6",
|
|
"video": "kimi-k2.6"
|
|
},
|
|
"autoPriority": {
|
|
"video": 20
|
|
}
|
|
}
|
|
},
|
|
"configGroups": [
|
|
{
|
|
"id": "web-search",
|
|
"title": "Web search",
|
|
"order": 10,
|
|
"properties": ["webSearch"]
|
|
}
|
|
],
|
|
"configSchema": {
|
|
"type": "object",
|
|
"additionalProperties": false,
|
|
"properties": {
|
|
"webSearch": {
|
|
"type": "object",
|
|
"additionalProperties": false,
|
|
"properties": {
|
|
"apiKey": {
|
|
"type": [
|
|
"string",
|
|
"object"
|
|
]
|
|
},
|
|
"baseUrl": {
|
|
"type": "string"
|
|
},
|
|
"model": {
|
|
"type": "string"
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|