mirror of
https://github.com/AgentSeal/codeburn.git
synced 2026-08-24 07:54:18 +00:00
Reasoning tokens are a subset of output_tokens for OpenAI models, not an extra bucket: on a 1,396-rollout corpus all 134,316 token_count events carrying a total satisfy input + output == total. CodeBurn added reasoning_output_tokens on top when pricing a codex call, in the cache-rehydration re-price, and in the models/audit display sums. That overstated codex cost by $166.03 (3.5%) and displayed output tokens by 34.6% on that corpus. Both cost sites and the display sums now go through one shared billableOutputTokens() so a cold parse and a warm read cannot drift apart. cache_write_input_tokens (codex PR #33454) was never read and cacheCreationInputTokens was hardcoded to 0. It is now carved out of the uncached-input bucket and clamped to it, but routed to the cache-write bucket ONLY when the pricing source publishes an explicit cache-write rate. buildCosts fabricates 1.25x input when a source omits one, which is correct for Anthropic and would have invented a surcharge OpenAI never charged on gpt-5.5 / 5.4 / 5.3-codex / gpt-5. ModelCosts now carries cacheWriteCostIsExplicit so that distinction survives getModelCosts. A cost change invalidates persisted output: codex-results.json v10 -> v11 (stores costUSD verbatim), the codex parse version moves (the token-bucket change does not self-heal on read), and the daily cache goes 20 -> 23 (21 is claimed by the #946 landing branch and 22 by PR #1056). The upgrade-path corpus asserts codex tokens and calls exactly and reports the repricing. Closes #1075 |
||
|---|---|---|
| .. | ||
| upgrade-path | ||
| bundle-litellm.mjs | ||