Cost
Cost is computed from the provider's reported token usage, priced per category: regular input, cache reads, cache writes (5-minute and 1-hour), and output. If a model has no known price, or usage is missing, the cost is unknown (shown as such), never guessed.
Prices ship in src/pricing/prices.json (dated 2026-09-23): current Anthropic models, and OpenAI's GPT-6, GPT-5.x, GPT-4.1, GPT-4o and o4-mini families. Where OpenAI shows no cache-read or cache-write price for a model, a call that uses one has unknown cost. Things to know:
- Short-context prices only. OpenAI also charges higher per-token prices above a context-size threshold; that tier is not modelled, so requests in it are under-priced. Supply an override if you use it.
- Promotions expire.
gpt-5.6-solis priced at its promotional rate through 2026-11-21 and at the standard rate afterwards (validUntilon the entry). If a promotion is extended, BehavTest will over-report cost until you override it. - OpenAI cache writes (
prompt_tokens_details.cache_write_tokens, GPT-5.6+) are priced at the cache-write rate and treated as a subset ofprompt_tokens, per OpenAI's usage format. OpenAI publishes no official cost formula from those fields, so treat OpenAI cost as an estimate.
Add or override prices in the suite:
"pricing": [{ "provider": "openai", "model": "my-model", "inputPerMTok": 2.5, "outputPerMTok": 10, "cachedInputPerMTok": 1.25, "validUntil": "2027-01-31" }]
or with --prices prices.json (an array or {entries: [...]}; validUntil is optional). Check the provider's pricing page: BehavTest's table is a convenience, not a bill.