Claude 5.1, Gemini 3.8 Flash, and a Qwen pricing overhaul
Steven · 2026-09-03
Five new models join the table today, alongside a full refresh of Alibaba's Qwen pricing, cache rates, output limits, and lifecycle status.
FEATAdded Claude Fable 5.1 (claude-fable-5-1) and Claude Mythos 5.1 (claude-mythos-5-1). Both have a 1M context window, 128K max output, and cost $10 input / $50 output per million tokens. Their published cache rates are $12.50 for a 5-minute write, $20 for a 1-hour write, and $0.25 for a cache read. Mythos remains limited to vetted partners.
FEATAdded Gemini 3.8 Flash (gemini-3.8-flash), released September 2 with a 1M context window and 64K max output. Introductory pricing is $0.75 input, $0.075 cached input, and $3.75 output per million tokens through 2026-12-31. Standard cache storage is $0.50 per million tokens per hour, with Batch, Flex, and Priority processing options also documented.
FEATAdded two hosted open-weight Qwen3.8 variants: qwen3.8-27b at $0.50 input / $3 output, and the frontier qwen3.8-2.4t-a95b at $2 input / $6 output. Both have 1M context windows, 128K max output, and published cache-read and cache-write pricing.
FIXqwen3.8-flash now reflects its live Model Studio listing: input drops from $0.16 to $0.15 per million, max output rises from 64K to 128K, and cache pricing is now included.
FIXqwen3.7-max no longer shows the expired 50% promotion. Its current Singapore rate is $2.50 input / $7.50 output, its max output is now 128K, cache pricing has been added, and the legacy alias is marked deprecated.
FIXqwen3.7-plus and qwen3.7-flash now show 128K max output and their complete cache pricing, including the higher rates applied at long-context tiers. Qwen3.7 Flash also moves from the former Beijing-only listing to published Singapore/International pricing.
FIXqwen3.6-flash is now marked deprecated.
The catalog now tracks 133 models across 26 vendors.