Tokens Boutique

Notes and updates from Tokens Boutique.

fair • transparent • no bs

11 new API models and 3 new vendors

Steven · 2026-08-25

Another vendor sweep is complete. The site now tracks 127 models across 26 vendors, with three vendors making their first appearance and several existing listings brought up to date.

FEATAdded Upstage with solar-pro4, a frontier model with 512K context and 128K max output. Its 90% launch promotion puts it at $0.03 input, $0.006 cached input, and $0.12 output per million tokens through September 10; the standard $0.30/$0.06/$1.20 rates are recorded too.

FEATAdded Sakana AI with sakana-namazu. It has 256K context and costs $0.95 input, $0.15 cached input, and $4 output. Availability currently excludes the EU/EEA, UK, and Switzerland, and its built-in web search and code execution are billed separately.

FEATAdded Thinking Machines Lab with Inkling, their 256K-context Tinker serverless route. It costs $1 input, $0.17 cached input, and $4.05 output, and remains a beta endpoint that the vendor does not yet recommend for intensive production use.

FEATAlibaba added qwen3.7-flash, with 1M context and three pricing tiers, plus qwen3-coder-next, a 256K coding model. Qwen3.7 Flash is currently a China (Beijing) deployment; Alibaba does not publish a Singapore or international rate for it yet.

FEATAdded the experimental multimodal deepseek-v4-flash-vision-exp and Z.ai's glm-5v-turbo. DeepSeek bills images as input tokens and keeps the same peak/off-peak schedule as V4 Flash.

FEATAdded MiniMax-M2.7-highspeed and Xiaomi's application-gated mimo-v2.5-pro-ultraspeed routes, plus the new perplexity/sonar Agent API model. The namespaced Perplexity model is separate from the legacy Sonar API route scheduled to retire September 27.

FEATAdded NVIDIA's nemotron-3.5-lightning-30b-a3b with its 1M context window. NVIDIA Build offers a free trial endpoint, but NVIDIA does not publish a durable production per-token rate, so the table keeps that caveat attached.

FIXOpenAI's gpt-5.6-sol now reflects its promotional $4 input, $0.40 cached input, and $20 output rates. Requests above 272K input tokens use the documented $8/$30 tier.

FIXminimax-m2.7 now shows its 200K context and output limits, along with explicit cache-read and cache-write pricing.

FIXdoubao-seed-2.1-pro now has the correct June 23 release date and ByteDance's published cache-read and storage rates.

FIXhunyuan-hy3 now includes Tencent's $0.035 per million cached-input rate.

Compare all 127 models.