Tokens Boutique

Notes and updates from Tokens Boutique.

fair • transparent • no bs

Qwen3.8 Flash pricing, context window, and API details

Steven · 2026-08-26

Qwen has released Qwen3.8-Flash-Next, an open-weight multimodal Mixture-of-Experts model and an early look at the architecture being developed for Qwen4. Its production API counterpart, qwen3.8-flash, is now tracked on Tokens Boutique.

According to Qwen's official launch post, the model has a 125B-parameter main network with 6B parameters activated per token, plus 51B N-gram embedding parameters. Qwen says it was trained at roughly one-ninth the cost of Qwen3.7 Plus while improving coding and office-task performance.

FEATAdded qwen3.8-flash to the calculator, model table, and pricing catalog at $0.16 per million input tokens and $0.47 per million output tokens.

FEATThe hosted QwenCloud version records a 1 million-token context window and 64K maximum output. The released Qwen3.8-Flash-Next weights natively support 262,144 tokens and can be extended to 1 million with YaRN.

FEATThe listing clearly marks the QwenCloud API as coming soon. Qwen announced the API model ID and pricing at launch, but we do not describe it as generally available until the vendor does.

FEATNo cached-input price is shown. Qwen has not published one for this model yet.

The addition brings Tokens Boutique to 128 models across 26 vendors.

Compare Qwen3.8 Flash with other Alibaba models.

Qwen3.8 Flash questions and answers

What is Qwen3.8-Flash-Next?

Qwen3.8-Flash-Next is Qwen's newly released open-weight multimodal MoE model. It previews architectural work intended for Qwen4, including Qwen Sparse Attention, Gated Residual, N-gram embeddings, and an updated optimization recipe.

What is the difference between Qwen3.8-Flash-Next and Qwen3.8-Flash?

Qwen3.8-Flash-Next is the name of the released open-weight architecture preview. qwen3.8-flash is the production QwenCloud API version, with a 1M default context window and official built-in tools. Tokens Boutique uses the API model ID because the site compares hosted API pricing.

How much does the Qwen3.8 Flash API cost?

Qwen announced pricing of $0.16 per million input tokens and $0.47 per million output tokens. No separate cached-input rate has been published.

What is the Qwen3.8 Flash context window?

The production qwen3.8-flash API version has a 1M token context window and a 64K max output. The open weights have a native 262,144-token context that is extensible to 1M with YaRN.

Is the Qwen3.8 Flash API available now?

Not yet according to the launch announcement. Qwen lists the API as coming soon. The model is included on Tokens Boutique now because the official API ID, limits, and prices have already been announced; its listing includes the availability warning.

Does Qwen3.8 Flash support images?

Yes. Qwen describes Qwen3.8-Flash-Next as multimodal, and its published QwenCloud configuration accepts text and image input.

Where can I compare Qwen3.8 Flash pricing?

Use the Tokens Boutique Alibaba comparison to compare its input price, output price, context window, and maximum output with the other tracked Qwen API models.