Alibaba's flash-tier multimodal MoE model and an early preview of the Qwen4 architecture. Use qwen3.8-flash in AdaL for fast, low-cost coding backed by a 262K native context window (extensible to 1M with YaRN).
Get StartedRun /model and pick Qwen3.8 Flash · $0.16/M input · $0.47/M output
Key capabilities and pricing
125B total parameters with just 6B activated per token — unmatched cost-efficiency, per Qwen's announcement.
GDN + QSA hybrid attention, gated residual, N-gram embedding, and the Muon optimizer — a precursor to Qwen4.
Native 262K-token context window, extensible to 1M tokens via YaRN scaling.
Same output ceiling as Qwen3.8 Max for substantial refactors.
Reasoning-capable for multi-step planning, coding, and analysis.
Processes images and visual context alongside code.
Cached input at $0.016/M reduces repeat-context cost.
Available in AdaL CLI and Desktop through the model selector
Use the coding environment where AdaL already understands your repo, tools, and task context.
Open the model selector and choose your model from the list.
Ask AdaL to inspect, plan, edit, test, and review with your chosen model powering the agent loop.
Key details about Qwen3.8 Flash
"Qwen3.8 Flash is a multimodal MoE model and an early preview of the Qwen4 architecture — open-weight, trained at 1/9 the cost of Qwen3.7-Plus while outperforming it, especially in coding and office tasks."
Qwen
"Per Qwen's announcement: 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI)."
Benchmarks
"Use Qwen3.8 Flash for fast iteration where a large native context window and vision support matter but budget is tight."
AdaL
"262K native context, extensible to 1M tokens with YaRN."
Context
"Available now in AdaL CLI and Desktop, 50% off launch week."
Availability
Open AdaL CLI or Desktop, run /model, and choose Qwen3.8 Flash.
Get Started with AdaL