Alibaba Qwen vs DeepSeek

Comparison

Alibaba Qwen vs DeepSeek compares the two Chinese model families that define the top of the open-weight field as of October 2026. Qwen, from Alibaba Cloud, is a broad family: the Qwen3.8 generation spans a 27-billion-parameter dense model under Apache 2.0, a sparse Flash-Next model, and a 2.4-trillion-parameter flagship whose weights are downloadable under a custom licence. DeepSeek ships a narrower line, currently DeepSeek-V4-Pro and DeepSeek-V4.1-Flash, and releases both under the MIT licence.

The headline differences are licensing and range. DeepSeek's largest models carry the most permissive licence available; Qwen's most permissive licence is on its smaller model, which is also the only one of the group sized for fine-tuning on modest hardware. On hosted APIs, both vendors offer one-million-token context and prices far below Western frontier models, with DeepSeek cheaper at the flagship tier and the two close at the low-cost tier. Benchmark figures below are vendor-reported from model cards and are not directly comparable across vendors.

Feature Comparison

DimensionAlibaba QwenDeepSeek
DeveloperAlibaba Cloud (Qwen team)DeepSeek
Hosted flagship (API)qwen3.8-maxdeepseek-v4-pro (DeepSeek-V4-Pro-0813)
Largest open weightsQwen3.8-2.4T-A95B: 2.4T parameters, 95B activated; text-only, thinking mode required (released 12 Aug 2026)DeepSeek-V4-Pro-0813: 1.7T parameters, mixture-of-experts (13 Aug 2026)
Licence, largest modelCustom "qwen3.8-max" licence; terms not reviewed hereMIT
Efficient open modelQwen3.8-Flash-Next: 125B with 6B activated; Qwen Community License 1.0DeepSeek-V4.1-Flash: 552B backbone, 8B active at prefill and 16B at decode; MIT (API release 10 Sep 2026)
Small dense open modelQwen3.8-27B: 27B, Apache 2.0, image and video input (14 Aug 2026)None in the current V4 line
Context length (open weights)262,144 tokens native, extensible to about 1,000,000Up to 1,000,000 tokens
API context / max outputPriced to 1M input tokens; max output not publicly documented in the pages reviewed1M context; 384K max output
Flagship API price per 1M tokensqwen3.8-max: $2 input / $6 output (Singapore region)deepseek-v4-pro: $1.32 input (cache miss) / $3.96 output at peak; half that off-peak
Low-cost API price per 1M tokensqwen3.8-flash: $0.15 input / $0.47 outputdeepseek-flash: $0.30 input (cache miss) / $1.20 output at peak; $0.15 / $0.60 off-peak
DiscountsContext caching; 50% batch discount where batch is supportedCache-hit input from $0.003 (flash) and $0.022 (v4-pro) off-peak; off-peak rates are half of peak
API formatsNot verified in the pages reviewedOpenAI and Anthropic formats; Responses API support added in 2026

Detailed Analysis

Model line-ups

Qwen3.8 arrived in August 2026. Its model cards describe a hybrid architecture that interleaves Gated DeltaNet linear-attention layers with gated attention, in three sizes. Qwen3.8-27B is a dense vision-language model with thinking mode on by default. Qwen3.8-Flash-Next is a sparse model listed at 125B parameters with 6B activated, plus a 51B n-gram embedding and a 4B multi-token-prediction module. Qwen3.8-2.4T-A95B activates 95B of its 2.4 trillion parameters. Alibaba states that the hosted Qwen3.8-Max is "the official version based on Qwen3.8-2.4T-A95B with more features", including vision input, a non-thinking mode, one-million-token context by default and built-in tools, so the downloadable weights are not identical to the product.

DeepSeek's cadence in 2026 was a V4 preview on 24 April, a public-beta V4-Flash update on 31 July, a generally available V4-Pro on 13 August and V4.1-Flash on 10 September. The V4 technical report (April 2026) gives V4-Pro as 1.6T parameters with 49B activated and V4-Flash as 284B with 13B activated; the August V4-Pro-0813 model card lists 1.7T. V4.1-Flash is a new design, a 40-layer causal encoder-decoder that accepts images and text, and the retired V4-Flash API names now route to it.

Licences

This is the clearest dividing line. Both current DeepSeek models are MIT-licensed, which permits commercial use, modification and redistribution with only a notice requirement. Qwen uses three licences across one generation. Qwen3.8-27B is Apache 2.0. Qwen3.8-Flash-Next uses the Qwen Community License 1.0, which requires a separate commercial licence for companies running a model-as-a-service or AI work-assistant business, and requires the model name to be displayed prominently in products exceeding 100 million monthly active users or US$20 million in monthly revenue. The 2.4T model carries a licence tagged "qwen3.8-max" whose text could not be opened for this comparison; anyone planning commercial use should read it first.

For most businesses the practical reading is that DeepSeek's weights and Qwen3.8-27B can be used without negotiation, while Qwen's larger weights need legal review, especially for anyone reselling inference.

API pricing

List prices as of 6 October 2026, per million tokens. On Alibaba Cloud Model Studio's Singapore deployment, qwen3.8-max is $2 input and $6 output; qwen3.7-plus is $0.40 and $1.60 up to 256K input tokens, rising to $1.20 and $4.80 beyond; qwen3.8-flash is $0.15 and $0.47. DeepSeek introduced peak and off-peak pricing in August 2026. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays. At peak, deepseek-v4-pro is $1.32 input on a cache miss and $3.96 output, and deepseek-flash is $0.30 and $1.20; off-peak prices are half of those.

So DeepSeek's flagship is about a third cheaper than Qwen's at peak and about two-thirds cheaper off-peak, while Qwen's flash tier undercuts DeepSeek's at peak and roughly matches it off-peak. DeepSeek's cache-hit input prices are far lower again, which favours agent loops that resend a long, stable prefix. Unit prices are not cost per task: a model that reasons at length before answering generates more billable tokens for the same job. See AI model pricing for how to compare on a per-task basis.

Context and efficiency

Both families now treat a one-million-token context as standard. DeepSeek's V4 report attributes this to architecture: hybrid compressed sparse attention that, by its account, needs 27% of the single-token inference FLOPs and 10% of the KV cache of its predecessor at million-token length. The V4.1-Flash card claims a further reduction to roughly a quarter of V4-Flash's KV cache. Qwen3.8 open weights are native to 262,144 tokens and extensible to about one million. All of these are vendor-reported.

Benchmarks and fine-tuning

Model-card scores should be read as claims. On Terminal Bench 2.1, Alibaba reports 73.0 for Qwen3.8-27B and DeepSeek reports 87.9 for V4-Pro-0813 and 90.6 for V4.1-Flash at maximum reasoning effort; Alibaba reports 86.6 on Terminal Bench for the 2.4T model. On SWE-bench Pro, Alibaba reports 61.7, 62.5 and 67.7 for its three sizes. Harnesses, effort settings and benchmark versions differ, and no independent evaluation was reviewed here.

For teams that want to adapt a model, size settles the matter. Qwen3.8-27B is the only current model from either vendor that is dense, permissively licensed and small enough for common fine-tuning and LoRA workflows. Its card recommends SGLang and vLLM for serving. DeepSeek's V4.1-Flash activates few parameters per token, but its 552B backbone still has to be held in memory, which puts it in multi-GPU server territory.

Best For

Fine-tuning on your own hardware

Alibaba Qwen

Qwen3.8-27B is dense, Apache 2.0 and 27B parameters. DeepSeek's current open models start at a 552B backbone.

Redistributing or reselling inference on large open weights

DeepSeek

V4-Pro and V4.1-Flash are MIT-licensed. Qwen's larger weights require a separate licence for model-as-a-service businesses.

Lowest-cost flagship API

DeepSeek

deepseek-v4-pro lists at $1.32 / $3.96 per million tokens at peak and half that off-peak, against $2 / $6 for qwen3.8-max.

High-volume, low-cost API calls

Depends on timing

qwen3.8-flash is $0.15 / $0.47 at all hours. deepseek-flash is $0.30 / $1.20 at peak and $0.15 / $0.60 off-peak, with very cheap cache hits.

Agent loops with long, repeated prefixes

DeepSeek

Cache-hit input is priced at a small fraction of cache-miss input, and the API accepts both OpenAI and Anthropic formats.

Image and video understanding in an open model

Alibaba Qwen

Qwen3.8-27B and Flash-Next accept image and video input. DeepSeek-V4.1-Flash accepts images and text; the V4-Pro card lists text generation only.

A range of sizes from one vendor

Alibaba Qwen

Three open Qwen3.8 sizes plus max, plus and flash API tiers, against two DeepSeek models.

Self-hosted million-token context

Either

Both publish open weights with roughly one-million-token context. Hardware budget and licence terms decide.

The Bottom Line

Qwen and DeepSeek have ended up with complementary strengths. DeepSeek offers the simplest proposition in open models: two current models, both MIT-licensed, both with one-million-token context, and the cheaper flagship API. Qwen offers breadth: a permissively licensed 27B model that ordinary teams can fine-tune and serve, larger weights for those who accept custom terms, and a three-tier commercial API.

The honest caveat is that almost everything quantitative here comes from the vendors. Prices are list prices on 6 October 2026 and have moved several times this year; DeepSeek cut API prices with V4.1-Flash in September. Benchmark numbers are self-reported under different conditions. Parameter counts for the same DeepSeek model differ between its April report and its August card.

A workable rule: choose Qwen3.8-27B when the plan is to own and adapt a model, DeepSeek when the plan is to call or host a large model with minimal licensing friction, and test both hosted flagships on real tasks, measuring cost per completed task, before committing to either.