Qwen 3.8 vs 3.5 generation time comparison on RTX 3090 (Test 3, Heavy 16K profile, NoThink mode)

Bar chart benchmark of text generation time for Qwen 3.8 9B, Qwen 3.5 27B and Qwen 3.8 27B in NoThink mode on an RTX 3090 24GB VRAM GPU across a short news article, a long-form piece and a vision task

The chart compares execution time for three task types — a short news item, a long-form article and a photo description — across three Qwen configurations in NoThink mode. On the RTX 3090 24GB the Qwen 3.8 27B model is the slowest while the compact 9B variant is markedly faster, a key factor for building a high-throughput article generation pipeline.

← Back to article