TL;DR: Qwen3.8-27B — the dense 27.8B that dropped on August 14 — fits on a single RTX 3090 at W4A16 and decodes at ~125–155 tokens/sec. Splitting it across two cards with tensor parallelism buys roughly 20% more decode (noisy: +13% in one run, +22% in the next) and 14% faster prefill on average on a cold prompt (23% at the top of the ladder: an... Weiterlesen
Intelligence View
Qwen3.8-27B on One RTX 3090 vs Two: +20% Decode, +14% Cold Prefill, and 3x on Cached Prompts
TL;DR: Qwen3.8-27B — the dense 27.8B that dropped on August 14 — fits on a single RTX 3090 at W4A16 and decodes at ~125–155 tokens/sec. Splitting it across two cards with tensor parallelism buys roughly 20% more decode (noisy: +13% in one r…
SOCIAL SHARE CARD GENERATOR