← Back to the board
qwen38-27b-rtx3090
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
github.com/syv-ai/qwen38-27b-rtx3090+232%
vs its 7-day baseline
- 24h signal
- 20
- Sources hot
- 1
- First seen
- 5h ago
Why it's blowing up
No explainer yet. They are generated for the board's top movers every few hours.
48h attention curve
Weighted engagement in the trailing 24h, sampled hourly.