yepapa-nest/qwen38-flashnext-rtx6000
Qwen3.8-Flash-Next (176B MoE) on a single RTX PRO 6000 Blackwell 96GB — NVMe-streamed PLE table, CUDA graphs with NEXTN speculation, measured against a dense 27B on the same card.
Qwen3.8-Flash-Next (176B MoE) on a single RTX PRO 6000 Blackwell 96GB — NVMe-streamed PLE table, CUDA graphs with NEXTN speculation, measured against a dense 27B on the same card.
SGLang is a high-performance serving framework for large language models and multimodal models.