← Lcoalhost
Pushing the Limits: Extreme Inference Speedup of Qwen 3.8 27B on NVIDIA B300 (100 to 10k+ tok/s)
Source :
DEV · #llm
See it live in context on Lcoalhost →