← Lcoalhost
Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe
Source :
DEV · #llm
See it live in context on Lcoalhost →