← Lcoalhost

Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe

Source : DEV · #llm

See it live in context on Lcoalhost →