πŸ” Search
Sign in to post
Show HN: Moe-Direct – MoE Models far larger than your RAM, on a consumer desktophttps://github.com/tmxkzm1925-max/MoE-Direct

I wanted to try using the larger models on my computer (32GB RAM, RTX 5080, Gen5 NVMe), but the best I could do was around 30B. So I started with the idea that it might be possible by taking advantage of the fact that MoE models use only some of the experts rather than all of them. MoE-Direct essentially uses the three layers of SSD, RAM, and VRAM instead of residing entirely in memory, caching only the necessary experts in RAM and making the model usable even with resources far smaller than required. In my environment, I obtained the following decode results: Kimi K2.6: 1.03 tok/s. Qwen3.5-1…

β†—

0trust.social media

Loading your media...

Pick a GIF β€” Giphy

Loading GIFs...