Hacker News·4 min read·hard
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
M
marcobambini✦AI Summary
Developers have created WASTE, an inference engine that allows a 2.78 trillion parameter model to run on a consumer laptop. By streaming experts from disk rather than loading the entire model into RAM, it achieves a functional, albeit slow, performance of 0.5 tokens per second.
Kimi K3 — 2.78 trillion parameters — running on a consumer laptop.
technologyai
✦
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in