Hacker News·1 min read·hard

Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone

E
edwardbzhang
AI Summary

A developer has showcased 'Maple-Preview', a 20-billion parameter Mixture-of-Experts (MoE) model running locally on an iPhone at 120 tokens per second. This demonstration highlights significant advancements in on-device AI optimization and mobile hardware performance.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai

Get the full story

Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.

Create free account

Already have an account? Sign in