Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
Maple-Preview, a machine learning model, has achieved a high performance of 120 tokens per second on an iPhone. This is significant as it demonstrates the potential for efficient AI processing on mobile devices. The model is based on a ternary 20B MoE architecture. Engineers may be interested in exploring this technology for future projects. Further details can be found at the provided article URL.