Bonsai Distillation Explained: From Qwen3.6-27B to a Phone-Friendly 3.9 GB Model
August 04, 2026A step-by-step look at how Prism ML distilled a 27B parameter model into a 3.9 GB phone-capable file while keeping 89.5% of full-precision reasoning.
A step-by-step look at how Prism ML distilled a 27B parameter model into a 3.9 GB phone-capable file while keeping 89.5% of full-precision reasoning.
I ran the 1-bit Bonsai 27B model locally on a laptop. Here's what happened when a 3.9 GB GGUF tried to write code, solve math problems, and keep up with full-precision models.
A technical walkthrough of how Prism ML distills Qwen3.6-27B into binary and ternary weights, keeps reasoning alive, and runs it on a phone.
Is Kimi distilled from Claude? We investigate the evidence behind Kimi distillation — output similarity, Anthropic's terms of service, community consensus, and what it means for developers using Kimi distilled models.
A technical dive into how Moonshot AI's Kimi-k models and the art of distillation are powering the near-instant coding experience in Cursor.