- PrismML has released Bonsai 27B, which includes 1-bit and ternary builds of the Qwen3.6-27B model that can run on laptops and iPhones.
- Ternary Bonsai 27B uses weights of {-1, 0, +1} at a true 1.71 bits per weight and has an ideal size of 5.9GB.
- 1-bit Bonsai 27B employs binary weights of {-1, +1} at 1.125 bits per weight, resulting in a size of 3.9GB.
- Both models are multimodal and maintain a context of 262K tokens, with approximately 75% of Qwen3.6-27B's attention being linear.
- Ternary Bonsai 27B retains 94.6% of the FP16 baseline, while 1-bit Bonsai 27B retains 89.5%.
- The 1-bit build allows for phone-local reasoning, measuring 672 tokens per 1% of iPhone battery.
PrismML has unveiled Bonsai 27B, a significant advancement in mobile AI technology, featuring both 1-bit and ternary Qwen3.6-27B models. The ternary model utilizes weights of {1, 0, +1} at a compact 1.71 bits per weight, with a total size of 5.9GB, while the 1-bit model employs binary weights of {1, +1} at 1.125 bits per weight, resulting in a size of 3.9GB.123
Both models are designed to run efficiently on iPhones, retaining up to 94.6% of the FP16 baseline performance for the ternary version and 89.5% for the 1-bit version. This is particularly crucial as iOS limits a single app's memory usage to about half of the device's physical memory, making the 5.9GB and 3.9GB sizes practical for deployment.6
The models support a context of 262K tokens, with the ternary model peaking at 14.7GB and the 1-bit model at 11.6GB during operation. Additionally, the 1-bit build measures 672 tokens1% of iPhone battery, showcasing its efficiency for mobile reasoning tasks.457
PrismML's innovations, including a DSpark drafter trained against the Bonsai 27B target, are set to enhance the capabilities of mobile AI applications significantly.
“The ternary Bonsai uses ±1,0 weights at 1.71 bits per weight (5.9GB ideal), and the 1-bit version uses binary ±1 weights at 1.125 bits (3.9GB) — both multimodal with 262K token context. On iPhone, the 1-bit build achieves 672 tokens per 1% battery; a DSpark drafter boosts H100 inference to 143.8 tok/s.”
