PrismML releases Bonsai 27B — 1-bit and ternary Qwen3.6-27B builds that run on iPhones, retaining up to 94.6% of baseline
PrismML

PrismML releases Bonsai 27B — 1-bit and ternary Qwen3.6-27B builds that run on iPhones, retaining up to 94.6% of baseline

PrismML has launched Bonsai 27B, featuring 1-bit and ternary Qwen3.6-27B models optimized for iPhones, achieving up to 94.6% retention of baseline performance. The models utilize low-bit representations, with the ternary version at 5.9GB and the 1-bit version at 3.9GB, enhancing mobile AI capabilities.

CuriousCats Full Story

PrismML has unveiled Bonsai 27B, a significant advancement in mobile AI technology, featuring both 1-bit and ternary Qwen3.6-27B models. The ternary model utilizes weights of {1, 0, +1} at a compact 1.71 bits per weight, with a total size of 5.9GB, while the 1-bit model employs binary weights of {1, +1} at 1.125 bits per weight, resulting in a size of 3.9GB.123

Both models are designed to run efficiently on iPhones, retaining up to 94.6% of the FP16 baseline performance for the ternary version and 89.5% for the 1-bit version. This is particularly crucial as iOS limits a single app's memory usage to about half of the device's physical memory, making the 5.9GB and 3.9GB sizes practical for deployment.6

The models support a context of 262K tokens, with the ternary model peaking at 14.7GB and the 1-bit model at 11.6GB during operation. Additionally, the 1-bit build measures 672 tokens1% of iPhone battery, showcasing its efficiency for mobile reasoning tasks.457

PrismML's innovations, including a DSpark drafter trained against the Bonsai 27B target, are set to enhance the capabilities of mobile AI applications significantly.

Key Insight
“The ternary Bonsai uses ±1,0 weights at 1.71 bits per weight (5.9GB ideal), and the 1-bit version uses binary ±1 weights at 1.125 bits (3.9GB) — both multimodal with 262K token context. On iPhone, the 1-bit build achieves 672 tokens per 1% battery; a DSpark drafter boosts H100 inference to 143.8 tok/s.”
CuriousCats studied:
1
MarkTechPostMarkTechPost
“PrismML just released . It is a low-bit representation of Qwen3.6-27B, not a new pretrain.”
MarkTechPost →
Ask CuriousCats
What is PrismML's Bonsai 27B?
How does Bonsai 27B benefit iPhone users?
Why are low-bit quantization methods important?
Are there comparable models from other developers?
How do Bonsai's performance metrics stack against traditional models?
Become the most informed
person in the room.
Personal AI agents scanning 100,000+ sources — news, video, and social media — delivered every morning.
Download the App Go to CuriousCats.ai
🇺🇸 US🇮🇳 India🇬🇧 UK🇨🇦 Canada🇸🇬 Singapore
One story brought you here.
CuriousCats brings you everything else worth knowing.
Get CuriousCats