WASTE Inference Engine Open Sourced: Running Kimi K3 on a 64GB MacBook
According to Dotion Beating monitoring, the edge database company SQLite AI has open-sourced the WASTE inference engine. It allows the preservation of all layers and the expert Kimi K3 to run on a 64GB MacBook Pro. The converted model is approximately 1TB in size, with a speed of 0.49 to 0.54 Tokens per second.
WASTE keeps around 27GB of the model backbone in memory. Over 80,000 experts are placed on the built-in SSD. For each generated Token, only the currently invoked expert is read from the disk.
This version does not involve distillation, pruning, or expert removal. However, the expert weights have been re-quantized to 3 bits, while the model backbone adopts 4-bit and 8-bit.