Show HN: Warp – Run the 313B GLM-5.3-Flash on a MacBook with 8GB RAM

A few months ago, I created the WARP engine (formerly WASTE) to run Kimi K3, the complete 2.78-trillion-parameter model, on macOS.

GLM-5.3-Flash shares many architectural similarities with Kimi K3, so I added support for it as well. It requires as little as 5.14 GB of RAM to run, and on a 64 GB MacBook Pro M5 Pro it reaches about 3.32 tok/s, or 3.86 tok/s on longer runs.

More memory means a larger expert cache, while higher storage and memory bandwidth can further improve performance.

The project is completely open-source and free to use: https://github.com/sqliteai/warp

2 points | by marcobambini 1 hour ago

0 comments