GLM-5.3-Flash shares many architectural similarities with Kimi K3, so I added support for it as well. It requires as little as 5.14 GB of RAM to run, and on a 64 GB MacBook Pro M5 Pro it reaches about 3.32 tok/s, or 3.86 tok/s on longer runs.
More memory means a larger expert cache, while higher storage and memory bandwidth can further improve performance.
The project is completely open-source and free to use: https://github.com/sqliteai/warp
0 comments