Settings

Theme

GLM-5.3-Flash at 3.3 tok/s

2 points by marcobambini · 0 comments · 1 min read


A few months ago, I created the WARP engine (formerly WASTE) to run Kimi K3, the complete 2.78-trillion-parameter model, on macOS. GLM-5.3-Flash shares many architectural similarities with Kimi K3, so I added support for it as well. It requires as little as 5.14 GB of RAM to run, and on a 64 GB MacBook Pro M5 Pro it reaches about 3.32 tok/s, or 3.86 tok/s on longer runs.

More memory means a larger expert cache, while higher storage and memory bandwidth can further improve performance.

The project is completely open-source and free to use: https://github.com/sqliteai/warp

No comments yet.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection