Settings

Theme

A new inference engine to run Kimi K3 2.78T parameter with 29GB of RAM

marcobambini.substack.com

4 points by marcobambini · 2 comments

Reader

1 thread
armchairhacker

Current speed is “approximately one third of a token per second”

  • marcobambiniOP

    Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection