Settings

Theme

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

arxiv.org

13 points by theanonymousone · 1 comment

Reader

1 thread
DiabloD3

The title of the paper is correct. The paper does not seem to actually get to the point in a generic way, but hyperfocuses on, effectively, one type of error compensation.

Highly quantized models, especially with highly quantized KV caches, will, effectively, attend to the wrong tokens and be unable to easily discern highly similar tokens. The bastardized way of explaining this is gradient descent techniques get stuck in localized minimum and global maximums, so what happens when you turn the slopes into hard stair steps?

We need to move to smaller models and smaller caches and better samplers, not new quant methods (although I'm willing to also take those too).

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection