Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs

6 points by vagabund 2 years ago · 1 comment

Reader

vagabundOP 2 years ago

"Hawk-3B exceeds the reported performance of Mamba-3B (Gu and Dao, 2023) on downstream tasks, despite being trained on half as many tokens. Griffin-7B and Griffin-14B match the performance of Llama-2 (Touvron et al., 2023) despite being trained on roughly 7 times fewer tokens."

Settings

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs

Keyboard Shortcuts