antirez (@antirez) on X

X (formerly Twitter) ·

1 min read Original article ↗

antirez on X: "That's a bit slow. Streaming directly the K3 official hugging face 1.6TB of weights in mxfp4 in an m5 max 128gb. https://t.co/E4YVbkSgoS"

  • user avatar

    That's a bit slow. Streaming directly the K3 official hugging face 1.6TB of weights in mxfp4 in an m5 max 128gb.

  • user avatar

    After minutes :D Btw this is just to get the inference graph dialed in. Q2 across two Mac Studios with 512GB can work at acceptable speed for chat at least. K3 was trained in MXFP4, I have the feeling it will quantize fine. I'll wait to have access to two Mac Studio 512GB.

  • user avatar

    Slow but works! Have you tried on M3 Ultra?