Junru Shao (@junrushao) on X

1 min read Original article ↗

Post

Post

  • user avatar

    (1/2) šŸ¦™ Buckle up and ready for a wild llama ride with 70B Llama-2 on a single MacBook šŸ’» 🤯 Now 70B Llama-2 can be run smoothly on an 64G M2 max with 4bit quantization. šŸ‘‰ Here is a step-by-step guide: mlc.ai/mlc-llm/docs/g… šŸš€ How about the performance? It's

  • user avatar

    Lots of folks will want to run an locally hosted model privately. All Macs can be a platform for private AI

  • user avatar

    What are the MMLU/HellaSwag benchmark scores for 4 bit quantized models like these. Maybe running a smaller model at higher precision is better than trying to cram a 70B model on a laptop.

  • user avatar

Don't miss what's happening

People on X are the first to know.