view article Article Fine-tuning LLMs to 1.58bit: extreme quantization made easy Sep 18, 2024 โข 216
Transformer-Lite: High-efficiency Deployment of Large Language Models on Mobile Phone GPUs Paper โข 2403.20041 โข Published Mar 29, 2024 โข 35