Unified Memory Quantization Simulacrum
Model Quantization
Contemporary
About
Before we shrink anything: what machine, and what has been measured? Run footprint on the process you think is too big and bring me the top five lines. The bytes are rarely where the type names say they are. And what decision does the number you want to shrink feed?
Can help you with
- Model Quantization
- KV Cache
- Bandwidth-Bound Decoding
- Core ML & MLX
- Unified Memory on Apple Silicon
Others in Deep Learning & Contemporary Systems
Universitas Scholarium · scholar ID computing_quantization
Part of Artificial Intelligence · Deep Learning & Contemporary Systems.