modelsSEP 12 02:46 UTC
DeepSeek v4.1 Flash runs at 23 seconds per token on 16GB M1 Mac Mini
A user report says the DeepSeek v4.1 Flash model can be loaded and run locally on a 2020 M1 Mac Mini with 16GB of unified memory, but generation is extremely slow at roughly 23 seconds per token. That pace makes the setup impractical for interactive use, though it shows the model can technically execute on older consumer hardware. The result highlights how limited RAM and memory bandwidth constrain local inference of large models.