Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The gap was MUCH larger in the past, but in my tests, oMLX and llama.cpp are now very similar (within 10%) in both prompt processing and generation speed. GGUF ecosystem provides a better selection of quants, in my experience Unsloth ones are excellent.


I thought the main advantage of oMLX is it's less likely to invalidate the KV cache when working with coding agents, which is key when working on a Mac because of the slower prompt processing.


llama-server also supports saving the kv cache to SSD. I had no issues with cache invalidation using pi.


TIL! When was this functionality added? It wasn’t in llamacpp when I looked in June




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: