Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

LLMs easily use a lot of RAM, and these systems are MUCH, MUCH cheaper (though slower) than a GPU setup with the equivalent RAM.

A 4-bit quantization of Llama-3.1 405b, for example, should fit nicely.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: