Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
HPsquared
on March 5, 2025
|
parent
|
context
|
favorite
| on:
Apple M3 Ultra
LLMs easily use a lot of RAM, and these systems are MUCH, MUCH cheaper (though slower) than a GPU setup with the equivalent RAM.
A 4-bit quantization of Llama-3.1 405b, for example, should fit nicely.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
A 4-bit quantization of Llama-3.1 405b, for example, should fit nicely.