Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> so you had to have lots of operand re-use to not be memory-bound

Looking at Nvidia's spec sheet, an H100 SXM can do 989 tf32 teraflops (or 67 non-tensor core fp32 teraflops?) and 3.35 TB/s memory (HBM) bandwidth, so ... similar problem?



There is caching today.


The cache hitrate is effectively 0 for LLMs since the datasets are so huge.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: