Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Something from previous discussions about the whole

"Rather than paying $50k up front for 8 x A100's, you can just rent some GPU's for $1.2k to train the whole thing in 4 days"

That feel off to me is that it completely ignores the compute time spent exploring new ideas, failing, tweaking the training data, etc.



Link for the previous discussion? Which model, dataset, training strategy?




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: