Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up?

Noam Brown (OpenAI) "Implications of Large-Scale Test-Time Compute" https://xcancel.com/i/article/2064210146558136827



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: