Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

they are not really optimized for 'wide variety of models' . what optimization did pi do for glm 5.3?


Tool calling success rate, in the case of omp.


Can't edit my post anymore, but here's the omp blog from February talking about improved tool calling rates across 15 models, with only the harness being tweaked to get the improvements.

https://stencil.so/blog/the-harness-problem

Three GLM models are mentioned, but so is Deepseek, Grok, Minimax, Kimi, and Gemini.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: