Why AI Can Fix Bugs but Still Struggle to Find Them

https://hackernoon.imgix.net/images/2jqChkrv03exBUgkLrDzIbfM99q2-c08223j.png

My experiments with GLM-5.3-flash, and the difference between following instructions and knowing what to question.

I had this crazy idea - what if I can find such algorithm which can find all the bugs in any software, literally ALL. And while going though it (quite successfully actually!), it helped me understand the difference in intelligence of small vs large models.

If TLDR - if you have clear task definition, or at least enough clues, cheaper models can do quite serious work now. If the task require to invent questions themselves - thats what require true intelligence, and cheap models fall short. Also majority of public benchmarks is a form of cheating.

I was working on bisecting on how exactly benchmarks like SWE bench work, how the eval looks like, what is in the golden data. And what I found is that it does not really ask a question to find...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more