Yup, this thing is basically useless. I just did /goal optimize the cuda kernel, and ai agent just proposed and tested a bunch of ideas by itself, and it profiled them using nsight ncu etc. by itself, which lead to one order of magnitude faster kernel.
respect for shipping something this technical solo, this is the kind of project that usually needs a team to even validate correctness. how are you handling regression testing across kernel variants, feels like the hardest part of an agentic optimizer isn't finding a faster kernel, it's proving the faster one didn't quietly break something
That is the funny part actually; we can either provide a reference kernel + input cases for correctness check. In this case at first it will use test harness to run reference kernel with reference inputs ad save the outputs as ground truth.
Or we can let AI to create a very basic reference implementation and input cases :D for my own experiments I used second one.
4 comments