1 points | by onnies 9 hours ago
2 comments
I might be dumb, can someone explain to me how these are supposed to be useful? What if I want to tackle one of these myself, manually?
I expand the sections and it’s all full of prompt slop
This benchmark measures an agents ability to solve SRE/On-call tasks!
Which expanded sections are slop? Happy to point you in the right direction - we also have a github repo: https://github.com/abundant-ai/incident-arena
which has the tasks in harbor format (instruction.md files etc)
I might be dumb, can someone explain to me how these are supposed to be useful? What if I want to tackle one of these myself, manually?
I expand the sections and it’s all full of prompt slop
This benchmark measures an agents ability to solve SRE/On-call tasks!
Which expanded sections are slop? Happy to point you in the right direction - we also have a github repo: https://github.com/abundant-ai/incident-arena
which has the tasks in harbor format (instruction.md files etc)