Human evals for AI
A structured way to run human evals on AI products, built because usability testing alone isn't enough.
Usability testing assumes a product behaves the same way twice. AI doesn't, so I've developed something else.
We are working on more and more AI products. Usability testing assumes a product behaves the same way each time, and AI doesn't. Ask it the same question twice and you can get two completely different answers. Which one is better, and why? It's hard to know without evals.
I didn't think "looks fine" was good enough. Before any automated eval can judge whether an AI response is good, someone has to decide what good actually means, in terms a reviewer can apply consistently, not just by feel. So I built a structured way to do that properly, starting with a human, before anything gets automated.
What I'm building
Right now it's a set of worksheets: a reviewer works through a fixed list of quick checks against real AI responses, so different people score consistently instead of each judging by feel. You can run it against one model, or compare several side by side.
What comes out the other end shapes two things: how the prompts get rewritten, and what the automated evals need to check for later.
I'm part of Torchbox's AI task force, and a smaller group inside it that focuses specifically on evals. It's a multidisciplinary team, working to make sure AI gets used well across Torchbox, for us and for our clients.
Where it's at
The worksheet version works. But it's still the first version, and it's slow: an expert has to work through a spreadsheet by hand.
The next stage is turning it into something faster and easier to use, so a client's own expert reviewers can pick it up directly. They won't need me to run it for them.
Where it's going
We work with charity and public sector clients across very different services, and more of them are bringing their own AI products to test. A product, not a one-off spreadsheet, means any of them can use it, and any researcher at Torchbox could run it too, without starting from scratch.
Responsible AI in research isn't a slogan for me. It's a framework other people can actually use to check their own work. I want to be one of the people building that future, not just reacting to it.