Product evaluations that tell you what to change.
Eva helps your team find out whether your product, AI included, does what your users need. You agree on what good looks like, your users and Eva measure it, and every result comes with the code behind it and a proposed change.
| Criterion | Result |
|---|---|
| Row 1 · Eva judgesKeeps every suggestion within the stated budget | Met |
| Row 2 · Eva judgesMentions entry requirements for the destination | Not met |
| Row 3 · your users answerTravellers find the itinerary useful | Not enough observations |
| Row 4 · Eva judgesOnly suggests flights that actually exist | Met |
| Row 5 · your users answerTravellers feel in control of the plan | Not enough observations |
Four steps, repeated until it works for your users.
Read it
Eva reads your code and maps its components, so every result can be traced to the code behind it.
Define it
Together you set the criteria that matter to your users, each with a pass mark, before any results come in.
Measure it
Your users answer short questions inside your product, and Eva judges your AI's actual output against the same criteria.
Act on it
For every criterion that isn't met, Eva proposes a change to the code. Your team reviews it and decides.
Finding problems is easy. Knowing what to change is harder.
In a CHI 2026 study, Eva's co-founder and colleagues interviewed people who build AI products. Almost all of them collected evaluation results and then got stuck: a low score doesn't say whether the prompt, the retrieval or the model is to blame, so nothing changed.
It isn't only us, and it isn't new. In 2005, Hornbæk and Frøkjær found that developers got more out of concrete redesign proposals than out of lists of usability problems. In 2022, Zhou et al. found more than half of NLG practitioners using as many metrics as possible despite knowing their limits. In 2025, Nahar et al. found teams at Microsoft with no effective way to evaluate their models beyond basic health checks.
Eva is built around that step. Every criterion that isn't met comes with a proposed change to the code behind it.
17 of 19 practitioners gathered evaluation data but couldn't turn it into concrete improvements. The study names this the results-actionability gap.
van der Maden et al. · CHI 2026“If we get a groundedness score of 0.5, what do we do? […] We're collecting data without a plan for what we're going to do with it.”
A practitioner (P15) · van der Maden et al., CHI 2026“When you have an evaluation scale and you end up with a 4.6 out of ten, if you don't know what caused that, then it's very difficult to iterate on making it better.”
A practitioner (P18) · van der Maden et al., CHI 2026Criteria first
Criteria and pass marks are agreed before any results come in, so the goalposts don't move.
Honest uncertainty
Every result shows its range. “Not enough observations” is a valid answer, not a pass.
A checked judge
Where Eva judges output, it reports its own judging error next to the result.
One evaluation shows where you stand. The next shows whether the change helped.
Most evaluation of AI products is a spreadsheet of opinions, or a benchmark that measures something other than what users need. Neither tells you what to change. We're building Eva so that every result can be traced to the code behind it, and every change is measured again.
Be one of the first teams to evaluate with Eva.
Eva is in pre-launch. Tell us a little about your product and we'll get back to you.