Pre-launch Opening to the first product teams

Product evaluations that tell you what to change.

Eva helps your team find out whether your product, AI included, does what your users need. You agree on what good looks like, your users and Eva measure it, and every result comes with the code behind it and a proposed change.

Eva's mascot: a small terminal on two legs, holding a screwdriver
How it works

Four steps, repeated until it works for your users.

01

Read it

Eva reads your code and maps its components, so every result can be traced to the code behind it.

02

Define it

Together you set the criteria that matter to your users, each with a pass mark, before any results come in.

03

Measure it

Your users answer short questions inside your product, and Eva judges your AI's actual output against the same criteria.

04

Act on it

For every criterion that isn't met, Eva proposes a change to the code. Your team reviews it and decides.

Built on peer-reviewed research

Finding problems is easy. Knowing what to change is harder.

In a CHI 2026 study, Eva's co-founder and colleagues interviewed people who build AI products. Almost all of them collected evaluation results and then got stuck: a low score doesn't say whether the prompt, the retrieval or the model is to blame, so nothing changed.

It isn't only us, and it isn't new. In 2005, Hornbæk and Frøkjær found that developers got more out of concrete redesign proposals than out of lists of usability problems. In 2022, Zhou et al. found more than half of NLG practitioners using as many metrics as possible despite knowing their limits. In 2025, Nahar et al. found teams at Microsoft with no effective way to evaluate their models beyond basic health checks.

Eva is built around that step. Every criterion that isn't met comes with a proposed change to the code behind it.

Criteria first

Criteria and pass marks are agreed before any results come in, so the goalposts don't move.

Honest uncertainty

Every result shows its range. “Not enough observations” is a valid answer, not a pass.

A checked judge

Where Eva judges output, it reports its own judging error next to the result.

Measure, act, measure again

One evaluation shows where you stand. The next shows whether the change helped.

Define it4 topics · 3 valuesNorth Star saved, v2.0Read it38 componentsscanned 11 days agoMeasure it9 rows · 212 of 225 observationscollecting, about 3 days leftAct on itNo actions yetproposed once results are inIteration 2Collecting
Iteration 1First measurementMatrix v1
4 met3 not met2 not enough observations
Iteration 2After the proposed changes were mergedMatrix v2
6 met ▲22 not met1 not enough observations
met not met not enough observations
Why we're building Eva
Most evaluation of AI products is a spreadsheet of opinions, or a benchmark that measures something other than what users need. Neither tells you what to change. We're building Eva so that every result can be traced to the code behind it, and every change is measured again.
The founders of Eva
Request access

Be one of the first teams to evaluate with Eva.

Eva is in pre-launch. Tell us a little about your product and we'll get back to you.

We only use this to contact you about Eva.