Taste, Trust, and Truth: How to Evaluate AI-Powered Discovery
Once a discovery system starts speaking, accuracy alone is not enough. You have to evaluate fit, trustworthiness, and product impact together.
Why the old evaluation model is not enough
For years, discovery systems were largely silent. Search results appeared. Recommendation rows updated. Ranking models made decisions behind the scenes. The product surfaced something compelling or it did not, but the system itself made few explicit claims.
That is changing.
AI-powered discovery can explain recommendations, ask follow-up questions, compare options, and guide a user through a decision. This can make discovery more useful, but it also changes what failure looks like.
Once the system starts speaking, correctness is no longer the only issue. Trust becomes visible.
A traditional recommendation can be irrelevant. An AI-powered recommendation can be irrelevant, misleading, overconfident, or inappropriate—and present the failure in fluent prose.
Clicks, starts, completions, conversion, and retention remain essential. They are no longer sufficient. A persuasive explanation can earn a click on a weak or misrepresented recommendation, allowing short-term engagement to conceal long-term trust damage.
In discovery, that is a dangerous trade.
Truth comes first
When I evaluate AI-powered discovery, I keep returning to three words: taste, trust, and truth.
Truth is the clearest place to begin. Is the item real and available? Is it described accurately? If the assistant says a show is a comedy-drama, has a strong female lead, or runs under two hours, can the product prove that claim? If it explains why something was recommended, is the explanation grounded in approved attributes and signals?
The fastest way to undermine an AI-powered discovery experience is to let it be eloquently wrong.
Taste is not the same as relevance
Factual correctness does not guarantee a good recommendation. The system still has to understand the user’s mood, context, and constraints. It needs enough range to support exploration without becoming generic, and enough focus to be useful without becoming repetitive.
Taste is hard because there is rarely one correct answer. A recommendation can be technically relevant and still feel wrong. That is why human judgment remains part of evaluation.
Trust sits across everything
Trust depends on both truth and taste. Does the assistant speak with appropriate confidence? Does it admit uncertainty? Does it know when to clarify and when to stop talking? Can the user correct it without fighting the interface? Does it respect age, tone, context, and control?
An AI discovery product earns trust by being consistently useful and appropriately honest. It loses trust when it overclaims, fabricates, or turns conversation into friction.
A four-layer evaluation framework
I would evaluate AI-powered discovery in four layers:
- Factual integrity: Every recommendation, explanation, and comparison should be testable against catalog reality, policy constraints, and product state.
- User fit: Was this a good answer for this user, in this moment, given the request they made? Fit includes relevance, nuance, diversity, novelty, and context sensitivity.
- Interaction trust: Was the assistant calibrated, clear, and controllable? Did it clarify only when useful and recover cleanly when misunderstood?
- Product outcome: Did the experience produce more satisfying discovery, fewer dead ends, better downstream engagement, and durable trust?
The order matters. Outcome metrics should not excuse failures in the layers above them.

What the framework catches
Consider a user asking: “I want something funny but not silly, under 90 minutes, that I can watch with my 12-year-old.”
A result can fail each layer differently:
- Truth: The assistant recommends a 104-minute title or describes mature content as family-friendly.
- Fit: Every constraint is technically satisfied, but the slate is dominated by broad children’s comedy and misses the request for something adults would enjoy too.
- Trust: The assistant claims, “You’ll love this because you watch it every summer,” without evidence, or hides uncertainty about the age rating.
- Outcome: The explanation earns a click, but the family abandons the title quickly and reformulates the request.
This is why a single engagement metric cannot tell the product team what went wrong. Retrieval may have produced a weak slate. The generated explanation may have misrepresented a good candidate. The assistant may have understood the request but handled uncertainty poorly. Each failure needs a different fix.
The scenario also shows why average scores are dangerous. A system may perform well on broad requests and still fail badly when age, availability, or duration becomes a hard constraint. Those cases should not disappear inside an overall relevance number. Some failures are quality problems; others make the response invalid enough that it should never reach the user.
How to operationalize the work
- Separate retrieval quality from response quality. A brilliant explanation cannot rescue a weak candidate set, and a strong slate still fails if the assistant misrepresents it.
- Build representative scenarios. Include ambiguous requests, cold-start users, group viewing, family contexts, unavailable titles, contradictory preferences, multilingual queries, and vague mood-based prompts.
- Use human evaluation with a real rubric. Score factual accuracy, fit, explanation quality, diversity, tone, and trustworthiness.
- Read transcripts and paths. Review where the assistant got close but missed. Look for overconfidence, unnecessary clarification, repetition, narrowing, and drift across turns.
- Run online experiments with guardrails. Measure reformulation, dead ends, fallback usage, dissatisfaction, and longer-term trust—not engagement lift alone.
The rubric should preserve those distinctions. Factual integrity can often be checked against structured sources. Taste requires calibrated human judgment and representative users. Trust needs transcript-level review because confidence and recovery emerge through interaction. Product outcome belongs in experiments, but only after the earlier layers establish that the experience is safe and meaningful enough to test.

Common mistakes
I would watch for four mistakes in AI discovery evaluation:
- Treating eloquence as quality. Fluent language can make a weak recommendation feel stronger than it is.
- Over-indexing on click-through rate. A persuasive explanation may increase clicks without increasing satisfaction. The metric will lie to you before the user tells you the truth.
- Ignoring catalog hygiene. Many apparent model failures are metadata, availability, knowledge, or instrumentation failures exposed by the conversational surface.
- Letting the model evaluate itself. Self-grading can assist diagnosis, but it cannot replace human judgment and real outcomes where fit and trust matter.
Automated graders still have a role. They can enforce schemas, compare outputs with catalog facts, identify missing constraints, and make large evaluation sets more manageable. The mistake is asking the same kind of model that produced a persuasive answer to become the sole authority on whether that answer was tasteful or trustworthy.

Evaluate the promise
Evaluation should be part of product design, not the final QA step. Teams need living scenarios, explicit failure taxonomies, human rubrics, and evidence that separates being helpful from being impressive.
Those scenarios should evolve with the product. New catalogs, markets, interaction patterns, and policy constraints create new failure modes. A frozen benchmark eventually rewards familiarity with yesterday’s problems instead of readiness for tomorrow’s users.
In content discovery, a great recommendation is not only one that gets the click. It is one that makes the user think, “This system gets me.” An AI-powered experience can strengthen that feeling or damage it quickly because the system is now making explicit claims.
Once a product begins to explain itself, it begins to make promises. Those promises need more rigorous evaluation, not less.
A recommendation must be true before it can be tasteful, trustworthy, or successful.