RD ← Back to portfolio

AI PRODUCT DISCOVERY ASSISTANT · PUBLIC DEMO

Turn feedback into findings—without losing the evidence.

This prototype shows how AI can help a product team analyse qualitative feedback while keeping source evidence, interpretation and human judgement visibly separate.

Evidence linked Human reviewed Synthetic data No customer information

Realistic signals, deliberately synthetic data.

The demonstration dataset contains 40 fictional comments from finance users, administrators and managers across support, interviews and surveys.

40feedback records
3research sources
3personas
“Month end takes forever because I have to export the payment report and reconcile it against our accounting system manually.” F001 · Support · Finance
“I found a failed payment three days later while doing reconciliation. We should have known about it immediately.” F010 · Support · Finance
“Please don’t automate everything. I need to be able to check and approve financial adjustments.” F033 · Survey · Finance
View all 40 records ↗

For transparency: this public demo loads a pre-generated sample analysis. It does not call a live model or send data to a server.

Strong evidence tracing. Incomplete theme coverage.

The same 40-record synthetic dataset, model and prompt were used for each run. Every generated theme was mapped to a documented human reference before scoring.

100% citation validity Every cited feedback ID existed in the source dataset.
95.5% mean evidence precision Citations usually supported the theme they were attached to.
66.7% mean theme coverage Two of the three human-reference themes were found consistently.
What the model missed

All three runs found the two dominant operational problems but missed the smaller payment communication and audit-history theme. Relevant records were often retrieved but grouped into a broader payment-status theme.

Product decision

Treat theme separation as the next experiment and keep RAG deferred. The evidence points to a grouping and coverage problem—not a failure to retrieve the source records.

Review every run and limitation ↗

AI assists. Product judgement decides.

The underlying prototype can analyse an uploaded CSV with an LLM. The public version focuses on the higher-value product question: can a reviewer trace, challenge and govern the output?

View the product brief, decision log and implementation on GitHub ↗