Accuracy on your questions. Checked against real people.

Put your questions to AI twins and the real people behind them. Review both sets of answers side by side, person by person, with every match and miss visible.

Your question

Which product concept would you most likely buy?

Human-to-twin validation study

1Bring your questions and audiences.
Bring the questions you want answered and the people you need to hear from. We build the test with you during kickoff.
2We ask real people in our panel.
We field those questions to the real people behind the twins, with video of every answer for you to review.
3See where the twins match.
Go through the results question by question and see which twins agreed with their human and which missed.

REAL-WORLD OUTCOMES

Predicted. Then observed.

Four results from the Replica Performance Whitepaper, shown only against what happened in the real world.

Most subscribers stayed put.

Twins forecast that 84% would absorb a $3 increase, versus 94% observed.

91% accuracy

OutcomeObservedTwins
Accept increase94%84%
Switch to ads<1%14%
Cancel5%2%

Travelers kept booking.

Twins predicted 93% would still book after a 10% fare increase, versus 91% observed.

98% accuracy

OutcomeObservedTwins
Still book91%93%
Skip9%7%

Matched the iPhone mix.

Which iPhone twins chose, compared with the published U.S. sales mix.

97.6% accuracy

OutcomeObservedTwins
Pro Max27%29.8%
Pro25%27.1%
Air6%0%
1722%22.7%
Older20%20.4%

Picked the winning ad.

Twins chose Ad B, the creative that delivered 3.53x ROAS versus 2.52x for Ad A.

72% picked the winner

OutcomeROASTwins
Ad A2.52x28%
Ad B3.53x72%
01 / 01

Human Preference Benchmark.

The first public and transparent human preference benchmark.

Coming

ModelScore
Rehearsals twins
94.0%
Claude Opus 5
78.7%
Gemini 3.8 Flash
77.5%
Gemini 3.1 Pro Preview
76.9%
GPT-5.6 Terra
76.0%
Grok 4.20 (Non-Reasoning)
71.6%

Trust and verification

Build trust with every Rehearsal.

Every validation adds evidence. Compare each new Rehearsal with relevant studies to see how much of that proof applies.

How this Rehearsal compares with measured validation evidence.

Evidence overlap

How closely the audience, task, stimulus, and question structure match evidence already in the library.

Accuracy score

Explicit human-agreement or observed-outcome accuracy from the comparable studies, weighted by relevance.

Overlap
Accuracy
Overlap
Accuracy

Audit the answer

Every claim keeps its receipts.

You can trace any score to its evidence. Read the underlying twin transcripts, open the rationale for a response, ask the agent about any claim, or launch a deeper analysis when the first pass is not enough.

Every report comes graded. Explore the validation evidence behind the grade and see where the findings have limits.

Read what every twin said. Open the full conversation for every member of the cohort. Expand why they said this to connect an answer back to the interview evidence and lived context that produced it.

Question any claim in place. Highlight a sentence, chart, or finding and ask the agent to explain, challenge, or re-cut it without leaving the report.

01 / 01

Prove it first

Bring the question you are least willing to get wrong.

Every four-week pilot includes a human-to-twin validation. We start it in week one and show you the misses as clearly as the matches.