> ## Documentation Index
> Fetch the complete documentation index at: https://runrehearsals.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Audience Size and Twin Fidelity

Across our white paper and validation studies, Rehearsals twins are 92% accurate to what real people say. This document covers a different question: how much your results change as you add more twins to a Rehearsal.

The answer is not much. For a quick task, like comparing two emails, you don't need a big audience. Going bigger mostly spends credits without changing the answer. Our general advice is 50-100 twins per audience, and usually 1-3 audiences per study.

## What audience size buys you

|                                 | Rehearsal Twin                               | Human Survey Taker                         | Human In-depth Interview |
| :------------------------------ | :------------------------------------------- | :----------------------------------------- | :----------------------- |
| Depth of answer                 | High-medium (conversational, with follow-up) | Low (primarily multiple choice)            | High                     |
| Recommended size for most tasks | 50-100 per audience, usually 1-3 audiences   | 500                                        | 10-20                    |
| What the size buys you          | A stable ranking of options                  | A precise percentage (about ±4 pts at 500) | Depth and themes         |

In this test, an audience of 100 twins ordered options the same way the full 500-twin audience did 95% of the time. Going 5x bigger, from 100 to 500, changes very little. Unless you need to separate options that sit within a few points of each other, the extra size is overkill.

The same test also checked twins against real people. Those numbers are below. They measure one narrow thing, the order of option pairs, and are separate from the 92% accuracy figure above.

## How we tested it

We ran the same Rehearsal, with the same 14 multiple-choice questions about everyday buying behavior, multiple times. Every number here comes from one of those runs: 500 twins from the US general population. To see how smaller audiences do, we drew 2,000 random audiences of each size from those 500 twins and scored each one. We did not run a separate simulation at each size.

For the human benchmark, 228 real US adults answered the same 14 questions with the same options. They were an independent group, not the people the twins were built from.

We score ordering, because most decisions come down to which option wins and by roughly how much. A question with options A, B and C gives three pairs (A vs B, A vs C, B vs C), and for each pair we check which option more twins chose. There are 200 pairs across the 14 questions.

The consistency score is the share of pairs a smaller audience puts in the same order as the full 500-twin audience. The "95% at 100 twins" figure is this measure. The human-accuracy score is the same idea, but the order to match is the one the 228 humans gave, counting only the 146 pairs where humans showed a clear preference (a gap of at least 1.96 standard errors, too big to be a fluke of who answered). Ties count as half.

## What we found

### Results by audience size

| Twins in audience | Same order as 500 twins | Same #1 answer as 500 twins | Same #1 answer as a second, separate audience of the same size | Same order as clear human preferences | Weakest 5% of audiences vs humans | Same #1 answer as humans | Avg. gap from human percentages |
| ----------------: | ----------------------: | --------------------------: | -------------------------------------------------------------: | ------------------------------------: | --------------------------------: | -----------------------: | ------------------------------: |
|                10 |                     82% |                         83% |                                                            76% |                                   82% |                               77% |                      60% |                        14.0 pts |
|                20 |                     87% |                         88% |                                                            84% |                                   85% |                               81% |                      65% |                        12.6 pts |
|                50 |                     92% |                         94% |                                                            90% |                                   89% |                               85% |                      69% |                        11.8 pts |
|               100 |                     95% |                         96% |                                                            94% |                                   90% |                               88% |                      69% |                        11.5 pts |
|               150 |                     96% |                         98% |                                                            95% |                                   91% |                               89% |                      68% |                        11.5 pts |
|               250 |                     98% |                         99% |                                                            98% |                                   91% |                               89% |                      67% |                        11.4 pts |
|               500 |    100% (by definition) |        100% (by definition) |                                                            n/a |                                   92% |                               92% |                      67% |                        11.4 pts |

Each row summarizes 2,000 random audiences of that size. The 500 row is the full audience.

Consistency with the 500-twin audience rises from 82% at 10 twins to 95% at 100. The last 400 twins add only 5 more points. The top answer holds steady: at 50 twins, 94 of 100 random audiences picked the same #1 answer as the full 500, and two separate 50-twin audiences agreed on the #1 answer 90% of the time.

Agreement with people levels off sooner still. In this test, the share of clear human pairs in the same order goes from 89% at 50 twins to 92% at 500. That 92% is a narrower measure (pair ordering on 14 questions) than the 92% accuracy figure from our validation work, and the two matching is a coincidence.

Some gaps do not close with size. Twins picked the same #1 answer as humans about two thirds of the time, and their percentages sat about 11 points from the human percentages, at 50 twins and at 500 alike. Those gaps come from how twins answer, so more twins do not fix them. About 8% of clear human preferences were missed even with all 500 twins, mostly from a lean toward "none of these" answers: for example, 70% of twins said they hadn't had a large medical bill, compared with 20% of humans. When a result matters and the gap is small, check it another way: rerun with different question wording, compare against something you already know about your customers, or confirm with a small human study.

### Clear differences show up with very few twins

The wider the real gap between two options, the fewer twins you need to see it. Agreement with human preferences in this test, by the size of the human gap:

| Gap between two options among humans | Clear human pairs in this group | 10 twins | 20 twins | 50 twins | 100 twins | 500 twins |
| :----------------------------------- | ------------------------------: | -------: | -------: | -------: | --------: | --------: |
| 40 points or more                    |                              20 |    99.5% |     100% |     100% |      100% |      100% |
| 20 to 40 points                      |                              50 |      87% |      91% |      95% |       96% |       98% |
| 10 to 20 points                      |                              59 |      76% |      79% |      83% |       84% |       86% |
| Under 10 points                      |                              17 |      69% |      74% |      79% |       82% |       85% |

If one option is clearly preferred (a gap of 20 points or more), 50 twins will find it about 95% of the time or better. Larger audiences mainly help with close calls, and even there the gain is small.

## Sizing by task

Our general advice is 50-100 twins per audience, and usually 1-3 audiences. Some tasks call for a different size:

| Task                                                 | Why                                                                                                                                                                                                                                                                                                                                                                                                                     | Suggested size                                                                       |
| :--------------------------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :----------------------------------------------------------------------------------- |
| Picking a winner among 2-3 ads, emails or concepts   | When one option leads by 20+ points, 50 twins find it about 95% of the time. If the top two land within about 10 points, call it a tie or rerun at 100.                                                                                                                                                                                                                                                                 | 50-75 in one audience                                                                |
| Comparing which audiences would buy at a given price | Each audience is scored on its own, so the numbers above apply to each audience separately. Use the result to rank audiences. Don't read the purchase rate as a forecast; twins run more positive than people on purchase intent.                                                                                                                                                                                       | 50-100 per audience, 1-3 audiences                                                   |
| Uncovering themes across a general population        | Size decides how rare a theme can be and still show up. With 100 twins, a theme held by 5% of people appears at least 3 times in 88% of runs. With 250 twins, a theme held by 2% appears that often in 88% of runs. If you split the population, each audience of about 100 will only reliably surface themes held by roughly 5% or more of its members. This row rests on sampling math. This test did not measure it. | 250 in one general-population audience; if you split, up to 3 audiences of about 100 |

## Limits of this test

* One study, 14 multiple-choice questions, US general population. We did not test open-ended questions, narrow or niche audiences, or other countries. The theme guidance rests on sampling math, not this test.
* Smaller audiences were drawn from the 500-twin run rather than run separately, so run-to-run variation is not captured. Comparing against 500 twins also flatters larger audiences, since a 250-twin draw shares half its twins with the 500. The "second, separate audience" column avoids that overlap and shows the same pattern.
* Humans clicked buttons. Twins answered in conversation, and we mapped each reply to the matching option (1 correction after spot-checking; 2 replies out of about 7,000 could not be mapped). The 228 humans are also a sample, and a different group might land borderline preferences differently.
