How many votes do you need for a reliable website test?

Why 10 votes is the minimum, what 20 and 50 votes add, and how to tell a real difference from noise. With the margins of error worked out.

Updated October 1, 2026

Ten votes tell you roughly where your homepage stands. Fifty votes tell you whether a change made a difference. Every vote is one person's impression, and people differ. The more votes, the closer the average gets to what visitors in general think. This guide shows how much closer, so you can pick the right test size.

Why one opinion is not enough

Ask one person whether they trust your website and you learn what that person thinks. The next person may answer the opposite. Only when you collect many answers does a pattern appear that you can act on. That is the idea behind StayOrBounce: many short votes on the same four questions, instead of one long review.

The margin of error, worked out

Any result based on a sample has a margin of error: the range in which the "true" score probably lies. A common choice is a 95% margin, which you calculate as 1.96 times the standard error.

For the stay score, the share of reviewers who would stay, the margin is largest when the score is around 50%. That is the worst case:

Votes Margin on a stay score of 50%
10 about ±31 percentage points
20 about ±22 percentage points
50 about ±14 percentage points

For the average score per question on the 1 to 5 scale, the margin depends on how much reviewers disagree. If their answers have a standard deviation of 1 point, the margins are:

Votes Margin on an average
10 about ±0.6
20 about ±0.4
50 about ±0.3

These are approximations, and with very few votes they are rough. But they show the pattern clearly: to halve the margin, you need four times as many votes.

What each test size is good for

  • 10 votes: a first reading. Enough to spot a big problem, such as an average of 2 on Do I understand what this is?. Not enough to tell a 3.4 from a 3.7. StayOrBounce marks a result as reliable from 10 votes; below that, you see it as an interim result.
  • 20 votes: a solid picture. The averages per question become stable enough to see which of the four is your weakest point.
  • 50 votes: comparing versions. If you want to know whether a new headline beats the old one, you need this many. Differences between two versions are usually smaller than the margin at 10 votes.

Comparing two results needs more votes

When you compare two campaigns, both results have a margin of error, so the margin on the difference is larger than on either result. With 10 votes on each side, two averages need to differ by almost a full point before you can call it a real difference. With 50 votes on each side, about 0.4 points is enough.

StayOrBounce does this calculation for you. When you compare two campaigns, each question is marked as clearly better, clearly worse or within the noise, based on the votes and the spread of both results. Read did your redesign work? for how to set up such a comparison.

Spread matters too

Two homepages can both average 3.5 on trust: one because everyone answered 3 or 4, the other because half answered 1 and half answered 5. The second one splits your audience, and its average is less certain. That is why every result shows the spread per question next to the average.

More votes are not the only way to more certainty

  • Target the right people. Votes from your actual audience are worth more than more votes from everyone. You can target by country, age group or role.
  • Keep the questions the same. Fixed questions make every vote comparable with every other vote. That is why StayOrBounce never changes them.
  • Filter out noise. Votes that look automatic are left out, and reviewers who often deviate from the group count for less. See how we score websites.