Market research professionals have long followed the guideline that you need to have a minimum of 30 samples to make any meaningful inferences and to perform statistical tests.

We’ve all worked with footnotes in small font sizes that warn us “small base, interpret with caution” when comparing results with N<30 samples.

But where does the magic number “30” come from?

The answer to this has its roots in the Central Limit Theorem

The Central Limit Theorem (CLT) is one of the most fundamental and widely used theorems in statistics with applications in market research (e.g., Z-tests, t-tests, etc.).

The theorem states that as your sample size increases to a sufficiently large number, the distribution of the sample mean will always be normally distributed.  

Now this may sound complex, but let’s unpack this first using a general example and then contextualize it to market research. 

First, what do we mean by normal distribution?

A normal distribution in its most simplest definition is a symmetrical bell shaped curve in which most of the data lies around the middle of the range, while the rest taper off gradually and equally towards both extremes.

The height of people is a common example of a normal distribution. This means that the majority of the population will have a height that falls around the average or mean height, with fewer people being either extremely tall or extremely short.

This is an example of a normal distribution. Most of the data is clustered near the mean, while the rest of the data become less frequent as we go farther away from the mean.

The shape of the normal distribution curve can differ depending on how spread out or varied the values are.

For instance, the normal distribution curve for newborn babies is narrow with a higher peak because most newborn babies fall within a smaller range of heights with little variation. However, for adults, there are more possibilities and greater variations in height, so the normal distribution curve will be wider and its peak shorter.

Now that we know what a normal distribution is, let’s take a closer look at the Central Limit Theorem.

Simplifying The Central Limit Theorem

Say you want to find out the average height of Singaporean women. You start by taking a random sample of N = 20 women, calculate the mean of these samples, and come up with a single average value of the N =20 women as seen below.

Now this doesn’t look too interesting and doesn’t tell us much. So let’s obtain more random N = 20 sample sets from the population, calculate means of each set, and plot the means on a histogram. We now start getting an interesting pattern! We see that the average heights range from 148 cm to 155 cm, but most cluster around 151 – 152 cm.

The plot of the means of a few different N=20 sample sets of women starts to paint a more interesting story

Now, as we continue to further draw more and more random samples of women from the population and continue plotting the sample means, we get the below graph as the outcome and it magically transforms into a normal distribution with the mean being 152 cm (corresponds to the peak at the center) which is a good approximation of the entire population of women in Singapore!

As you continue to randomly draw more samples from the population and plot the means, the curve starts to look like a normal distribution curve

This in essence is the Central Limit Theorem. The magical thing about CLT is that even if the original distribution is not normal (for instance it may be right or left skewed – we will see an example from market research below), the plot of the means will be normally distributed.

The Central Limit Theorem states that if we take a large number of repeated samples from a population, calculate the mean of each sample, and then plot the distribution of those means, the resulting distribution will be approximately normal.

Now the important thing to remember is that the Central Limit theorem refers to the plot of the MEANS OR AVERAGE of the samples and not a plot of the raw data points. In our example it means that we plot the average of the height of different samples, not the raw heights.

But why do I care about the data being normally distributed in market research?

The most important practical application of normal distribution in market research is that it allows us to do statistical tests such as t-tests, Z-tests, regression and a whole suite of other statistical analyses.

Statistical tests are used to determine whether the differences between two or more groups are statistically significant or not. For example, you may want to know if there is a statistically significant difference in the satisfaction levels of two different groups of customers. To determine this, you could use a statistical test such as the Z-test, which assumes a normal distribution.

The assumption of normal distribution is important because it allows us to make certain assumptions about the data that we are analyzing.

However, how good a normal distribution curve it is to derive precise estimates of the target population will depend on the sample size. And this is where the magic number 30 comes in. Let’s look at an example contextualised to Market Research below.

How does this apply to Market Research?

Let’s say there are N = 10,000 subscribers of a streaming platform. This is your entire target population of interest. You now want to measure the satisfaction levels of these 10,000 subscribers on a 5-point Likert scale.

In an ideal world you could survey the entire target population (N = 10,000) and derive a distribution of the satisfaction scores which might look something like the chart below.

Hypothetically, let’s say the mean satisfaction score of the entire target population is around 6.78, with the majority scoring it 7 or 8. However, there’s a “long tail” of people who give it a lower rating, such as 3, 2 or even 1. This is not normally distributed (as is the case with most satisfaction ratings that skew more positive).

This is a hypothetical distribution of the satisfaction score for the entire target population (N = 10,000). Satisfaction scores among users typically skew towards the positive end.

In market research, surveying the entire population is not feasible. Hence, the approach we take is to randomly draw a sample and then estimate the satisfaction score for the target population.

Imagine that you take a small sample of the population. You randomly select N = 10 subscribers and ask them their satisfaction score. You might get something like this, with the mean score being 3.7. Now this is very far off from the population mean and is not a good estimate but that’s expected since the sample size is only N = 10.

Suppose that you repeat this procedure 10 times, taking samples of N = 10 subscribers each time, and calculating the mean of each sample. This is a sampling distribution of the mean.

If you repeat this process many more times, the distribution will look something like this:

The sampling distribution isn’t normally distributed because the sample size of each set isn’t sufficiently large for the central limit theorem to apply.

However, as the sample size increases, the sampling distribution looks increasingly similar to a normal distribution, and the spread decreases:

So why does N = 30 come into picture?

The number 30 has an interesting (and debated!) relation with normal distributions. Remember at the, the Central Limit Theorem states that no matter what the distribution of the actual population is (it may be totally random and may not follow any distribution), the sample means will follow normal distribution if the sample size is large enough. This “large enough” has been statistically accepted to be 30 – meaning that if we draw samples of size 30 or more (n >30), the sample means shall follow a normal distribution, irrespective of how random the original population was.

The sampling distribution of the mean for samples with n = 30 approaches normality.

When the sample size is increased further to n = 100, the sampling distribution follows a normal distribution.

The moment it hits the threshold of 30 it enters something known as standardised normal distribution