4 min read
What is significance?
(Statistical) significance, also called the P(robability) value, stands for the probability that the result found was found while the null hypothesis (there is no effect, association or difference) cannot be rejected. In other words, significance stands for the probability that the result found is based on chance, and when this probability is small enough, the null hypothesis can be rejected and the alternative hypothesis accepted.
This is a follow-up to 'What is a hypothesis?'

Science is all about providing evidence for rejecting or not rejecting the null hypothesis (there is no effect, association or difference) and accepting the alternative hypothesis (there is an effect, association or difference).
Do you want to learn what significance is with the help of a clear example? Check: https://significantie.coenfirmationbias.nl/
Calculating the P-value
To decide whether or not we reject the null hypothesis, the statistical significance of the result is calculated. Statistical significance is also called the P(robability) value. A statistical program calculates the significance on the basis of two principles:
-
The expected value is calculated. In other words, if the result were purely down to chance, what would the expected value be? And how large would the spread around it be? These are the expected values.
-
Then it is calculated how far the results found deviate from the expected value.
The further the results found deviate from the expected value, the smaller the p-value, and the greater the chance that the null hypothesis can be rejected and the alternative hypothesis accepted. In other words, the result found is not chance and is therefore significant.
Science is uncertainty and probabilities
This establishes a fundamental principle of science: we never know anything 100% for certain. It is always about probabilities and uncertainties. Scientists have agreed among themselves that we accept 5% uncertainty. When the p-value (probability of chance) is smaller than 5% (P<0.05), we say: the probability is large enough (>0.95) to state that there is a real effect, association or difference, so the null hypothesis can be rejected and the alternative hypothesis accepted. This means that there is a 5% chance that you reject the null hypothesis while it actually cannot be rejected. In other words, the evidence shows that someone is guilty (accepting the alternative hypothesis) while that is not the case. We call this a Type 1 error: we think there is an association, difference or effect while there is not. When we do not reject the null hypothesis while we should have, we make a Type 2 error.
Example
Let's take a hypothesis as an example.
‘People who take a vitamin C pill catch a cold less often than people who do not.’
Suppose you have 200 people and make two groups: one group takes a vitamin C pill every day (100) and the other group does not (100). The people are then followed for 5 years, and every year it is measured how many people had a cold.
First we look at the expected value if it were down to chance. Let's assume that about 25% of people catch a cold each year. This means that we expect 25 people in both groups to catch a cold each year. But we can expect a natural spread in this: one year 20 people catch a cold and another year 15. Still, on average 25 people catch a cold per year. The result:
| | Vitamin C | Control |
|---|---|---|
| Had a cold | 25 | 25 |
| No cold | 75 | 75 |
You can see this as the null hypothesis: there is no difference between the vitamin C group and the control group.
But when we carry out the study, we find the following:
| | Vitamin C | Control |
|---|---|---|
| Had a cold | 15 | 25 |
| No cold | 85 | 75 |
Now we certainly see a difference between the two groups. However, to calculate whether this is a significant difference, we do not only look at the differences between the groups. We also look at the differences within the groups. (Were there on average 15 people with a cold in the vitamin C group, but in some years also 35? Then there is a large spread.)
To keep it short: if the differences within the groups are small (enough) and the differences between the groups are large (enough), the P-value will come out low (and thus indicate significance). We then state that the difference (the effect of the vitamin C) is not down to chance.
The null hypothesis can be rejected and the alternative hypothesis accepted.
Confidence interval
A confidence interval is a statistical measure that provides a range of values within which we expect to find the true value of a parameter in the population, with a certain degree of certainty. The interval gives not only an estimate of the parameter, but also an indication of the uncertainty around that estimate.
How does it work?
When calculating a confidence interval, a sample of data is usually used. On the basis of this sample, a point estimate is made (for example the mean) and then a margin of error is added, which depends on the standard deviation and the sample size. The result is an interval that indicates the probability that the true value lies within this range.
Confidence
The confidence interval is usually presented with a confidence level, such as 95% or 99%. A 95% confidence interval means that if we were to repeat the same sampling procedure many times, 95% of the calculated intervals would contain the true parameter value.
Example
Suppose we want to estimate the mean weight of a population and we calculate a 95% confidence interval of 70 kg to 80 kg. This means that we can say with 95% certainty that the true mean weight of the population lies between these two values.
Conclusion
A confidence interval is a useful tool in statistics that helps to quantify the uncertainty of an estimate. It offers not only an estimate, but also a context for the reliability of that estimate.
