Why We Need Central Tendency and Variability in Statistics
In everyday life we are confronted with a lot of data – test scores, salaries, water use
and this can be confusing – therefore we need statistics to help us understand data
and make sense of these numbers.
Let’s say that we are looking at student marks in a classroom – we need to look at
what a typical value is – this is called central tendency. We also look at how different
these values are from each other – this is called variability.
Before we go further into this let us understand a few statistical terms.
Mean – The Mathematical Average
To calculate the mean, we take the sum of all the given values and divide it by the
number of values. This makes more sense when the values are all close to each
other and not at extremes. It could go wrong however when there is one huge
variation. For example, we are calculating the mean income of a neighbourhood –
the one billionaire who lives there will change the average by a lot. Or if we take a
cricket team, one bad series could lower the team’s average significantly.
Median – The Middle Value
Finding the median basically means picking the number that is neither very high or
low – this is useful because it gives a more realistic picture than the mean would if
we’re looking at what people earn in a neighbourhood.
Next we come to the Mode, which is the Most Frequently Occurring Value This
works very well to find popularity or categories for example the most popular pizza
topping. This does not give us the average but helps us understand what people
prefer.
Variability – How Spread Out is the Data?
Let’s go back to the earlier example of a class’s test scores. Now 2 classes may
have similar average scores. But let’s say one has everyone scoring around the
same number while the other shows quite a few differences. This is where variability
matters.
Range – Difference Between Highest and Lowest
The difference between the highest and lowest values in a given data set is called
range. To find it we subtract the minimum value from the maximum value. That is
Range = Maximum value – minimum value. Lets say that in a class test scores the
highest is 90 and the lowest is 40 – the range therefore is 90-40= 50. However this
can be distorted if there is a very high score or very low score that does not really
reflect the performance of the class.
To adjust for this we ignore the extremes and look for data in the interquartile range
or middle 50%. This is more reliable than range for real-life data. For example in the
class test scores, we remove the highest and lowest scores and look for data that is
between them.
Variance – How Far Values Spread from the Mean
Here we find the squared distance from the mean to show the spread of data. For
example lets say that height of 5 students in a group is (in cm)
160, 162, 163, 165, 180
First, we find the mean (average height):
160+162+163+165+180/5= 830/5 = 166 cm.
Then we check how far each student is from the mean and we square it
Height
160
162
163
165
180
x-mean
-6
-4
-3
-1
14
(x-mean)square
36
16
9
1
196
258
Now we divide by n-1 5-1=4
So 258/4= 64.5 cm
The variance of the height in this group is 64.5 which shows that it is quite spread
out – the lower the variance the closer the students are in height.
Standard Deviation – The Spread in Original Units
SD is merely the square root of the variance so using the above example – the
square root of 64.5 is 8.03 cam
This means that each students height is about 8cm away from the mean height of
166 cm. If the SD is smaller lets say 3 cm then everyone isi more or less the same
height – if it is bigger then there is a larger variation. This is an important measure of
variability.
Why Both Central Tendency & Variability Are Needed Together
A mean alone does not give us the real picture.
For example Class A marks: 70, 71, 72 (mean 71)
Class B marks: 20, 71, 99 (mean 71)
The mean is the same but the students performance is very different.
This is why we need central tendency – because it tells us how far each student is
from the mean.
Skewness – Why Mean Sometimes Lies
Skewness helps us understand whether the data is concentrated on one side of the
mean. The skew can be as follows:
Positive tail: mean gets pulled right
Negative tail: mean pulled left
Zero skewness means that the data is perfectly symmetrical.
Skewness helps to:
1. assess deviations
2.Identify where outliers could occur
3. Interpret results
Some real-life applications of what we’ve discussed could be government measuring
income inequality, hospitals studying recovery time variations etc.
Central tendency and variations therefore provide a balanced picture of data sets.
Central tendency simplifies data while variability gives us a more nuanced view of
this data. Together, they help us to compare groups fairly and make better decisions
based on evidence and data.