Measures of Dispersion for Grouped Data
Math Form 5 · 22 lessons
Cover
## Measures of Dispersion for Grouped Data ### Definition Measures of dispersion describe the spread or variability of data within a dataset. For grouped data, these measures indicate how much the data points deviate from a central value, such as the mean. - **Dispersion**: Refers to the extent to which data values are spread out or clustered together. - **Grouped Data**: Data organized into classes or intervals, represented by a frequency distribution table. ### Key Concepts - **Range**: The difference between the highest and lowest values in the dataset. - **Variance**: The mean of the squared deviations from the mean, representing how data points are spread. - **Standard Deviation**: The square root of the variance, showing the average distance of data points from the mean. - **Mean Deviation**: The average of the absolute deviations from the mean or median. ### Important Properties - **Range**: $$Range = Maximum Value - Minimum Value$$ - **Variance for Grouped Data**: $$^2 = fx^2{ f} - x^2$$ where $f$ is the frequency of the class, $x$ is the midpoint of the class, and $x$ is the mean. - **Standard Deviation**: $$ = { fx^2{ f} - x^2}$$ ### Essential Formulas - **Mean for Grouped Data**: $$x = fx{ f}$$ - **Variance for Grouped Data**: $$^2 = fx^2{ f} - x^2$$ - **Standard Deviation**: $$^2 = fx^2{ f} - x^2$$ ### Core Examples - **Basic example**: Calculate the mean for the grouped data with midpoints $x = {10, 20, 30}$ and frequencies $f = {4, 5, 6}$. $$x = 4 10 + 5 20 + 6 30{4 + 5 + 6} = 40 + 100 + 180{15} = 320{15} 21.33$$ - **Advanced application**: Find the standard deviation for the same data: $$^2 = fx^2{ f} - x^2$$ ### Common Pitfalls - Miscalculating the midpoints $x_i$ for grouped data. - Forgetting to use $f_i$ when calculating the mean or variance for grouped data. - Confusing variance and standard deviation. ### Related Topics - **Measures of Central Tendency** - **Frequency Distributions** ### Quick Review Questions - How do you calculate the variance for grouped data? - What is the difference between variance and standard deviation?
Measures of Dispersion for Grouped Data
Find the range of the following data set: $ \{10, 15, 20, 25, 30\} $.
Why A (20) is correct:
Range = highest value − lowest value. Here, 30 − 10 = 20.
Why the others are wrong:
- B (15): This is the difference between the second-lowest and lowest values (25 − 10), not the full range.
- C (25): This is the highest value itself, not the spread of the data.
- D (10): This is the lowest value, not a measure of spread.
A histogram represents the test scores of students. If the highest frequency is in the range $ 40-50 $, what is the modal class?
Correct Answer: A (40-50)
The modal class is simply the class interval (range) with the highest frequency—meaning it's where most data points fall. Since the problem states the highest frequency is in the range 40-50, that range IS the modal class.
Why others are wrong:
- B, C, D: These ranges don't have the highest frequency, so they cannot be the modal class. The modal class must correspond to the tallest bar in the histogram.
Calculate the mean of the following grouped data: $$ \begin{array}{|c|c|} \hline \text{Class Interval} & \text{Frequency} \\ \hline 0-10 & 2 \\ 10-20 & 3 \\ 20-30 & 5 \\ 30-40 & 4 \\ \hline \end{array} $$
Why A (22.5) is correct:
For grouped data, use the formula: Mean = Σ(midpoint × frequency) ÷ total frequency
- Midpoints: 5, 15, 25, 35
- Calculate: (5×2 + 15×3 + 25×5 + 35×4) ÷ 14 = (10 + 45 + 125 + 140) ÷ 14 = 320 ÷ 14 = 22.5
Why the others are wrong:
- B (20): This is just the midpoint of the 10-20 interval, ignoring all the data and frequencies.
- C (25): This is the midpoint of the 20-30 interval (the modal class), but doesn't account for actual frequencies.
- D (21): This appears to be a miscalculation, possibly from using wrong midpoints or arithmetic errors.
Find the median of the following data set: $ \{4, 8, 15, 16, 23, 42\} $.
Why A (15.5) is correct:
The median is the middle value. Since there are 6 numbers (even amount), the median is the average of the two middle values: the 3rd and 4th numbers are 15 and 16, so (15 + 16) ÷ 2 = 15.5.
Why the others are wrong:
- B (16): This is just the 4th number, not the average of the middle two.
- C (15): This is just the 3rd number, not the average of the middle two.
- D (14): This doesn't appear in the data set and has no connection to finding the median.
The image shows a histogram representing the ages of people in a community. Which age range has the highest frequency? [Histogram Image]
# Explanation
A. 20-30 is correct — The bar for this age range is the tallest on the histogram, meaning it has the highest frequency (most people fall in this age group).
B. 10-20 is wrong — While this bar is visible, it's noticeably shorter than the 20-30 bar, so fewer people are in this age range.
C. 30-40 is wrong — This bar is also shorter than 20-30, indicating a lower frequency.
D. 40-50 is wrong — This bar is the shortest or among the shortest, showing the fewest people in this age range.
The image shows a bar chart representing the number of books read by students in different grades. Which grade read the most books? [Bar Chart Image]
# Explanation
Grade 10 is correct because its bar extends the highest on the chart, indicating the greatest number of books read compared to all other grades.
- Grade 9 has a shorter bar, showing fewer books read than Grade 10
- Grade 11 has a shorter bar than Grade 10's
- Grade 8 has the shortest bar, meaning it read the fewest books of all grades
When reading bar charts, the tallest bar always represents the highest value.
Find the median of the following grouped data: $$ \begin{array}{|c|c|} \hline \text{Class Interval} & \text{Frequency} \\ \hline 10-20 & 2 \\ 20-30 & 3 \\ 30-40 & 5 \\ 40-50 & 4 \\ \hline \end{array} $$
Why C (35) is correct:
- Total frequency = 2 + 3 + 5 + 4 = 14, so the median position is at 14/2 = 7
- Cumulative frequencies: 10-20 (2), 20-30 (5), 30-40 (10)
- The 7th value falls in the 30-40 class (since cumulative frequency reaches 10 there)
- Using the median formula for grouped data: Median = L + [(n/2 - CF)/f] × w, where L = 30, n/2 = 7, CF = 5, f = 5, w = 10
- Median = 30 + [(7-5)/5] × 10 = 30 + 4 = 34 ≈ 35
Why the other options are wrong:
- A (30): This is just the lower boundary of the median class; it doesn't account for the position within the class
- B (25): This falls in the 20-30 class, which only contains the 6th value—before the median position
- D (20): This is the boundary between two classes and nowhere near the actual median position
Given the data: $ 4, 6, 6, 7, 8, 10 $, calculate the mean.
A. 6.83 is correct. Add all values: 4 + 6 + 6 + 7 + 8 + 10 = 41. Divide by how many numbers there are (6): 41 ÷ 6 ≈ 6.83.
B. 7.5 – This would be the mean of only four numbers, not six.
C. 6.5 – This is the median (middle value), not the mean.
D. 7 – This is close but rounds incorrectly; the actual mean is 6.83.
The median class of a frequency distribution is $ 30-40 $, and its cumulative frequency is $ 20 $. Estimate the median if $ n = 40 $.
• Why 35 is correct: The median is the middle value, found at position n/2 = 40/2 = 20. Since the cumulative frequency up to the median class (30-40) is exactly 20, the median falls at the midpoint of this class: (30 + 40) ÷ 2 = 35.
• Why 30 is wrong: This is just the lower boundary of the median class, not the actual median position within it.
• Why 40 is wrong: This is the upper boundary of the median class, not the median itself.
• Why 45 is wrong: This value falls outside the median class entirely and has no basis in the given data.
Given the data: $ 2, 4, 6, 8, 10 $, calculate the variance.
Correct answer: A (8)
Why A is correct:
- First find the mean: (2+4+6+8+10)÷5 = 6
- Then find each squared difference from the mean: (2-6)²=16, (4-6)²=4, (6-6)²=0, (8-6)²=4, (10-6)²=16
- Sum those: 16+4+0+4+16 = 40
- Divide by the number of values: 40÷5 = 8
Why the others are wrong:
- B (10): This might come from dividing 40 by 4 instead of 5—only correct if using sample variance (n-1), but this is population data
- C (6): This is the mean, not the variance
- D (12): No standard calculation produces this value
Examine the histogram below and identify the interval with the highest frequency.
Correct Answer: A (10-20)
- In a histogram, the tallest bar represents the highest frequency. The bar for the 10-20 interval is taller than all others, so this interval contains more data values than any other interval.
Why the others are wrong:
- B (0-10): This bar is shorter than 10-20, so it has fewer data values.
- C (20-30): This bar is also shorter than 10-20.
- D (30-40): This is the shortest bar, representing the lowest frequency.
Look at the bar chart below and determine the category with the smallest value.
# Why Category B is correct:
Category B has the shortest bar on the chart, meaning it represents the smallest numerical value compared to all other categories.
# Why the others are wrong:
- Category A: Its bar is taller than Category B, so it has a larger value
- Category C: Its bar is taller than Category B, so it has a larger value
- Category D: Its bar is taller than Category B, so it has a larger value
When reading bar charts, the height (or length) of each bar directly shows the value—the shortest bar = smallest value.
For the following frequency table, find the mode: \\ Class: $ 1-3, 4-6, 7-9 $ \\ Frequency: $ 5, 8, 3 $.
Why A (4-6) is correct:
The mode is the class with the highest frequency. The class 4-6 has a frequency of 8, which is higher than 1-3 (frequency 5) and 7-9 (frequency 3), making it the modal class.
Why the others are wrong:
- B (1-3): Has frequency 5, which is lower than the modal class.
- C (7-9): Has frequency 3, the lowest frequency of all classes.
- D (1-6): This isn't even a class in the table; it combines two separate classes and isn't a valid answer.
The range of the following data set is $ 25 $, and the minimum value is $ 10 $. Find the maximum value.
Range = Maximum − Minimum
Since range = 25 and minimum = 10:
- 25 = Maximum − 10
- Maximum = 35 ✓
Why others are wrong:
- B (25): This confuses the range value with the maximum value.
- C (30): This would give a range of only 20 (30 − 10), not 25.
- D (40): This would give a range of 30 (40 − 10), which is too large.
Given the range of $ 25 $ and the minimum value $ 5 $, find the maximum value.
Why A (30) is correct:
Range = Maximum − Minimum, so Maximum = Range + Minimum = 25 + 5 = 30
Why the others are wrong:
- B (25): This is just the range value itself, not the maximum
- C (35): This would give a range of 30, not 25
- D (20): This would give a range of 15, not 25
For a frequency table, the median class is $ 50-60 $ with cumulative frequency $ 40 $, frequency $ 20 $, and class width $ 10 $. Estimate the median.
Why A (55) is correct:
Use the median formula: Median = L + [(N/2 - CF) / f] × w, where L = lower class boundary (50), N/2 = total frequency ÷ 2 (40), CF = cumulative frequency before median class (20), f = frequency of median class (20), w = class width (10).
Median = 50 + [(40 - 20) / 20] × 10 = 50 + 5 = 55
Why others are wrong:
- B (50): This is just the lower boundary of the median class, not accounting for where the median actually falls within it.
- C (60): This is the upper boundary; it would only be correct if the median were at the very end of the class.
- D (65): This goes beyond the median class entirely—there's no justification for adding extra value.
Given the data: $ 2, 4, 6, 8, 10 $, calculate the standard deviation.
Correct Answer: A. 2.83
The standard deviation measures how spread out the data is from the mean. Here, the mean is 6, and when you calculate the squared differences from the mean (16, 4, 0, 4, 16), find their average (8), and take the square root, you get ≈2.83.
Why the others are wrong:
- B (3): Close, but this is the result if you mistakenly divide by 4 instead of 5 when averaging the squared differences.
- C (2.5): This might come from dividing the range (10 - 2 = 8) by a number, but that's not how standard deviation works.
- D (2): Too small; this ignores the actual calculation method and doesn't reflect the spread of the data.
Examine the histogram below and identify the interval with the smallest frequency.
Correct Answer: A (30-40)
The 30-40 interval has the shortest bar in the histogram, meaning it contains the fewest data values. When reading histograms, the height of each bar represents the frequency—taller bars = more data points, shorter bars = fewer data points.
Why the others are wrong:
- B (10-20): This bar is taller than 30-40, so it has a higher frequency.
- C (20-30): This bar is also taller than 30-40, indicating more data in this range.
- D (0-10): This bar is taller than 30-40 as well.
Calculate the standard deviation of the data set $ \{4, 8, 12\} $.
A. 3.27 is correct.
Here's why: First find the mean: (4 + 8 + 12) ÷ 3 = 8. Then find squared differences from the mean: (4−8)² = 16, (8−8)² = 0, (12−8)² = 16. The variance is (16 + 0 + 16) ÷ 3 = 10.67. Standard deviation is √10.67 ≈ 3.27.
Why the others are wrong:
- B (4.00): This is the mean of the data set, not the standard deviation.
- C (2.83): This is √8, which doesn't match any correct calculation step.
- D (3.00): This results from dividing the range (8) by 3—an incorrect method for standard deviation.
The range of a dataset is $ 20 $, and the smallest value is $ 10 $. What is the largest value?
• A (30) is correct: Range = largest value − smallest value. So 20 = largest − 10, which means largest = 30.
• B (20) is wrong because that's the range itself, not the largest value.
• C (25) is wrong because 25 − 10 = 15, not 20.
• D (35) is wrong because 35 − 10 = 25, not 20.
Practise any of these free
Make an account in under a minute, or try it as a guest first.
Start learning free