Study Class 10 Mathematics Statistics with clear notes on mean, median and mode of grouped data, cumulative frequency tables and the three methods for finding the mean.
Statistics is the branch of mathematics that deals with collecting, organising, presenting and interpreting numerical data. In earlier classes you learned how to classify data into ungrouped and grouped frequency distributions, draw bar graphs, histograms and frequency polygons, and find the mean, median and mode of ungrouped data. This chapter extends those three measures of central tendency to grouped data, where observations are arranged in class intervals with frequencies. You will learn the direct method, the assumed mean method and the step-deviation method for finding the mean of grouped data, the formula for the mode of grouped data, and the formula for the median of grouped data using cumulative frequency. You will also learn how to build cumulative frequency tables of the less than type and the more than type, and how these ideas help in locating the median class and interpreting what the mean, median and mode actually tell us about a data set.
What you'll learn
1Recall how the mean, median and mode of ungrouped data are found
2Convert raw data into a grouped frequency distribution with suitable class intervals
3Find the class mark of a class interval and use it to represent the whole class
4Calculate the mean of grouped data by the direct, assumed mean and step-deviation methods
5Determine the modal class and compute the mode of grouped data using the mode formula
6Build cumulative frequency tables of the less than type and the more than type
7Locate the median class and compute the median of grouped data using the median formula
8Compare the mean, median and mode and decide which measure suits a given situation
Chapter at a glance
01Mean, Median, and Mode of Ungrouped Data
02Grouped Data and Class Intervals
03Mean of Grouped Data
04Mean of Grouped Data
05Mode and Median of Grouped Data
06Graphical Representation of Data
Detailed chapter notes
01
From Ungrouped to Grouped Data
In Class IX you studied how data is classified into ungrouped and grouped frequency distributions and how it is shown pictorially through bar graphs, histograms and frequency polygons. You also learned the three measures of central tendency for ungrouped data: the mean, the median and the mode. In this chapter the same three measures are extended to grouped data, where observations are placed in class intervals and each interval has a frequency. When data is very large, it is condensed into grouped form so that a meaningful study becomes possible. For example, marks of 30 students can be listed one by one, or the same marks can be arranged into class intervals such as 10-25, 25-40 and so on, with the number of students in each interval written as the frequency.
Grouped dataobservations arranged in class intervals with frequencies
Frequency of a classthe number of observations falling in that class
02
Mean of Grouped Data: Direct Method
The mean is the sum of the values of all observations divided by the total number of observations. For grouped data we do not know each individual observation, so we assume that the frequency of a class is centred at its mid-point, called the class mark. The class mark is found by averaging the upper and lower limits of the class. If the class marks are x₁, x₂, …, xₙ with corresponding frequencies f₁, f₂, …, fₙ, the mean is the sum of the products fᵢxᵢ divided by the sum of the frequencies. This is called the direct method. Note that while forming class intervals, a value equal to the upper class limit is counted in the next class, for example a student scoring 40 marks goes into the class 40-55 and not into 25-40.
Class mark = (upper class limit + lower class limit) ÷ 2
Direct methodx̄ = Σfᵢxᵢ ÷ Σfᵢ
The mean obtained from grouped data is approximate because of the mid-point assumption
03
Assumed Mean and Step-Deviation Methods
When the class marks and frequencies are large numbers, multiplying them becomes tedious. In the assumed mean method we choose one class mark as the assumed mean, denoted by a, and find the deviation dᵢ = xᵢ − a for every class. Then the mean is x̄ = a + (Σfᵢdᵢ ÷ Σfᵢ). The value of the mean does not depend on which class mark is chosen as a. If all the deviations share a common factor, we can divide each deviation by the class size h and write uᵢ = (xᵢ − a) ÷ h. The mean then becomes x̄ = a + h × (Σfᵢuᵢ ÷ Σfᵢ), which is the step-deviation method. All three methods give the same mean; only the amount of calculation differs.
Assumed mean methodx̄ = a + (Σfᵢdᵢ ÷ Σfᵢ), where dᵢ = xᵢ − a
Step-deviation methodx̄ = a + h × (Σfᵢuᵢ ÷ Σfᵢ), where uᵢ = (xᵢ − a) ÷ h
Choose the direct method for small numbers and the other two for large numbers
04
Mode of Grouped Data
The mode is the value that occurs most often, that is, the observation with the maximum frequency. For ungrouped data the mode can be read directly from the frequency table, but for grouped data we can only locate the class with the maximum frequency, called the modal class. The mode is a value inside the modal class and is calculated using Mode = l + [(f₁ − f₀) ÷ (2f₁ − f₀ − f₂)] × h, where l is the lower limit of the modal class, h is the class size, f₁ is the frequency of the modal class, f₀ is the frequency of the class preceding it and f₂ is the frequency of the class succeeding it. If more than one value has the same maximum frequency the data is called multimodal, but here we consider only single-mode problems.
Modal classthe class with the highest frequency
Mode = l + [(f₁ − f₀) ÷ (2f₁ − f₀ − f₂)] × h
The mode can be less than, equal to, or greater than the mean
05
Cumulative Frequency and Cumulative Frequency Distribution
The cumulative frequency of a class is obtained by adding the frequencies of all the classes preceding that class. A table showing cumulative frequencies is called a cumulative frequency distribution. There are two types. In the less than type, we record how many observations are less than each upper class limit; for example, the cumulative frequency of the class 10-20 tells us how many students scored less than 20 marks. In the more than type, we record how many observations are more than or equal to each lower class limit. Both tables describe the same data and either of them can be used to find the median of grouped data.
Cumulative frequency = sum of the frequencies of all preceding classes plus the frequency of the given class
Less than typeuses the upper limits of the classes
More than typeuses the lower limits of the classes
06
Median of Grouped Data
The median is the middle-most observation of the data. For ungrouped data, if n is odd the median is the ((n + 1) ÷ 2)th observation, and if n is even it is the average of the (n ÷ 2)th and ((n ÷ 2) + 1)th observations. For grouped data we first compute n ÷ 2 and locate the class whose cumulative frequency is greater than and nearest to n ÷ 2; this is the median class. Then Median = l + [(n ÷ 2 − cf) ÷ f] × h, where l is the lower limit of the median class, n is the total number of observations, cf is the cumulative frequency of the class preceding the median class, f is the frequency of the median class and h is the class size. The class intervals must be continuous before applying this formula.
Median classthe class whose cumulative frequency is greater than and nearest to n ÷ 2
Median = l + [(n ÷ 2 − cf) ÷ f] × h
About 50% of the observations lie below the median and about 50% above it
07
Choosing the Right Measure of Central Tendency
The mean is the most frequently used measure because it takes all observations into account and lies between the smallest and largest values, which makes it useful for comparing two or more distributions. However, extreme values can pull the mean away from the typical observation. When individual observations are not important and we want a typical value, the median is more appropriate, as in finding a typical productivity rate or average wage where extreme values may exist. When we need the most frequent or most popular item, the mode is the best choice, for example the most watched television programme or the most demanded consumer item. There is also an empirical relationship between the three measures: 3 Median = Mode + 2 Mean.
Meanuses all observations, good for comparison, affected by extreme values
Medianuseful when extreme values are present and a typical value is needed
Modebest for finding the most frequent or most popular item
Empirical relationship3 Median = Mode + 2 Mean
Want the complete chapter resources?Topic notes, quizzes and flashcards for Statistics.
What is the median class of a frequency distribution if the cumulative frequency just exceeds N/2 at that class?
AThe class where cumulative frequency equals N/2
BThe class where cumulative frequency first exceeds N/2
CThe class with the highest frequency
DThe class with the lowest frequency
Show answer
Answer: (B) The class where cumulative frequency first exceeds N/2
The median class is identified as the class interval where the cumulative frequency first becomes greater than or equal to N/2.
Question 04
What is a histogram used to represent?
AContinuous grouped data
BOnly names of students
CDays of the week
DCricket team players
Show answer
Answer: (A) Continuous grouped data
A histogram is a graphical representation used specifically for continuous grouped data where class intervals are shown on the x-axis and frequencies on the y-axis.
Question 05
Find the median of: 3, 1, 4, 1, 5, 9, 2
A1
B3
C4
D5
Show answer
Answer: (C) 4
Arrange in order: 1, 1, 2, 3, 4, 5, 9. Middle value (4th position) is 4.
Ready for more practice?Unlock the full quiz for this chapter.
Q1. The marks obtained by 10 students in a test are: 12, 15, 18, 20, 22, 25, 28, 30, 32, 35. Find the median marks.
Show model answer
Model answer
Arrange marks in ascending order: 12, 15, 18, 20, 22, 25, 28, 30, 32, 35. Number of observations n = 10 (even). Median = average of (n/2)th and (n/2 + 1)th observations = average of 5th and 6th observations = (22 + 25)/2 = 23.5.
Sample question3 marks
Q2. Explain why the mean calculated from grouped data using the direct method is an approximate value, not the exact mean.
Show model answer
Model answer
In grouped data, the exact values of observations are not known; we assume that all observations in a class are centered at its class mark (mid-point). This assumption leads to a loss of precision, making the calculated mean an approximation of the true mean.
Sample question3 marks
Q3. Calculate the mean of the following grouped data using the direct method: Class interval: 0-10, 10-20, 20-30; Frequency: 2, 5, 3.
Q4. For a grouped frequency distribution, the modal class is 40-55 with frequencies: preceding class = 3, modal class = 7, succeeding class = 6. If the class size is 15, find the mode.
Show model answer
Model answer
Using the formula: Mode = l + [(f1 - f0) / (2f1 - f0 - f2)] × h. Here l = 40, f1 = 7, f0 = 3, f2 = 6, h = 15. So Mode = 40 + [(7-3)/(14-3-6)] × 15 = 40 + (4/5)×15 = 40 + 12 = 52.
Sample question3 marks
Q5. What is an ogive? Explain the difference between 'less than' and 'more than' ogives.
Show model answer
Model answer
An ogive is a cumulative frequency curve. A 'less than' ogive plots cumulative frequencies against upper class limits, showing the number of observations less than or equal to each upper limit. A 'more than' ogive plots cumulative frequencies against lower class limits, showing the number of observations greater than or equal to each lower limit.
Want more questions with answers?Get the full practice set for this chapter.
What is the difference between grouped and ungrouped data?
Ungrouped data lists each observation separately with its frequency, such as the marks of 30 students written one by one. Grouped data arranges observations into class intervals, such as 10-25, 25-40, and gives the number of observations in each interval. Grouped data is used when the data is large and needs to be condensed for a meaningful study.
How do you find the mean of grouped data?
The mean of grouped data can be found by three methods. In the direct method, x̄ = Σfᵢxᵢ ÷ Σfᵢ, where xᵢ is the class mark. In the assumed mean method, x̄ = a + (Σfᵢdᵢ ÷ Σfᵢ) with dᵢ = xᵢ − a. In the step-deviation method, x̄ = a + h × (Σfᵢuᵢ ÷ Σfᵢ) with uᵢ = (xᵢ − a) ÷ h. All three give the same result.
What is the modal class and how is the mode of grouped data calculated?
The modal class is the class interval with the highest frequency. The mode is a value inside this class, given by Mode = l + [(f₁ − f₀) ÷ (2f₁ − f₀ − f₂)] × h, where l is the lower limit of the modal class, h is the class size, f₁ is the frequency of the modal class, and f₀ and f₂ are the frequencies of the preceding and succeeding classes.
How do you find the median of grouped data?
First compute n ÷ 2 and find the class whose cumulative frequency is greater than and nearest to n ÷ 2; this is the median class. Then use Median = l + [(n ÷ 2 − cf) ÷ f] × h, where l is the lower limit of the median class, cf is the cumulative frequency of the preceding class, f is the frequency of the median class and h is the class size.
What is cumulative frequency and what are the two types of cumulative frequency distribution?
The cumulative frequency of a class is the sum of the frequencies of all classes up to and including that class. In the less than type distribution, we record how many observations are less than each upper class limit. In the more than type distribution, we record how many observations are more than or equal to each lower class limit.
Which measure of central tendency should be used, mean, median or mode?
The mean is used when all observations matter and two or more distributions are to be compared, though extreme values affect it. The median is better when extreme values are present and a typical value is needed. The mode is best when the most frequent or most popular item has to be identified.