Which of the following is not an absolute measure of dispersion ?
- (a)Range
- (b)Mean Deviation
- (c)Quartile Deviation
- (d)Coefficient of Variation
Answer
Why
Correct — D, (d) Coefficient of Variation.
READ THE ASK FIRST. The word "not" is printed in BOLD ITALIC in the booklet, and that emphasis does not survive into plain text. Three of the four options ARE absolute measures of dispersion; the question wants the one that is not.
THE TEST THAT DECIDES IT is a test of UNITS. An ABSOLUTE measure of dispersion is expressed in the same units as the data themselves — if the observations are wages in rupees, the measure is in rupees; if they are heights in centimetres, it is in centimetres. A RELATIVE measure is a pure number, obtained by dividing an absolute measure by an appropriate average, so the units cancel and what remains is a ratio, usually written as a percentage. Ask of each option: with the data in rupees, is this quantity in rupees ?
RANGE = largest value − smallest value. Rupees. ABSOLUTE. MEAN DEVIATION = the average of the absolute deviations from the mean or the median. Rupees. ABSOLUTE. QUARTILE DEVIATION = (Q3 − Q1)/2, the semi-interquartile range. Rupees. ABSOLUTE. COEFFICIENT OF VARIATION = (standard deviation ÷ mean) × 100. Rupees divided by rupees. A PURE NUMBER, hence RELATIVE.
So the odd one out is the coefficient of variation, and the answer is option (d).
THE WORD "COEFFICIENT" IS THE SIGNPOST, and it is the most useful thing to take away from this item. Every absolute measure of dispersion has a relative partner formed by dividing it by an average, and the partner is named by putting "coefficient of" in front of it:
Range → Coefficient of Range = (L − S) / (L + S) Quartile Deviation → Coefficient of Quartile Deviation = (Q3 − Q1) / (Q3 + Q1) Mean Deviation → Coefficient of Mean Deviation = Mean Deviation ÷ the average it was taken about Standard Deviation → Coefficient of Variation = (σ ÷ mean) × 100
Read down that list and the structure of the option set becomes plain. The paper has offered three absolute measures from the left-hand column and one relative measure from the right-hand column, and the giveaway is in the name. The only irregularity is that the partner of the standard deviation is not called the "coefficient of standard deviation" but the coefficient of variation, which is precisely why it is the one chosen for the question — its name hides its family less well than the others reveal theirs, but "coefficient" is still there at the front of it.
WHY THE DISTINCTION EXISTS AT ALL. An absolute measure answers the question "how spread out are these numbers ?" and is meaningful only within one data set. A relative measure answers "how spread out are these numbers COMPARED WITH THEIR OWN SIZE ?", and that is what allows two different data sets to be compared. A standard deviation of ₹ 20 means something quite different among wages averaging ₹ 200 and among wages averaging ₹ 2,000; the coefficients of variation, 10 per cent and 1 per cent, say so at once. Comparing variability across data sets with different units — heights against weights, rainfall against temperature — is impossible with an absolute measure and routine with a relative one.
A NOTE ON THIS ITEM'S FORM. Only two questions on this whole paper ask a negative question, and this is one of them, which is why the setter has emphasised the word in bold italic rather than trusting it to be noticed.
Why the others are wrong
- (a)Range — The range is the simplest absolute measure of dispersion there is: the largest observation minus the smallest. If the data are wages in rupees the range is in rupees, which settles its classification at once. It is worth knowing its properties, since papers test them. It is very quick to compute and easy to explain, which is why it is used in quality control charts and in weather reporting; but it uses only two of the observations and ignores everything between them, so it is extremely sensitive to a single extreme value and gives no information about how the bulk of the data are distributed. Two data sets with the same range can be quite differently spread. Its relative partner is the coefficient of range, (L − S)/(L + S), which is the pure number formed from it — and the presence of that partner is itself a reminder that the range is the absolute member of the pair.
- (b)Mean Deviation — The mean deviation is the arithmetic mean of the ABSOLUTE deviations of the observations from a central value, usually the mean or the median. Because each deviation is measured in the units of the data and the averaging does not change the units, the mean deviation is in rupees when the data are in rupees — an absolute measure. It improves on the range by using every observation, and taking the deviations in absolute value is what prevents them from cancelling to zero, which is what they would otherwise do about the mean. A property worth carrying: the mean deviation is at its smallest when it is taken about the MEDIAN rather than about the mean. Its relative partner is the coefficient of mean deviation, formed by dividing it by the average about which it was computed, and once again it is the coefficient, not the deviation itself, that is the relative measure.
- (c)Quartile Deviation — The quartile deviation, also called the semi-interquartile range, is half the difference between the third and the first quartiles, (Q3 − Q1)/2. It is measured in the units of the data and is therefore absolute. Its distinctive merit is that it depends only on the middle half of the observations, so it is unaffected by extreme values at either end and can be computed for a distribution with open-ended classes — an income table whose top class is "₹ 1,00,000 and above", for instance, has no range and no standard deviation but does have a quartile deviation. Its weakness is the mirror of that merit: it ignores the tails altogether and so tells nothing about the extremes. Its relative partner is the coefficient of quartile deviation, (Q3 − Q1)/(Q3 + Q1), and the naming pattern holds here exactly as it does for the range and the mean deviation.
Concept
DISPERSION IS THE SECOND THING TO KNOW ABOUT ANY DATA SET, after central tendency. An average tells where the observations are centred; it says nothing about whether they cluster tightly around that centre or scatter widely from it. Two factories may pay the same average wage and be entirely different places to work. Measures of dispersion supply what the average leaves out.
THE TWO FAMILIES.
ABSOLUTE MEASURES are expressed in the units of the data. Range = L − S. Uses two observations only; sensitive to extremes; very quick. Quartile Deviation = (Q3 − Q1)/2. Uses the middle half; immune to extremes; usable with open-ended classes. Mean Deviation = the average of the absolute deviations about the mean or the median. Uses all observations; least when taken about the median; the absolute-value sign makes it awkward algebraically. Standard Deviation = the square root of the average of the squared deviations about the mean. Uses all observations; algebraically tractable; the foundation of most of statistical theory. Its square is the variance.
RELATIVE MEASURES are pure numbers, formed by dividing an absolute measure by an average. Coefficient of Range = (L − S)/(L + S) Coefficient of Quartile Deviation = (Q3 − Q1)/(Q3 + Q1) Coefficient of Mean Deviation = Mean Deviation ÷ the average used Coefficient of Variation = (standard deviation ÷ mean) × 100
WHAT THE COEFFICIENT OF VARIATION IS FOR. It is the standard tool for comparing the variability of two data sets that are not directly comparable — because they are measured in different units, or in the same units but at very different levels. The set with the LOWER coefficient of variation is the more consistent, uniform or stable of the two, and the one with the higher is the more variable. Examination questions almost always ask it in that comparative form: which of two batsmen is more consistent, which of two factories has the steadier wage structure, which of two crops the more reliable yield.
TWO PROPERTIES WORTH MEMORISING, because they separate the families cleanly. Add a constant to every observation: all the absolute measures are UNCHANGED, since every observation moves together and the gaps between them do not alter — but the coefficient of variation CHANGES, because the mean has moved while the standard deviation has not. Multiply every observation by a positive constant: every absolute measure is multiplied by that constant, while the coefficient of variation is UNCHANGED, since numerator and denominator scale together. Those two facts are the mathematical content of the absolute-relative distinction.
A CAUTION. The coefficient of variation is meaningful only where the mean is positive and the data are measured on a ratio scale. Where the mean is near zero it becomes unstable, and where the zero point is arbitrary — temperature in degrees Celsius, for instance — it has no interpretation at all.
Statistics on this paper is asked at the level of classification and definition, and this item is the clearest example of it. Nothing is computed; the candidate is asked whether he knows which family each of four named measures belongs to. That is a fair thing to ask of an officer who will read wage data, contribution data and compliance data, and who needs to know when a spread can be compared across two populations and when it cannot.
The construction is a straightforward odd-one-out with a negative ask, and the negation is where the difficulty is deliberately placed. Only two questions on this whole paper ask a negative question, so the habit of scanning for "not" is not one that this paper's rhythm builds up in a candidate — which is exactly why the setter has printed the word in bold italic. When a paper emphasises a word, it is not decoration; it marks the pivot of the item. Since bold and italic do not survive into plain text, a reader working from a transcription has to supply the care that the typography was supplying.
The item is also a good illustration of how far a naming convention can carry a candidate who has never memorised the classification. Three options are named for what they measure — a range, a deviation, a deviation — and one is named a coefficient. In this subject "coefficient of" almost always signals a ratio, and a ratio of two quantities in the same units is a pure number, hence relative. A candidate who has only that much can still answer correctly, which is worth knowing on a paper where time is short and partial knowledge has to be made to work.
The four options are short noun phrases in title case, printed as the booklet sets them.
Key facts
- An absolute measure of dispersion is expressed in the units of the data; a relative measure is a pure number formed by dividing an absolute measure by an average.
- The range, the mean deviation, the quartile deviation and the standard deviation are absolute measures; each has a relative partner named "coefficient of".
- The coefficient of variation is (standard deviation ÷ mean) × 100 — rupees divided by rupees — so it is a pure number and therefore relative.
- The range is L − S, uses only two observations and is highly sensitive to extreme values; its relative partner is the coefficient of range, (L − S)/(L + S).
- The quartile deviation is (Q3 − Q1)/2, depends only on the middle half of the data, is unaffected by extremes and can be computed for open-ended distributions.
- The mean deviation is the average of the absolute deviations about the mean or the median, and it is smallest when taken about the median.
- A relative measure is what makes two data sets comparable when they are in different units or at very different levels of magnitude.
- Adding a constant to every observation leaves every absolute measure unchanged but alters the coefficient of variation, because the mean moves while the spread does not.
- Multiplying every observation by a positive constant multiplies every absolute measure by it while leaving the coefficient of variation unchanged.
- The set with the LOWER coefficient of variation is the more consistent or uniform of two, which is the form in which examination questions usually put it.
Study next
Common traps
- Missing the negative ask. The word "not" is emphasised in the booklet but the emphasis disappears in plain text, so the ask must be read deliberately.
- Treating the quartile deviation as relative because it is built from quartiles. It is half the interquartile range and carries the units of the data.
- Confusing the mean deviation with its coefficient. The deviation is absolute; dividing it by an average makes it relative.
- Assuming the variance is relative because its units are squared. It is classed with the absolute measures, and its square root restores the data's own units.
- Comparing two data sets by their standard deviations alone. Without dividing by the means, the comparison says nothing about relative variability.
- Applying the coefficient of variation where the mean is near zero or the scale has an arbitrary zero, in which case the ratio has no useful meaning.
Statistics items on EPFO papers cluster around definitions, classifications and properties, and dispersion is one of the two or three topics that recur most reliably. The shapes to expect are the odd-one-out, as here; the property question — what happens to a statistic when every observation is doubled, or when a constant is added; the comparison question, in which two data sets are described by their means and standard deviations and the more consistent one is wanted; and the occasional small computation on five or six round numbers.
The negative form is used sparingly on this paper, so a candidate cannot rely on his eye being trained to it by repetition. What he can rely on is the setter marking the pivot, usually by emphasis. When a stem carries an emphasised word, read the whole stem again before looking at the options.
The preparation that pays here is a single table held in the head: the four absolute measures down the left, their four coefficients down the right, one line each for what the measure uses and what it ignores. That table answers classification questions outright, answers property questions once the two constant-effect rules are added to it, and turns comparison questions into a single division. It is a small amount of learning for a topic that appears in some form on every paper.
Related PYQs
EPFO_APFC_2016_Q75The mean and standard deviation of a set of 16 non-zero positive numbers in an observation are 26 and 3·5 respectively. The mean and standard deviation of another set of 24 non-zero positive numbers without changing the circumstances of both sets of observations, are 29 and 3, respectively. The mean and standard deviation of their combined set of observations will respectively be
- (a) 27·8 and 3·21
- (b) 26·2 and 3·32
- (c) 27·8 and 3·32
- (d) 26·2 and 3·21
Answer(a) 27·8 and 3·21
The mean-and-standard-deviation item on this same paper involving two sets of observations — dispersion computed rather than classified.
EPFO_APFC_2016_Q95Four quantities are such that their arithmetic mean (A.M.) is the same as the A.M. of the first three quantities. The fourth quantity is
- (a) Sum of the first three quantities
- (b) A.M. of the first three quantities
- (c) (Sum of the first three quantities)/4
- (d) (Sum of the first three quantities)/2
Answer(b) A.M. of the first three quantities
The arithmetic-mean property item on this paper, the central-tendency companion to this question on spread.
EPFO_APFC_2016_Q94For which time intervals, is the percentage rise of population the same for the following data ? Period | Population 1970 | 40,000 1980 | 50,000 1990 | 60,000 2000 | 72,000 2010 | 80,000
- (a) 1970 – 80 and 1980 – 90
- (b) 1980 – 90 and 1990 – 2000
- (c) 2000 – 2010 and 1990 – 2000
- (d) 1980 – 90 and 2000 – 2010
Answer(b) 1980 – 90 and 1990 – 2000
The population-table item on this paper, where the absolute-against-proportional distinction decides the answer in the setting of growth rather than of spread.
Practice
- practice — not a real PYQ
In factory A the wages have a mean of ₹ 200 and a standard deviation of ₹ 20, while in factory B the mean is ₹ 400 and the standard deviation ₹ 30. Which factory shows the greater relative variability of wages ?
- (a)Factory A
- (b)Factory B
- (c)The two are equally variable
- (d)It cannot be judged from these figures
Answer(a) Factory A — the coefficients of variation are (20/200) × 100 = 10 per cent for A and (30/400) × 100 = 7·5 per cent for B, so A is the more variable and B the more consistent, even though B has the larger standard deviation in rupees.
- practice — not a real PYQ
If the same constant is added to every observation in a data set, which one of the following measures changes as a result ?
- (a)The range
- (b)The quartile deviation
- (c)The standard deviation
- (d)The coefficient of variation
Answer(d) The coefficient of variation — adding a constant shifts every observation equally, so the gaps between them and hence all the absolute measures are untouched, but the mean rises while the standard deviation does not, and the ratio of the two therefore changes.