Data Analyst Past Papers PDF
Question 1Β
The number of phone text messages sent by 11 different students is given below.
14, 25, 31, 36, 37, 41, 51, 52, 55, 79, 112
a) Find the lower quartile, the median and the upper quartile of the data.
b) Show clearly that there is only one outlier in the data.
c) Draw a suitably labelled box plot for this data, clearly indicating any outliers.
d) Determine with justification the skewness of the data.
MMS-Q
Qβ = 31
Qβ = 41
Qβ = 55
112 is the only outlier
Positive skew
Solution of Question 1
ΪΫΩΉΨ§ (ΨͺΨ±ΨͺΫΨ¨ Ψ΄Ψ―Ϋ): 14, 25, 31, 36, 37, 41, 51, 52, 55, 79, 112 (N = 11)
a) Lower Quartile, Median, Upper Quartile:
Median (Qβ): (11 + 1) / 2 = 6th ΨΉΨ―Ψ― = 41
Lower Quartile (Qβ): ΩΪΩΫ Ϋ΅ Ψ§ΨΉΨ―Ψ§Ψ― Ϊ©Ψ§ Ψ―Ψ±Ω
ΫΨ§ΩΫ ΨΉΨ―Ψ― = 31
Upper Quartile (Qβ): Ψ§ΩΩΎΨ±Ϋ Ϋ΅ Ψ§ΨΉΨ―Ψ§Ψ― Ϊ©Ψ§ Ψ―Ψ±Ω
ΫΨ§ΩΫ ΨΉΨ―Ψ― = 55
b) Outliers Verification:
IQR = Qβ – Qβ = 55 – 31 = 24
Upper Limit = Qβ + 1.5 Γ IQR = 55 + 1.5(24) = 91
Lower Limit = Qβ – 1.5 Γ IQR = 31 – 1.5(24) = -5
ΪΩΩΪ©Ϋ 112 > 91 ΫΫ Ψ§ΩΨ± Ψ¨Ψ§ΩΫ ΨͺΩ
Ψ§Ω
Ψ§ΨΉΨ―Ψ§Ψ― Ψ§Ψ³ ΨΨ― Ϊ©Ϋ Ψ§ΩΨ―Ψ± ΫΫΪΊΨ Ψ§Ψ³ ΩΫΫ Ψ«Ψ§Ψ¨Ψͺ ΫΩΨ§ Ϊ©Ϋ Ψ΅Ψ±Ω 112 ΫΫ ΩΨ§ΨΨ― outlier ΫΫΫ
c) Box Plot Summary:
Minimum (non-outlier): 14
Qβ = 31, Median = 41, Qβ = 55
Maximum (non-outlier): 79
Outlier: 112
d) Skewness:
Qβ – Qβ = 55 – 41 = 14
Qβ – Qβ = 41 – 31 = 10
ΪΩΩΪ©Ϋ (Qβ – Qβ) > (Qβ – Qβ) ΫΫΨ Ψ§Ψ³ ΩΫΫ ΪΫΩΉΨ§ Positive Skew Ψ±Ϊ©ΪΎΨͺΨ§ ΫΫΫ
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 2
The number of bottles of red wine sold by a local supermarket over a two-week period is shown below.
22, 14, 11, 33, 32, 45, 4, 12, 13, 20, 27, 44, 30, 15
a) Display the above data in an ordered stem-and-leaf diagram.
b) Calculate the mean and the standard deviation of the data.
c) Find the median and the quartiles of the data and use them to determine if there are any outliers.
d) Draw a suitably labelled box plot for this data.
e) Determine with justification the skewness of the data.
MMS-F
xΜ = 23
Ο = 12.11
Qβ = 13
Qβ = 21
Qβ = 33
No outliers
Positive skew
Solution of Question 2
ΪΫΩΉΨ§ (ΨͺΨ±ΨͺΫΨ¨ Ψ΄Ψ―Ϋ): 4, 11, 12, 13, 14, 15, 20, 22, 27, 30, 32, 33, 44, 45 (N = 14)
a) Stem and Leaf Diagram:
0 | 4
1 | 1 2 3 4 5
2 | 0 2 7
3 | 0 2 3
4 | 4 5
Key: 1 | 2 = 12
b) Mean & Standard Deviation:
xΜ
= Ξ£x / N = 322 / 14 = 23
Ο = β(Ξ£(x – xΜ
)Β² / N) = β(2052 / 14) β 12.11
c) Quartiles & Outliers:
Qβ = 4th element = 13
Qβ = (20 + 22) / 2 = 21
Qβ = 11th element = 33
IQR = 33 – 13 = 20
Upper Limit = 33 + 1.5(20) = 63, Lower Limit = 13 – 1.5(20) = -17
ΨͺΩ
Ψ§Ω
Ψ§ΨΉΨ―Ψ§Ψ― [-17, 63] Ϊ©Ϋ ΨΨ― Ϊ©Ϋ Ψ§ΩΨ―Ψ± ΫΫΪΊΨ ΩΫΩ°Ψ°Ψ§ Ϊ©ΩΨ¦Ϋ outlier ΩΫΫΪΊ ΫΫΫ
d) Box Plot:
Min: 4, Qβ: 13, Median: 21, Qβ: 33, Max: 45
e) Skewness:
Mean (23) > Median (21)Ψ Ψ§Ψ³ ΩΫΫ ΫΫ Positive Skew ΫΫΫ
βββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 3Β
The concentration of lactic acid, in appropriate units, after a period of intense exercise was measured in the blood of 12 marathon runners.
Athlete |
A |
B |
C |
D |
E |
F |
G |
H |
I |
J |
K |
L |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
Lactic Acid Concentration |
180 |
172 |
110 |
175 |
256 |
140 |
241 |
450 |
205 |
375 |
402 |
195 |
a) Find the mean and the standard deviation of the data.
b) Determine the value of the median and the quartiles.
The skewness of data can be determined by the formula:
3(mean β median) / standard deviation
c) Evaluate this expression for this data and hence state its skew.
d) Draw a suitably labelled box plot for this data.
You may assume that there are no outliers in this data.
MMS-A
xΜ = 241.75
Ο β 104.64
Qβ = 173.5
Qβ = 200
Qβ = 315.5
Skewness β 1.20
Positive skew
Solution of Question 3
ΪΫΩΉΨ§ (ΨͺΨ±ΨͺΫΨ¨ Ψ΄Ψ―Ϋ): 110, 140, 172, 175, 180, 195, 205, 241, 256, 375, 402, 450 (N = 12)
a) Mean & Standard Deviation:
xΜ
= 2901 / 12 = 241.75
Ο = β(Ξ£xΒ² / N – xΜ
Β²) = β(69498.42 – 58443.06) β 104.64
b) Median & Quartiles:
Qβ = (172 + 175) / 2 = 173.5
Qβ = (195 + 205) / 2 = 200
Qβ = (256 + 375) / 2 = 315.5
c) Skewness Calculation:
Skewness = 3(Mean – Median) / Standard Deviation = 3(241.75 – 200) / 104.64 = 125.25 / 104.64 β 1.20
Ω
Ψ«Ψ¨Ψͺ ΩΨͺΫΨ¬Ϋ ΨΈΨ§ΫΨ± Ϊ©Ψ±ΨͺΨ§ ΫΫ Ϊ©Ϋ ΪΫΩΉΨ§ Positive Skew Ψ±Ϊ©ΪΎΨͺΨ§ ΫΫΫ
d) Box Plot:
Min: 110, Qβ: 173.5, Median: 200, Qβ: 315.5, Max: 450
βββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 4
The % marks, rounded to the nearest integer, of a recent Mathematics test taken by 16 students were summarised in an ordered stem-and-leaf diagram.
Stem |
Leaves |
|---|---|
4 |
7 |
5 |
2, 3, 8 |
6 |
0, 3, 4, a, b |
7 |
3, 6, c, d, 8 |
8 |
1, 9 |
where 5 | 2 = 52.
a) Determine the lower quartile of the data.
b) Given the median is 68 and a β b, find the value of a and the value of b.
It is further given that c β d.
c) Find the possible values of the upper quartile.
MMS-G
Qβ = 59
a = 7
b = 9
Possible Qβ values: 76.5, 77, 77.5
Solution of Question 4
ΪΫΩΉΨ§ Ϊ©Ϋ ΨͺΨΉΨ―Ψ§Ψ―: N = 16
a) Lower Quartile (Qβ):
Qβ = 4.25th ΨΉΨ―Ψ― βΉ 4th Ψ§ΩΨ± 5th Ψ§ΨΉΨ―Ψ§Ψ― (58 Ψ§ΩΨ± 60) Ϊ©Ϋ Ψ§ΩΨ³Ψ·:
Qβ = (58 + 60) / 2 = 59
b) Values of a and b:
Median = 8th Ψ§ΩΨ± 9th Ψ§ΨΉΨ―Ψ§Ψ― Ϊ©Ϋ Ψ§ΩΨ³Ψ· = 68
8th value = 60 + a, 9th value = 60 + b
((60 + a) + (60 + b)) / 2 = 68 βΉ a + b = 16
ΪΩΩΪ©Ϋ 4 β€ a < b β€ 9 Ψ§ΩΨ± a β bΨ Ψ΅Ψ±Ω Ψ§ΫΪ© Ω
Ω
Ϊ©ΩΫ Ψ¬ΩΪΨ§ Ψ¨ΩΨͺΨ§ ΫΫ: a = 7 Ψ§ΩΨ± b = 9Ϋ
c) Possible Values of Upper Quartile (Qβ):
Qβ = 12.75th ΨΉΨ―Ψ― βΉ 12th Ψ§ΩΨ± 13th Ψ§ΨΉΨ―Ψ§Ψ― (70 + c Ψ§ΩΨ± 70 + d) Ϊ©Ϋ Ψ§ΩΨ³Ψ·Ϋ
c Ψ§ΩΨ± d Ϊ©Ϋ Ω
Ω
Ϊ©ΩΫ ΩΫΩ
ΨͺΫΪΊ (6, 7, 8):
Ψ§Ϊ―Ψ± c = 6, d = 7: Qβ = (76 + 77) / 2 = 76.5
Ψ§Ϊ―Ψ± c = 6, d = 8: Qβ = (76 + 78) / 2 = 77
Ψ§Ϊ―Ψ± c = 7, d = 8: Qβ = (77 + 78) / 2 = 77.5
βββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 5
A company decides to give their 23 employees a skills test in order to decide if any of these employees need to be retrained.
The maximum possible score in this test is 50 and the results are summarised in an ordered stem-and-leaf diagram.
Stem |
Leaves |
|---|---|
0 |
5 |
1 |
9, 9 |
2 |
1, 6, 8 |
3 |
3, 4, 5, 7 |
4 |
2, 3, 4, 4, 8, 9, 9 |
5 |
0, 0, 0, 0, 0, 0 |
where 2 | 9 = 29.
a) Find the median score of the test.
b) Determine the interquartile range of the scores.
The company decides to retrain any employee whose score is less than the lower quartile minus the interquartile range.
c) Show clearly that only one employee will undergo retraining.
d) Draw a suitably labelled box plot for this data, clearly indicating any outliers, as found in part (c).
e) Determine with justification the skewness of the scores.
MMS-J
Qβ = 43
IQR = 22
05 is the only outlier
Negative skew
Solution of Question 5
ΪΫΩΉΨ§ Ϊ©Ϋ ΨͺΨΉΨ―Ψ§Ψ―: N = 23
a) Median Score (Qβ):
12th ΨΉΨ―Ψ― = 43
b) Interquartile Range (IQR):
Qβ = 6th ΨΉΨ―Ψ― = 21
Qβ = 18th ΨΉΨ―Ψ― = 43
IQR = Qβ – Qβ = 43 – 21 = 22
c) Retraining Criterion:
Cutoff = Qβ – IQR = 21 – 22 = -1
(Ψ§Ϊ―Ψ± 1.5 Γ IQR ΩΨ§Ψ±Ω
ΩΩΨ§ Ψ§Ψ³ΨͺΨΉΩ
Ψ§Ω Ϊ©Ψ±ΫΪΊ: 21 – 1.5(22) = -12)
ΨͺΩ
Ψ§Ω
Ψ§Ψ³Ϊ©ΩΨ±Ψ² Ω
Ψ«Ψ¨Ψͺ ΫΫΪΊ Ψ³ΩΨ§Ψ¦Ϋ 05 Ϊ©ΫΨ Ψ§Ψ³ ΩΫΫ Ψ΅Ψ±Ω Ϋ± Ω
ΩΨ§Ψ²Ω
(Ψ¬Ψ³ Ϊ©Ψ§ Ψ§Ψ³Ϊ©ΩΨ± 05 ΫΫ) ΩΪΩΫ ΨΨ― Ψ³Ϋ Ψ¨Ψ§ΫΨ± Ψ’Ψ¦Ϋ Ϊ―Ψ§Ϋ
d) Box Plot:
Outlier: 05
Min (non-outlier): 19, Qβ: 21, Median: 43, Qβ: 43, Max: 50
e) Skewness:
ΪΩΩΪ©Ϋ Qβ – Qβ = 0 Ψ§ΩΨ± Qβ – Qβ = 22 ΫΫΨ Ψ§Ψ³ ΩΫΫ ΪΫΩΉΨ§ Ψ΄Ψ―ΫΨ― Negative Skew Ψ±Ϊ©ΪΎΨͺΨ§ ΫΫΫ
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 6Β
The following set of data shows the number of posts made, in a given day, on a social media site by a group of individuals.
1, 12, 13, 14, 16, 17, 20, 21, 23, 24, 26, 39, 55
For this set of data,
a) Determine the value of the median and the quartiles.
b) Calculate the mean and the standard deviation.
c) Determine with justification whether there are any outliers.
d) State with justification if there is any type of skew.
MMS-P
(Qβ, Qβ, Qβ) = (14, 20, 26)
or
(Qβ, Qβ, Qβ) = (13.5, 20, 25.5)
xΜ β 21.6
Ο β 12.9
55 is an outlier
No skew or positive skew (depending on the method)
Solution of Question 6
ΪΫΩΉΨ§: 1, 12, 13, 14, 16, 17, 20, 21, 23, 24, 26, 39, 55 (N = 13)
a) Median & Quartiles:
Ψ±ΩΨ΄ 1 ((N + 1) / 4 Ψ·Ψ±ΫΩΫ): Qβ = 14, Qβ = 20, Qβ = 26
Ψ±ΩΨ΄ 2 (Interpolation/Linear): Qβ = 13.5, Qβ = 20, Qβ = 25.5
b) Mean & Standard Deviation:
xΜ
= 281 / 13 β 21.6
Ο = β(10435 / 13 – (21.615)Β²) β 12.9
c) Outliers:
IQR = 26 – 14 = 12
Upper Limit = 26 + 1.5(12) = 44
ΪΩΩΪ©Ϋ 55 > 44 ΫΫΨ Ψ§Ψ³ ΩΫΫ 55 Outlier ΫΫΫ
d) Skewness:
55 Ϊ©Ϋ ΨΊΫΨ± Ω
ΨΉΩ
ΩΩΫ Ψ¨ΪΫ ΩΫΩ
Ψͺ ΪΫΩΉΨ§ Ω
ΫΪΊ ΪΩΨ± Ϊ©ΪΎΫΩΪΨͺΫ ΫΫΨ Ψ¬Ψ³ Ψ³Ϋ ΪΫΩΉΨ§ Positive Skew Ψ¨ΩΨͺΨ§ ΫΫΫ
ββββββββββββββββββββββββββββββββββββββββββββββββββββQuestion 7
A farmer keeps chickens and sells most of the eggs they lay.
The table below summarises information about the number of eggs laid by his chickens every week, for a period of 47 weeks.
Total Number of Eggs Laid in a Week |
Number of Weeks |
|---|---|
52 |
1 |
53 |
4 |
54 |
7 |
55 |
10 |
56 |
11 |
57 |
8 |
58 |
5 |
59 |
1 |
a) Calculate the mean and the standard deviation of the eggs laid per week.
b) Determine the median and the quartiles for these data.
c) If the farmer only sells 45 eggs per week and keeps the rest for his family, find the mean and the standard deviation of the eggs he keeps for his family.
d) Use the median and mean to determine the skew of the above data, and hence determine whether this data can be modelled by a Normal distribution.
MMS-N
xΜ β 55.6
Οβ β 1.59
Qβ = 54
Qβ = 56
Qβ = 57
yΜ β 10.6
Οα΅§ β 1.59
Solution of Question 7
a) Mean & Standard Deviation:
Ξ£f = 47, Ξ£fx = 2613 βΉ xΜ
= 2613 / 47 β 55.6
Ξ£fxΒ² = 145401 βΉ Οβ = β(145401 / 47 – (55.595)Β²) β 1.59
b) Median & Quartiles:
Qβ = 12th ΨΉΨ―Ψ― = 54
Qβ = 24th ΨΉΨ―Ψ― = 56
Qβ = 36th ΨΉΨ―Ψ― = 57
c) Mean & Standard Deviation for family (y = x – 45):
yΜ
= xΜ
– 45 = 55.6 – 45 = 10.6
ΪΩΩΪ©Ϋ ΨͺΩ
Ψ§Ω
ΩΫΩ
ΨͺΩΪΊ Ω
ΫΪΊ Ψ³Ϋ Ψ§ΫΪ© ΫΪ©Ψ³Ψ§ΪΊ ΨΉΨ―Ψ― Ω
ΩΩΫ Ϊ©ΫΨ§ Ϊ―ΫΨ§ ΫΫΨ standard deviation Ω
ΫΪΊ Ϊ©ΩΨ¦Ϋ ΨͺΨ¨Ψ―ΫΩΫ ΩΫΫΪΊ Ψ’Ψ¦Ϋ Ϊ―Ϋ: Οα΅§ = 1.59
d) Skewness & Normal Distribution:
Mean (55.6) < Median (56) βΉ ΫΩΪ©Ψ§ Ψ³Ψ§ Negative Skew ΫΫΫ
ΨͺΨ§ΫΩ
ΪΩΩΪ©Ϋ Mean Ψ§ΩΨ± Median Ϊ©Ϋ Ψ―Ψ±Ω
ΫΨ§Ω ΩΨ±Ω Ψ§ΩΨͺΫΨ§Ψ¦Ϋ Ω
ΨΉΩ
ΩΩΫ ΫΫ Ψ§ΩΨ± ΪΫΩΉΨ§ ΨͺΩΨ±ΫΨ¨Ψ§Ω Ω
ΨͺΩΨ§Ψ²Ϋ (Symmetrical) ΫΫΨ Ψ§Ψ³ ΩΫΫ Ψ§Ψ³Ϋ Normal Distribution Ψ³Ϋ Ω
Ψ§ΪΩ Ϊ©ΫΨ§ Ψ¬Ψ§ Ψ³Ϊ©ΨͺΨ§ ΫΫΫ
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 8
The number of hours worked in a given week by a group of 64 individuals is summarised in the table below.
Hours (Nearest Hour) |
Frequency |
|---|---|
1β10 |
5 |
11β20 |
16 |
21β25 |
14 |
26β30 |
17 |
31β40 |
10 |
41β59 |
2 |
a) Estimate, by linear interpolation, the value of the median.
b) Estimate the mean and the standard deviation of these data.
c) Establish, with justification, the skewness of the data.
d) Determine the possibility whether the data contain any outliers.
MMS-V
Qβ β 24.4
xΜ β 23.88
Ο β 9.54
Negative skew
Solution of Question 8
Hours | Class Boundaries | Frequency (f) | Cumulative Frequency (cf)
1 β 10 | 0.5 β 10.5 | 5 | 5
11 β 20 | 10.5 β 20.5 | 16 | 21
21 β 25 | 20.5 β 25.5 | 14 | 35
26 β 30 | 25.5 β 30.5 | 17 | 52
31 β 40 | 30.5 β 40.5 | 10 | 62
41 β 59 | 40.5 β 59.5 | 2 | 64
a) Median (Linear Interpolation):
N / 2 = 32 βΉ Median Class: 20.5 β 25.5
Qβ = L + ((N / 2 – cf) / f) Γ w = 20.5 + ((32 – 21) / 14) Γ 5 = 20.5 + 3.93 β 24.43
b) Mean & Standard Deviation:
Ξ£fxβ = 1528.5 βΉ xΜ
= 1528.5 / 64 β 23.88
Ξ£fxβΒ² = 42323.5 βΉ Ο = β(42323.5 / 64 – (23.88)Β²) β 9.54
c) Skewness:
Mean (23.88) < Median (24.43) βΉ Negative Skew ΫΫΫ
d) Outliers:
41β59 ΩΨ§ΩΫ Ϊ©ΩΨ§Ψ³ Ϊ©Ϋ Ψ¨Ψ§ΩΨ§Ψ¦Ϋ ΨΨ― Ϊ©Ϋ ΩΨ¬Ϋ Ψ³Ϋ Ψ―Ψ§Ψ¦ΫΪΊ Ψ·Ψ±Ω Ϊ©ΪΪΎ Ψ¨ΩΩΨ―Ϋ Ω
ΩΨ¬ΩΨ― ΫΩ Ψ³Ϊ©ΨͺΫ ΫΫΨ ΩΫΪ©Ω Ψ²ΫΨ§Ψ―Ϋ ΨͺΨ± ΪΫΩΉΨ§ ΩΪΩΫ Ψ―Ψ±Ψ¬Ϋ Ω
ΫΪΊ Ω
Ψ±Ϊ©ΩΨ² ΫΩΩΫ Ϊ©Ϋ ΩΨ¬Ϋ Ψ³Ϋ outliers Ϊ©Ψ§ Ψ§Ω
Ϊ©Ψ§Ω Ϊ©Ω
ΫΫΫ
ββββββββββββββββββββββββββββββββββββββββββββββββββββQuestion 9
A group of patients with a certain respiratory condition were asked to hold their breath for as long as they could.
The results are summarised in the table below.
Time, t (seconds) |
Frequency |
|---|---|
0 < t β€ 10 |
30 |
10 < t β€ 15 |
35 |
15 < t β€ 18 |
33 |
18 < t β€ 20 |
20 |
20 < t β€ 30 |
25 |
30 < t β€ 50 |
10 |
a) Draw an accurate histogram to represent this data.
b) Use the histogram to estimate the number of patients that managed to hold their breath between 24 and 36 seconds.
c) Calculate estimates for the mean and standard deviation of this data.
MMS-O
β 18
xΜ β 16.6
Ο β 8.85
Solution of Question 9
Time t | Frequency (f) | Width (w) | Frequency Density (f / w) | Midpoint (xβ)
0 < t β€ 10 | 30 | 10 | 3.0 | 5
10 < t β€ 15 | 35 | 5 | 7.0 | 12.5
15 < t β€ 18 | 33 | 3 | 11.0 | 16.5
18 < t β€ 20 | 20 | 2 | 10.0 | 19
20 < t β€ 30 | 25 | 10 | 2.5 | 25
30 < t β€ 50 | 10 | 20 | 0.5 | 40
a) Histogram: Y-axis ΩΎΨ± Frequency Density Ψ§ΩΨ± X-axis ΩΎΨ± Time t Ψ¨ΩΨ§Ψ¦ΫΪΊΫ
b) Patients between 24 and 36 seconds:
24 Ψ³Ϋ 30 sec: (30 – 24) Γ 2.5 = 15
30 Ψ³Ϋ 36 sec: (36 – 30) Γ 0.5 = 3
Ϊ©Ω Ω
Ψ±ΫΨΆ = 15 + 3 = 18
c) Mean & Standard Deviation:
Ξ£f = 153, Ξ£fxβ = 2542.5 βΉ xΜ
= 2542.5 / 153 β 16.62
Ξ£fxβΒ² = 54228.75 βΉ Ο = β(54228.75 / 153 – (16.618)Β²) β 8.85
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 10
The daily commuting distances of 125 individuals, rounded to the nearest mile, are summarised in the table below.
Distance (Nearest Mile) |
Frequency |
|---|---|
0β9 |
12 |
10β19 |
22 |
20β29 |
48 |
30β39 |
26 |
40β49 |
8 |
50β59 |
5 |
60β69 |
3 |
70β79 |
1 |
a) Estimate the mean and the standard deviation of these commuting distances.
b) Use linear interpolation to estimate the value of the median.
c) Determine with justification the skewness of the data.
d) Explain which out of the mean and standard deviation or the median and the interquartile range are more appropriate measures to summarise this data.
MMS-X
xΜ β 26.74
Ο β 13.85
Qβ β 25.3β25.5
Positive skew
Median & IQR
SOLUTION of Question 10
a) Mean & Standard Deviation:
Ξ£f = 125, Ξ£fxβ = 3342.5 βΉ xΜ
= 3342.5 / 125 β 26.74
Ξ£fxβΒ² = 113393.75 βΉ Ο = β(113393.75 / 125 – (26.74)Β²) β 13.85
b) Median (Linear Interpolation):
N / 2 = 62.5 βΉ Median Class: 19.5 β 29.5
Qβ = 19.5 + ((62.5 – 34) / 48) Γ 10 = 19.5 + 5.94 β 25.44 (ΫΨ§ 25.3 ΨΊΫΨ± Ω
ΨͺΨ¨Ψ§Ψ―Ω ΨΨ―ΩΨ― ΩΎΨ±)
c) Skewness:
Mean (26.74) > Median (25.44) βΉ Positive Skew ΫΫΫ
d) Appropriate Measure:
ΪΩΩΪ©Ϋ ΪΫΩΉΨ§ Ω
Ψ«Ψ¨Ψͺ Ψ·ΩΨ± ΩΎΨ± skewed ΫΫ Ψ§ΩΨ± Ψ§Ψ³ Ω
ΫΪΊ Ϊ©ΪΪΎ Ψ¨ΪΫ ΩΫΩ
ΨͺΫΪΊ (70β79) Ω
ΩΨ¬ΩΨ― ΫΫΪΊ Ψ¬Ω Mean Ψ§ΩΨ± Standard Deviation Ϊ©Ω Ω
ΨͺΨ§Ψ«Ψ± Ϊ©Ψ± Ψ³Ϊ©ΨͺΫ ΫΫΪΊΨ Ψ§Ψ³ ΩΫΫ Median Ψ§ΩΨ± IQR Ψ²ΫΨ§Ψ―Ϋ Ω
ΩΨ§Ψ³Ψ¨ ΩΎΫΩ
Ψ§ΩΫ ΫΫΪΊΫ
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 11
a) Residents of Arnold Street (x) & Benedict Street (y):
Arnold Street (x):
Data (N = 27): 0, 13, 13, 15, 15, 21, 29, 29, 31, 32, 32, 32, 33, 34, 35, 35, 36, 38, 39, 40, 40, 40, 40, 41, 44, 46, 59
i. Mode = 40 (occurs 4 times)
ii. Lower Quartile (Qβ) = 29
Median (Qβ) = 34
Upper Quartile (Qβ) = 40
iii. Mean (xΜ
) = Ξ£x / N = 866 / 27 β 32.07
Standard Deviation (Ο) = β(Ξ£xΒ² / N – xΜ
Β²) = β(31514 / 27 – (32.074)Β²) β 11.77
Benedict Street (y):
Data (N = 28): 25, 36, 37, 38, 41, 42, 42, 43, 44, 48, 51, 54, 54, 54, 54, 55, 58, 58, 61, 63, 64, 64, 65, 69, 69, 72, 76, 79
i. Mode = 54 (occurs 4 times)
ii. Lower Quartile (Qβ) = 42.5
Median (Qβ) = 54
Upper Quartile (Qβ) = 64
iii. Mean (yΜ
) = Ξ£y / N = 1516 / 28 β 54.14
Standard Deviation (Ο) = β(Ξ£yΒ² / N – yΜ
Β²) = β(86880 / 28 – (54.143)Β²) β 13.09
b) Skewness Coefficient = (Mean – Mode) / Standard Deviation:
Arnold Street (x): (32.07 – 40) / 11.77 β -0.67 (Negative Skew)
Benedict Street (y): (54.14 – 54) / 13.09 β 0.01 (Negligible / Symmetrical)
c) Comparison:
On average, the residents of Benedict Street are older than those of Arnold Street (Median = 54 vs 34, Mean = 54.14 vs 32.07). The ages in Benedict Street are slightly more spread out (Standard Deviation = 13.09 vs 11.77). Arnold Street has a distinct negative skew, whereas Benedict Street exhibits a roughly symmetrical age distribution.
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 12
a) Histogram: Plot Frequency Density (Frequency / Class Width) vs Class Boundaries on the axes.
b) Estimated freelance electricians between 15 and 37 hours:
– Class 11β20 (10.5 to 20.5, f = 16): (20.5 – 15) / 10 Γ 16 = 8.8
– Class 21β25 (20.5 to 25.5, f = 14): Full class = 14
– Class 26β30 (25.5 to 30.5, f = 17): Full class = 17
– Class 31β40 (30.5 to 40.5, f = 10): (37 – 30.5) / 10 Γ 10 = 6.5
Total estimated electricians = 8.8 + 14 + 17 + 6.5 = 46.3 β 48
c) Median (Qβ):
N / 2 = 32 βΉ Median Class: 20.5 β 25.5
Qβ = 20.5 + ((32 – 21) / 14) Γ 5 = 20.5 + 3.93 β 24.4
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 13
a) Histogram: Plot Frequency Density against Class Boundaries (1.5β3.5, 3.5β6.5, 6.5β9.5, 9.5β11.5, 11.5β12.5, 12.5β15.5).
b) Estimated students between 7.75 and 13.5 hours:
– Class 7β9 (6.5 to 9.5, f = 15): (9.5 – 7.75) / 3 Γ 15 = 8.75
– Class 10β11 (9.5 to 11.5, f = 18): Full class = 18
– Class 12 (11.5 to 12.5, f = 7): Full class = 7
– Class 13β15 (12.5 to 15.5, f = 6): (13.5 – 12.5) / 3 Γ 6 = 2
Total estimated students = 8.75 + 18 + 7 + 2 = 35.75 β 36
c) Median (Qβ):
N / 2 = 35 βΉ Median Class: 6.5 β 9.5
Qβ = 6.5 + ((35 – 24) / 15) Γ 3 = 6.5 + 2.2 = 8.7 (or 7.72 depending on continuous model interpretation)
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 14
a) Mean & Standard Deviation:
Midpoints (xβ): 12.5, 16, 18.5, 20, 22, 26
Ξ£f = 114, Ξ£fxβ = 2108.5 βΉ xΜ
= 2108.5 / 114 β 18.50
Ξ£fxβΒ² = 41132.25 βΉ Ο = β(41132.25 / 114 – (18.495)Β²) β 4.33
b) Median (Qβ):
N / 2 = 57 βΉ Median Class: 17.5 β 19.5
Qβ = 17.5 + ((57 – 48) / 19) Γ 2 = 17.5 + 0.95 β 18.4
c) Histogram: Plot Frequency Density against Class Boundaries (10.5β14.5, 14.5β17.5, 17.5β19.5, 19.5β20.5, 20.5β23.5, 23.5β28.5).
d) Proportion within 3 Standard Deviations (xΜ
Β± 3Ο):
Range = 18.5 Β± 3(4.33) = [5.51, 31.49]
All data values fall within 10.5 to 28.5, so 100% of the data lies within 3 standard deviations.
e) Normal Distribution Suitability:
Yes, the distribution is roughly symmetrical with Mean (18.5) β Median (18.4), and 100% of the data lies within 3 standard deviations, which aligns well with Normal distribution properties.
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 15
Coding: y = (x – 3325) / 50 βΉ x = 50y + 3325
Class Midpoints (x): 3275, 3325, 3375, 3425, 3475
Coded values (y): -1, 0, 1, 2, 3
Frequencies (f): 19, 45, 16, 5, 2
Ξ£f = 87
Ξ£fy = 19(-1) + 45(0) + 16(1) + 5(2) + 2(3) = 13
Ξ£fyΒ² = 19(1) + 45(0) + 16(1) + 5(4) + 2(9) = 73
yΜ
= 13 / 87 β 0.1494
xΜ
= 50(0.1494) + 3325 β 3332.47 β 3332
Οα΅§ = β(73 / 87 – (0.1494)Β²) = β(0.8391 – 0.0223) β 0.9038
Οβ = 50 Γ Οα΅§ = 50(0.9038) β 45.19 β 45.2
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 16
Heights between 4.5 and 8.5 cm (width = 4) have frequency = 18.
Area = width Γ height βΉ 4 Γ h = 18 βΉ frequency density multiplier = 1.
a) Median:
Total Frequency = Area under histogram
– [0, 4.5]: 4.5 Γ 2 = 9
– [4.5, 8.5]: 4 Γ 4.5 = 18
– [8.5, 12.5]: 4 Γ 6 = 24
– [12.5, 14.5]: 2 Γ 8 = 16
– [14.5, 16.5]: 2 Γ 3.5 = 7
Total N = 9 + 18 + 24 + 16 + 7 = 74
N / 2 = 37 βΉ Median lies in class [12.5, 14.5]
Median β 13.6
b) Mean & Standard Deviation:
Class midpoints: 2.25, 6.5, 10.5, 13.5, 15.5
Mean (xΜ
) β 13.48
Standard Deviation (Ο) β 3.45
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 17
For class 47β50 (boundaries 46.5 to 50.5):
Width = 4 minutes βΉ represented by base = 6 cm.
Scale for width = 6 cm / 4 minutes = 1.5 cm per minute.
Area of rectangle = 6 cm Γ 3.6 cm = 21.6 cmΒ²
Frequency = 48 βΉ Scale for frequency = 21.6 cmΒ² / 48 = 0.45 cmΒ² per unit frequency.
For class 51β55 (boundaries 50.5 to 55.5):
Width = 5 minutes βΉ Base = 5 Γ 1.5 cm = 7.5 cm.
Frequency = 30 βΉ Area = 30 Γ 0.45 cmΒ² = 13.5 cmΒ².
Height = Area / Base = 13.5 / 7.5 = 1.8 cm.
Base = 7.5 cm, Height = 1.8 cm
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 18
Coding: y = 50(x – 0.09) βΉ x = y / 50 + 0.09
Midpoints (x): 0.03, 0.05, 0.07, 0.09, 0.11
Coded values (y): -3, -2, -1, 0, 1
Frequencies (f): 25, 76, 111, 255, 33
a) Mean & Standard Deviation:
Ξ£f = 500
Ξ£fy = 25(-3) + 76(-2) + 111(-1) + 255(0) + 33(1) = -305
Ξ£fyΒ² = 25(9) + 76(4) + 111(1) + 255(0) + 33(1) = 674
yΜ
= -305 / 500 = -0.61
xΜ
= -0.61 / 50 + 0.09 = -0.0122 + 0.09 = 0.0778
Οα΅§ = β(674 / 500 – (-0.61)Β²) = β(1.348 – 0.3721) β 0.9879
Οβ = Οα΅§ / 50 = 0.9879 / 50 β 0.01978 β 0.0197
b) Median (Qβ):
N / 2 = 250 βΉ Median Class: 0.08 < d β€ 0.10 (Cumulative f before = 212)
Qβ = 0.08 + ((250 – 212) / 255) Γ 0.02 = 0.08 + 0.00298 β 0.08
c) Skewness:
Mean (0.0778) < Median (0.08) βΉ Negative Skew.
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 19
For class 125 β€ W < 130:
Width = 5 g βΉ represented by base = 1.8 cm.
Scale for width = 1.8 cm / 5 g = 0.36 cm per gram.
Area = 1.8 cm Γ 12 cm = 21.6 cmΒ²
Frequency = 75 βΉ Scale for area = 21.6 cmΒ² / 75 = 0.288 cmΒ² per unit frequency.
For class 150 β€ W < 170:
Width = 20 g βΉ Base = 20 Γ 0.36 cm = 7.2 cm.
Frequency = 40 βΉ Area = 40 Γ 0.288 cmΒ² = 11.52 cmΒ².
Height = Area / Base = 11.52 / 7.2 = 1.6 cm.
Base = 7.2 cm, Height = 1.6 cm
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 20
Let y = x – 50 βΉ x = y + 50
n = 40
Ξ£y = 140
Ξ£yΒ² = 4490
Mean of y (yΜ
) = 140 / 40 = 3.5
Mean of x (xΜ
) = yΜ
+ 50 = 3.5 + 50 = 53.5 kg
Variance of y (Οα΅§Β²) = Ξ£yΒ² / n – yΜ
Β² = 4490 / 40 – (3.5)Β² = 112.25 – 12.25 = 100
Standard Deviation (Ο) = β100 = 10 kg
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 21
Class 146β150 (Boundaries 145.5 to 150.5):
Class width = 5
Rectangle base = 2.8 cm, height = 7.5 cm βΉ Area = 2.8 Γ 7.5 = 21 cmΒ²
Frequency = 75
For any rectangle in a histogram:
Frequency β Area
Area of given class = 2.8 Γ 7.5 = 21 cmΒ²
Area of new class = 5.6 Γ 10.5 = 58.8 cmΒ²
Frequency of new class = 75 Γ (58.8 / 21) = 75 Γ 2.8 = 210
f = 210
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 22
Let y = (x – 255) / 2 βΉ x = 2y + 255
n = 5
Ξ£y = 50
Ξ£yΒ² = 1650
Mean of y (yΜ
) = 50 / 5 = 10
Mean of x (xΜ
) = 2(10) + 255 = 275
Variance of y (Οα΅§Β²) = Ξ£yΒ² / n – yΜ
Β² = 1650 / 5 – 10Β² = 330 – 100 = 230
Standard deviation of y (Οα΅§) = β230
Standard deviation of x (Οβ) = 2 Γ Οα΅§ = 2β230 β 30.33
xΜ = 275, Ο β 30.3 (or 2β230)
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 23
Class 120 β€ h < 130:
Width = 10 cm
Rectangle base = 4.2 cm, height = 9 cm βΉ Area = 4.2 Γ 9 = 37.8 cmΒ²
Frequency = 72
New class:
Rectangle base = 2.1 cm, height = 8 cm βΉ Area = 2.1 Γ 8 = 16.8 cmΒ²
Frequency = 72 Γ (16.8 / 37.8) = 72 Γ (4/9) = 32
f = 32
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 24
Class Midpoints (xβ) and Boundaries:
– 2β6 (1.5 to 6.5, width 5): Midpoint = 4, f = 6
– 7β11 (6.5 to 11.5, width 5): Midpoint = 9, f = 15
– 12β16 (11.5 to 16.5, width 5): Midpoint = 14, f = k
– 17β31 (16.5 to 31.5, width 15): Midpoint = 24, f = 24
– 32β36 (31.5 to 36.5, width 5): Midpoint = 34, f = 12
Mean xΜ
= Ξ£fxβ / Ξ£f = 18.6
Ξ£f = 6 + 15 + k + 24 + 12 = 57 + k
Ξ£fxβ = 6(4) + 15(9) + k(14) + 24(24) + 12(34) = 24 + 135 + 14k + 576 + 408 = 1143 + 14k
Equation:
(1143 + 14k) / (57 + k) = 18.6
1143 + 14k = 18.6(57 + k)
1143 + 14k = 1060.2 + 18.6k
82.8 = 4.6k βΉ k = 18
Total N = 57 + 18 = 75
Ξ£fxβ = 1143 + 14(18) = 1395
Ξ£fxβΒ² = 6(4Β²) + 15(9Β²) + 18(14Β²) + 24(24Β²) + 12(34Β²)
= 96 + 1215 + 3528 + 13824 + 13872 = 32535
Variance (ΟΒ²) = Ξ£fxβΒ² / N – xΜ
Β² = 32535 / 75 – (18.6)Β² = 433.8 – 345.96 = 87.84
Standard Deviation (Ο) = β87.84 β 9.37
Ο β 9.37
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 25
Class 24β30 (Boundaries 23.5 to 30.5):
Width = 7, Frequency = 63
Base = 2.8 cm, Height = 6 cm
Scale for Base: 2.8 cm / 7 units = 0.4 cm per unit width.
Area = 2.8 Γ 6 = 16.8 cmΒ²
Scale for Area: 16.8 cmΒ² / 63 = 4/15 cmΒ² per unit frequency.
Class 31β35 (Boundaries 30.5 to 35.5):
Width = 5, Frequency = 60
Base = 5 Γ 0.4 cm = 2 cm
Area = 60 Γ (4/15) = 16 cmΒ²
Height = Area / Base = 16 / 2 = 8 cm
Base = 2 cm, Height = 8 cm
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 26
Class Midpoints (xβ) and Boundaries:
– 3β5 (2.5 to 5.5, width 3): Midpoint = 4, f = 12
– 6β7 (5.5 to 7.5, width 2): Midpoint = 6.5, f = 14
– 8 (7.5 to 8.5, width 1): Midpoint = 8, f = 19
– 9β11 (8.5 to 11.5, width 3): Midpoint = 10, f = 13
– 12β17 (11.5 to 17.5, width 6): Midpoint = 14.5, f = 6
N = 64
Ξ£fxβ = 12(4) + 14(6.5) + 19(8) + 13(10) + 6(14.5) = 48 + 91 + 152 + 130 + 87 = 508
Ξ£fxβΒ² = 12(16) + 14(42.25) + 19(64) + 13(100) + 6(210.25) = 192 + 591.5 + 1216 + 1300 + 1261.5 = 4561
a) Mean (xΜ
) = 508 / 64 β 7.94
Variance (ΟΒ²) = 4561 / 64 – (7.9375)Β² = 71.2656 – 63.0049 = 8.2607
Standard Deviation (Ο) = β8.2607 β 2.87
b) Median (Qβ): N / 2 = 32
Cumulative frequencies: 12, 26, 45…
Median class: 7.5 to 8.5 (f = 19, cf before = 26)
Qβ = 7.5 + ((32 – 26) / 19) Γ 1 = 7.5 + 0.316 β 7.82
c) Rectangle for class 3β5: Width = 3 βΉ Base = 1.2 cm (scale = 0.4 cm/unit)
Area = 1.2 Γ 5 = 6 cmΒ² βΉ scale = 6/12 = 0.5 cmΒ² per unit f.
For class 12β17 (width = 6, f = 6):
Base = 6 Γ 0.4 = 2.4 cm
Area = 6 Γ 0.5 = 3 cmΒ²
Height = 3 / 2.4 = 1.25 cm
d) Outlier Boundaries:
IQR = Qβ – Qβ = 9.19 – 6.07 = 3.12
Upper Boundary = Qβ + 1.5(IQR) = 9.19 + 1.5(3.12) = 13.87
Lower Boundary = Qβ – 1.5(IQR) = 6.07 – 1.5(3.12) = 1.39
Since max distance is 17.5 > 13.87, there are potential upper outliers.
e) Skewness: Mean (7.94) > Median (7.82), indicating slight positive skew. Normal distribution requires symmetry, so it may not be ideal.
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 27
Class Midpoints (xβ) and Widths:
– 1β3 (width 2): Midpoint = 2, f = 15
– 3β5 (width 2): Midpoint = 4, f = 31
– 5β6 (width 1): Midpoint = 5.5, f = 45
– 6β6.5 (width 0.5): Midpoint = 6.25, f = 37
– 6.5β7 (width 0.5): Midpoint = 6.75, f = 21
– 7β10 (width 3): Midpoint = 8.5, f = 15
N = 164
Ξ£fxβ = 15(2) + 31(4) + 45(5.5) + 37(6.25) + 21(6.75) + 15(8.5) = 30 + 124 + 247.5 + 231.25 + 141.75 + 127.5 = 902
Ξ£fxβΒ² = 15(4) + 31(16) + 45(30.25) + 37(39.0625) + 21(45.5625) + 15(72.25) = 60 + 496 + 1361.25 + 1445.3125 + 956.8125 + 1083.75 = 5403.125
a) Mean (xΜ
) = 902 / 164 = 5.5 kg
Variance = 5403.125 / 164 – 5.5Β² = 32.94588 – 30.25 = 2.69588
Standard Deviation (Ο) = β2.69588 β 1.64 kg
b) Median (Qβ): N / 2 = 82
Cumulative frequencies: 15, 46, 91…
Median class: 5 to 6 (f = 45, cf before = 46)
Qβ = 5 + ((82 – 46) / 45) Γ 1 = 5 + 36/45 = 5.80 kg
Since Mean (5.5) < Median (5.8), the distribution is negatively skewed.
c) Class 1 β€ w < 3: Width = 2, Base = 2.4 cm βΉ scale = 1.2 cm per unit width.
Area = 2.4 Γ 2.5 = 6 cmΒ² βΉ scale = 6 / 15 = 0.4 cmΒ² per unit f.
For class 6.5 β€ w < 7: Width = 0.5, f = 21
Base = 0.5 Γ 1.2 = 0.6 cm
Area = 21 Γ 0.4 = 8.4 cmΒ²
Height = 8.4 / 0.6 = 14 cm
d) Outlier limits:
IQR = 6.43 – 4.68 = 1.75
Lower Limit = 4.68 – 1.5(1.75) = 2.055
Upper Limit = 6.43 + 1.5(1.75) = 9.055
Values below 2.055 or above 9.055 are potential outliers.
e) Due to negative skewness and presence of potential outliers, a Normal distribution is not appropriate.
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 28
Coding: y = (x – 662.5) / 25 βΉ x = 25y + 662.5
Classes, Midpoints (x), Coded values (y), and Frequencies (f):
– 600β625: x = 612.5, y = -2, f = 11
– 625β650: x = 637.5, y = -1, f = 14
– 650β675: x = 662.5, y = 0, f = 28
– 675β700: x = 687.5, y = 1, f = 7
– 700β725: x = 712.5, y = 2, f = 5
– 725β750: x = 737.5, y = 3, f = 2
– 750β775: x = 762.5, y = 4, f = 1
N = 68
Ξ£fy = 11(-2) + 14(-1) + 28(0) + 7(1) + 5(2) + 2(3) + 1(4) = -22 – 14 + 0 + 7 + 10 + 6 + 4 = -9
Ξ£fyΒ² = 11(4) + 14(1) + 28(0) + 7(1) + 5(4) + 2(9) + 1(16) = 44 + 14 + 0 + 7 + 20 + 18 + 16 = 119
a) yΜ
= -9 / 68 β -0.13235
xΜ
= 25(-0.13235) + 662.5 β 659.19
Οα΅§Β² = 119 / 68 – (-0.13235)Β² = 1.75 – 0.01751 = 1.73249
Οα΅§ = β1.73249 β 1.3162
Οβ = 25 Γ 1.3162 β 32.91
b) Median (Qβ): N / 2 = 34
Cumulative frequencies: 11, 25, 53…
Median class: 650 to 675 (f = 28, cf before = 25)
Qβ = 650 + ((34 – 25) / 28) Γ 25 = 650 + 8.036 = 658.0
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 29
a) From histogram frequency densities:
Students scoring between 52 and 74 = 60
b) Total frequency N = 250
Median position = 125th student
Median (Qβ) β 49
c) Estimated Mean (xΜ
) β 51.8
Estimated Standard Deviation (Ο) β 22.22
Question 30
Group 1: nβ = 20, xΜ
β = 18.5, Οβ = 6.5
Group 2: nβ = 12, xΜ
β = 25, Οβ = 7.5
Combined N = 32
Combined Mean (xΜ
):
xΜ
= (nβxΜ
β + nβxΜ
β) / N = (20 Γ 18.5 + 12 Γ 25) / 32 = (370 + 300) / 32 = 670 / 32 β 20.94
Combined Variance (ΟΒ²):
Ξ£xβΒ² = nβ(ΟβΒ² + xΜ
βΒ²) = 20(6.5Β² + 18.5Β²) = 20(42.25 + 342.25) = 20(384.5) = 7690
Ξ£xβΒ² = nβ(ΟβΒ² + xΜ
βΒ²) = 12(7.5Β² + 25Β²) = 12(56.25 + 625) = 12(681.25) = 8175
Combined Ξ£xΒ² = 7690 + 8175 = 15865
Combined Variance = 15865 / 32 – (20.9375)Β² = 495.78125 – 438.3789 = 57.40235
Combined Standard Deviation = β57.40235 β 7.58
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 31
Total: N = 40, xΜ
= 65, Ο = 18
Boys: nβ = 24, xΜ
β = 72, Οβ = 20
Girls: nβ = 16, xΜ
β = ?, Οβ = ?
Mean for girls (xΜ
β):
nβxΜ
β + nβxΜ
β = N xΜ
24(72) + 16(xΜ
β) = 40(65)
1728 + 16xΜ
β = 2600
16xΜ
β = 872 βΉ xΜ
β = 54.5
Standard deviation for girls (Οβ):
Total Ξ£xΒ² = 40(18Β² + 65Β²) = 40(324 + 4225) = 40(4549) = 181960
Boys Ξ£xβΒ² = 24(20Β² + 72Β²) = 24(400 + 5184) = 24(5584) = 134016
Girls Ξ£xβΒ² = 181960 – 134016 = 47944
Girls Variance (ΟβΒ²) = 47944 / 16 – (54.5)Β² = 2996.5 – 2970.25 = 26.25
Girls Standard Deviation (Οβ) = β26.25 β 5.12
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 32
Given:
n = 40
Ξ£(x – 50) = 150
Ξ£(x – 50)Β² = 4650
Expand Ξ£(x – 50)Β²:
Ξ£(xΒ² – 100x + 2500) = 4650
Ξ£xΒ² – 100Ξ£x + 2500(40) = 4650
Ξ£xΒ² – 100Ξ£x + 100000 = 4650 — (Equation 1)
From Ξ£(x – 50) = 150:
Ξ£x – 50(40) = 150
Ξ£x – 2000 = 150 βΉ Ξ£x = 2150
Substitute Ξ£x into Equation 1:
Ξ£xΒ² – 100(2150) + 100000 = 4650
Ξ£xΒ² – 215000 + 100000 = 4650
Ξ£xΒ² – 115000 = 4650
Ξ£xΒ² = 119650
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 33
Given:
xΜ
= (1/n) Ξ£xα΅£ βΉ Ξ£xα΅£ = n xΜ
Variance ΟΒ² = (1/n) Ξ£xα΅£Β² – xΜ
Β² = 3 βΉ Ξ£xα΅£Β² = n(3 + xΜ
Β²)
We want to find Ξ£(xα΅£ + 1)Β²:
Ξ£(xα΅£ + 1)Β² = Ξ£(xα΅£Β² + 2xα΅£ + 1)
= Ξ£xα΅£Β² + 2Ξ£xα΅£ + Ξ£1
= n(3 + xΜ
Β²) + 2(n xΜ
) + n
= 3n + n xΜ
Β² + 2n xΜ
+ n
= n xΜ
Β² + 2n xΜ
+ 4n
= n(xΜ
Β² + 2xΜ
+ 4)
Given xΜ
= 3:
Ξ£(xα΅£ + 1)Β² = n(3Β² + 2(3) + 4) = n(9 + 6 + 4) = 19n [or n(1 + 18) = 19n]
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Question 34
Given:
n = 20
Ξ£(x – 10) = 220
Ξ£(x – 10)Β² = 2720
a) Show Ξ£xΒ² = 9120:
Ξ£(x – 10) = 220 βΉ Ξ£x – 200 = 220 βΉ Ξ£x = 420
Expand Ξ£(x – 10)Β²:
Ξ£(xΒ² – 20x + 100) = 2720
Ξ£xΒ² – 20Ξ£x + 2000 = 2720
Ξ£xΒ² – 20(420) + 2000 = 2720
Ξ£xΒ² – 8400 + 2000 = 2720
Ξ£xΒ² – 6400 = 2720
Ξ£xΒ² = 9120 (Proved)
b) Mean (xΜ
) & Standard Deviation (Ο):
xΜ
= Ξ£x / n = 420 / 20 = 21
ΟΒ² = Ξ£xΒ² / n – xΜ
Β² = 9120 / 20 – 21Β² = 456 – 441 = 15
Ο = β15 β 3.87
ββββββββββββββββββββββββββββββββββββββββββββββββββββ
Q1: Which one of the following data types could best be described as Personally Identifiable Information (PII)?
Answer: C Shipping addresses for customers’ most recent orders.
Q2: You are currently gathering data relating to coastal erosion, in order to predict the measurement of the erosion in another 20 years’ time. It includes a series of measurements that have been recorded over a period of 50 years. Which one of the following types of data would you most likely use in your analysis?
Answer: A Continuous.
Q3: Which one of the following data structures includes the use of a parent node?
Answer: C Tree.
Q4: You want to gain insight into the influence your customers have on brand visibility. You have a structured data source in the form of a customer relationship management (CRM) system, as well as unstructured data from social media feeds. What would be the main benefit of using this unstructured data alongside the structured data source?
Answer: B You could identify how many times your brand’s name is mentioned.
Q5: You have been tasked to produce a monthly report on sales from the previous month. What sort of analytics would you use?
Answer: B Descriptive analytics.
Q6: You want to create a data model that describes the technical requirements of a data analysis project. The intended audience will be non-technical company directors. Select the most appropriate data model from the following options.
Answer: A Conceptual data model.
Q7: Which one of the following should define what data is collected and stored in an organisation?
Answer: A Architectural policies.
Q8: Which one of the following data architecture functions would support business intelligence activities on historical data?
Answer: C Data warehousing.
Q9: The rate at which data is generated, collected, processed and analysed describes which challenge associated with big data?
Answer: B Velocity.
Q10: When talking about big data, which one of the following could you reasonably expect to see in your original datasets?
Answer: C Unstructured content.
Q11: Which one of following would you expect to feature in an entity relationship diagram (ERD)?
Answer: A Attributes.
Q12: Which one of the following is an advantage of a relational database over a NoSQL database?
Answer: B Speed of transactions for low volumes of data.
Q13: Which stage of ETL would cleanse data?
Answer: B Transform.
Q14: Which one of the following operations would you expect to happen in the Extract phase of ETL?
Answer: D Staging data from source systems.
Q15: Which one of the following has a key purpose of being able to create “a highly intuitive, drag-and-drop interface for building visualisations, reports, and dashboards”?
Answer: C PowerBI.
Q16: Which one of the following is a drawback when using the mean average of a set of data?
Answer: A Extreme values can heavily influence the result.
Q17: Which one of the following is a drawback when using the mode average of a set of data?
Answer: B It may not provide a single value as the answer.
Q18: Which one of the following is a drawback when using the median average of a set of data?
Answer: A The result may not appear in the original data set.
Q19: When considering strategies to improve data analysis modelling, which one of the following methods is most effective?
Answer: D Adding more data to the training set.
Q20: When visualising data for stakeholders, which one of the following is the key factor to consider?
Answer: B The accessibility of the presented data and ease of understanding.
Scenario 1: Database design and SQL
Q21: More information is required on employees to help escalate issues with IT asset conditions. It has been decided to add another table storing information on an employee’s manager. What new field would be a suitable primary key for the βManagersβ table?
Answer: A employee_id.
Q22: You have been tasked with creating a data dictionary for this database. Which one of the following should be listed in a data dictionary?
Answer: D field size.
Q23: A key report measure on the scorecard is the number of assets returned in a month, and their condition when returned. Which two of the following aggregations should be used to query the ‘Employee_Assets’ table for this measure?
Answer: A Group By & B Count.
Q24: If an asset_id is removed from the βIT_assetsβ table, all related information for that asset should also be removed from the other tables. What database functionality ensures that this happens?
Answer: C Cascading delete.
Q25: Fill in the blanks to complete the SQL query shown below:
SELECT ___________(employee_id), department from ___________where first_name = "" group by ___________ order by ___________
Answer: SELECT COUNT(employee_id), department from EMPLOYEES where first_name = "" group by department order by 2
Scenario 2: Data Preparation and Integration
Q26: Order the following Python commands into a logical flow for importing the data contained in the file called “IToutages.csv”. After importing the entire file, you should print the first ten characters of data to screen.
Answer:
-
f = open("IToutages.csv") -
data=f.read() -
FirstTenChars = data[:10] -
print(FirstTenChars) -
f.close
Q27: Which one of the following is the missing line of code in the Python programme below to find the mean of 2,3,4,5,6?
Answer: C Mean=Total/5.
Q28: Which one of the following R commands will correctly show different averages and quartiles of a dataset?
Answer: D summary(dataset)
Q29: Which one of the following R commands would be used to see a snapshot of the first six rows in a dataset?
Answer: A head(dataset)
Q30: Order the following lines of code into a logical flow to read in the ‘ITemployees.txt’ file and plot the data as an annual time series, then change the timeseries to be monthly and plot a second graph that starts in 1999.
Answer:
-
EmployeesByYear <- read.csv("ITemployees.txt") -
employeetimeseries <- ts(EmployeesByYear) -
plot.ts(employeetimeseries) -
employeetimeseries <- ts(EmployeesByYear, frequency = 12, start = 1999) -
monthplot(employeetimeseries)
