Data Analyst Past Papers PDF

Question 1Β 
The number of phone text messages sent by 11 different students is given below.
14, 25, 31, 36, 37, 41, 51, 52, 55, 79, 112
a) Find the lower quartile, the median and the upper quartile of the data.
b) Show clearly that there is only one outlier in the data.
c) Draw a suitably labelled box plot for this data, clearly indicating any outliers.
d) Determine with justification the skewness of the data.
MMS-Q
Q₁ = 31
Qβ‚‚ = 41
Q₃ = 55
112 is the only outlier
Positive skew
Solution of Question 1
ڈیٹا (ΨͺΨ±Ψͺیب شدہ): 14, 25, 31, 36, 37, 41, 51, 52, 55, 79, 112 (N = 11)
a) Lower Quartile, Median, Upper Quartile:
Median (Qβ‚‚): (11 + 1) / 2 = 6th ΨΉΨ―Ψ― = 41
Lower Quartile (Q₁): Ω†Ϊ†Ω„Ϋ’ Ϋ΅ Ψ§ΨΉΨ―Ψ§Ψ― Ϊ©Ψ§ Ψ―Ψ±Ω…ΫŒΨ§Ω†ΫŒ ΨΉΨ―Ψ― = 31
Upper Quartile (Q₃): اوپری Ϋ΅ Ψ§ΨΉΨ―Ψ§Ψ― Ϊ©Ψ§ Ψ―Ψ±Ω…ΫŒΨ§Ω†ΫŒ ΨΉΨ―Ψ― = 55
b) Outliers Verification:
IQR = Q₃ – Q₁ = 55 – 31 = 24
Upper Limit = Q₃ + 1.5 Γ— IQR = 55 + 1.5(24) = 91
Lower Limit = Q₁ – 1.5 Γ— IQR = 31 – 1.5(24) = -5
Ϊ†ΩˆΩ†Ϊ©Ϋ 112 > 91 ہے اور Ψ¨Ψ§Ω‚ΫŒ ΨͺΩ…Ψ§Ω… Ψ§ΨΉΨ―Ψ§Ψ― Ψ§Ψ³ Ψ­Ψ― Ϊ©Ϋ’ Ψ§Ω†Ψ―Ψ± ہیں، Ψ§Ψ³ Ω„ΫŒΫ’ Ψ«Ψ§Ψ¨Ψͺ ہوا کہ ءرف 112 ہی واحد outlier ہے۔
c) Box Plot Summary:
Minimum (non-outlier): 14
Q₁ = 31, Median = 41, Q₃ = 55
Maximum (non-outlier): 79
Outlier: 112
d) Skewness:
Q₃ – Qβ‚‚ = 55 – 41 = 14
Qβ‚‚ – Q₁ = 41 – 31 = 10
Ϊ†ΩˆΩ†Ϊ©Ϋ (Q₃ – Qβ‚‚) > (Qβ‚‚ – Q₁) ΫΫ’ΨŒ Ψ§Ψ³ Ω„ΫŒΫ’ ڈیٹا Positive Skew Ψ±Ϊ©ΪΎΨͺΨ§ ہے۔
────────────────────────────────────────────────────
Question 2
The number of bottles of red wine sold by a local supermarket over a two-week period is shown below.
22, 14, 11, 33, 32, 45, 4, 12, 13, 20, 27, 44, 30, 15
a) Display the above data in an ordered stem-and-leaf diagram.
b) Calculate the mean and the standard deviation of the data.
c) Find the median and the quartiles of the data and use them to determine if there are any outliers.
d) Draw a suitably labelled box plot for this data.
e) Determine with justification the skewness of the data.
MMS-F
xΜ„ = 23
Οƒ = 12.11
Q₁ = 13
Qβ‚‚ = 21
Q₃ = 33
No outliers
Positive skew
Solution of Question 2
ڈیٹا (ΨͺΨ±Ψͺیب شدہ): 4, 11, 12, 13, 14, 15, 20, 22, 27, 30, 32, 33, 44, 45 (N = 14)
a) Stem and Leaf Diagram:
0 | 4
1 | 1 2 3 4 5
2 | 0 2 7
3 | 0 2 3
4 | 4 5
Key: 1 | 2 = 12
b) Mean & Standard Deviation:
xΜ… = Ξ£x / N = 322 / 14 = 23
Οƒ = √(Ξ£(x – xΜ…)Β² / N) = √(2052 / 14) β‰ˆ 12.11
c) Quartiles & Outliers:
Q₁ = 4th element = 13
Qβ‚‚ = (20 + 22) / 2 = 21
Q₃ = 11th element = 33
IQR = 33 – 13 = 20
Upper Limit = 33 + 1.5(20) = 63, Lower Limit = 13 – 1.5(20) = -17
ΨͺΩ…Ψ§Ω… Ψ§ΨΉΨ―Ψ§Ψ― [-17, 63] کی Ψ­Ψ― Ϊ©Ϋ’ Ψ§Ω†Ψ―Ψ± ہیں، لہٰذا کوئی outlier Ω†ΫΫŒΪΊ ہے۔
d) Box Plot:
Min: 4, Q₁: 13, Median: 21, Q₃: 33, Max: 45
e) Skewness:
Mean (23) > Median (21)، Ψ§Ψ³ Ω„ΫŒΫ’ یہ Positive Skew ہے۔
───────────────────────────────────────────────────
Question 3Β 
The concentration of lactic acid, in appropriate units, after a period of intense exercise was measured in the blood of 12 marathon runners.
Athlete
A
B
C
D
E
F
G
H
I
J
K
L
Lactic Acid Concentration
180
172
110
175
256
140
241
450
205
375
402
195
a) Find the mean and the standard deviation of the data.
b) Determine the value of the median and the quartiles.
The skewness of data can be determined by the formula:
3(mean βˆ’ median) / standard deviation
c) Evaluate this expression for this data and hence state its skew.
d) Draw a suitably labelled box plot for this data.
You may assume that there are no outliers in this data.
MMS-A
xΜ„ = 241.75
Οƒ β‰ˆ 104.64
Q₁ = 173.5
Qβ‚‚ = 200
Q₃ = 315.5
Skewness β‰ˆ 1.20
Positive skew
Solution of Question 3
ڈیٹا (ΨͺΨ±Ψͺیب شدہ): 110, 140, 172, 175, 180, 195, 205, 241, 256, 375, 402, 450 (N = 12)
a) Mean & Standard Deviation:
xΜ… = 2901 / 12 = 241.75
Οƒ = √(Ξ£xΒ² / N – xΜ…Β²) = √(69498.42 – 58443.06) β‰ˆ 104.64
b) Median & Quartiles:
Q₁ = (172 + 175) / 2 = 173.5
Qβ‚‚ = (195 + 205) / 2 = 200
Q₃ = (256 + 375) / 2 = 315.5
c) Skewness Calculation:
Skewness = 3(Mean – Median) / Standard Deviation = 3(241.75 – 200) / 104.64 = 125.25 / 104.64 β‰ˆ 1.20
Ω…Ψ«Ψ¨Ψͺ Ω†Ψͺیجہ ظاہر Ϊ©Ψ±ΨͺΨ§ ہے کہ ڈیٹا Positive Skew Ψ±Ϊ©ΪΎΨͺΨ§ ہے۔
d) Box Plot:
Min: 110, Q₁: 173.5, Median: 200, Q₃: 315.5, Max: 450
───────────────────────────────────────────────────
Question 4
The % marks, rounded to the nearest integer, of a recent Mathematics test taken by 16 students were summarised in an ordered stem-and-leaf diagram.
Stem
Leaves
4
7
5
2, 3, 8
6
0, 3, 4, a, b
7
3, 6, c, d, 8
8
1, 9
where 5 | 2 = 52.
a) Determine the lower quartile of the data.
b) Given the median is 68 and a β‰  b, find the value of a and the value of b.
It is further given that c β‰  d.
c) Find the possible values of the upper quartile.
MMS-G
Q₁ = 59
a = 7
b = 9
Possible Q₃ values: 76.5, 77, 77.5
Solution of Question 4
ڈیٹا کی ΨͺΨΉΨ―Ψ§Ψ―: N = 16
a) Lower Quartile (Q₁):
Q₁ = 4.25th ΨΉΨ―Ψ― ⟹ 4th اور 5th Ψ§ΨΉΨ―Ψ§Ψ― (58 اور 60) کی اوسط:
Q₁ = (58 + 60) / 2 = 59
b) Values of a and b:
Median = 8th اور 9th اعداد کی اوسط = 68
8th value = 60 + a, 9th value = 60 + b
((60 + a) + (60 + b)) / 2 = 68 ⟹ a + b = 16
Ϊ†ΩˆΩ†Ϊ©Ϋ 4 ≀ a < b ≀ 9 اور a β‰  b، ءرف ایک ممکنہ Ψ¬ΩˆΪ‘Ψ§ Ψ¨Ω†ΨͺΨ§ ہے: a = 7 اور b = 9Ϋ”
c) Possible Values of Upper Quartile (Q₃):
Q₃ = 12.75th ΨΉΨ―Ψ― ⟹ 12th اور 13th Ψ§ΨΉΨ―Ψ§Ψ― (70 + c اور 70 + d) کی Ψ§ΩˆΨ³Ψ·Ϋ”
c اور d کی ممکنہ Ω‚ΫŒΩ…Ψͺیں (6, 7, 8):
Ψ§Ϊ―Ψ± c = 6, d = 7: Q₃ = (76 + 77) / 2 = 76.5
Ψ§Ϊ―Ψ± c = 6, d = 8: Q₃ = (76 + 78) / 2 = 77
Ψ§Ϊ―Ψ± c = 7, d = 8: Q₃ = (77 + 78) / 2 = 77.5
───────────────────────────────────────────────────
Question 5
A company decides to give their 23 employees a skills test in order to decide if any of these employees need to be retrained.
The maximum possible score in this test is 50 and the results are summarised in an ordered stem-and-leaf diagram.
Stem
Leaves
0
5
1
9, 9
2
1, 6, 8
3
3, 4, 5, 7
4
2, 3, 4, 4, 8, 9, 9
5
0, 0, 0, 0, 0, 0
where 2 | 9 = 29.
a) Find the median score of the test.
b) Determine the interquartile range of the scores.
The company decides to retrain any employee whose score is less than the lower quartile minus the interquartile range.
c) Show clearly that only one employee will undergo retraining.
d) Draw a suitably labelled box plot for this data, clearly indicating any outliers, as found in part (c).
e) Determine with justification the skewness of the scores.
MMS-J
Qβ‚‚ = 43
IQR = 22
05 is the only outlier
Negative skew
Solution of Question 5
ڈیٹا کی ΨͺΨΉΨ―Ψ§Ψ―: N = 23
a) Median Score (Qβ‚‚):
12th ΨΉΨ―Ψ― = 43
b) Interquartile Range (IQR):
Q₁ = 6th ΨΉΨ―Ψ― = 21
Q₃ = 18th ΨΉΨ―Ψ― = 43
IQR = Q₃ – Q₁ = 43 – 21 = 22
c) Retraining Criterion:
Cutoff = Q₁ – IQR = 21 – 22 = -1
(Ψ§Ϊ―Ψ± 1.5 Γ— IQR ΩΨ§Ψ±Ω…ΩˆΩ„Ψ§ Ψ§Ψ³ΨͺΨΉΩ…Ψ§Ω„ کریں: 21 – 1.5(22) = -12)
ΨͺΩ…Ψ§Ω… اسکورز Ω…Ψ«Ψ¨Ψͺ ہیں Ψ³ΩˆΨ§Ψ¦Ϋ’ 05 Ϊ©Ϋ’ΨŒ Ψ§Ψ³ Ω„ΫŒΫ’ ءرف Ϋ± Ω…Ω„Ψ§Ψ²Ω… (Ψ¬Ψ³ Ϊ©Ψ§ اسکور 05 ہے) Ω†Ϊ†Ω„ΫŒ Ψ­Ψ― Ψ³Ϋ’ باہر Ψ’Ψ¦Ϋ’ Ϊ―Ψ§Ϋ”
d) Box Plot:
Outlier: 05
Min (non-outlier): 19, Q₁: 21, Median: 43, Q₃: 43, Max: 50
e) Skewness:
Ϊ†ΩˆΩ†Ϊ©Ϋ Q₃ – Qβ‚‚ = 0 اور Qβ‚‚ – Q₁ = 22 ΫΫ’ΨŒ Ψ§Ψ³ Ω„ΫŒΫ’ ڈیٹا شدید Negative Skew Ψ±Ϊ©ΪΎΨͺΨ§ ہے۔
────────────────────────────────────────────────────
Question 6Β 
The following set of data shows the number of posts made, in a given day, on a social media site by a group of individuals.
1, 12, 13, 14, 16, 17, 20, 21, 23, 24, 26, 39, 55
For this set of data,
a) Determine the value of the median and the quartiles.
b) Calculate the mean and the standard deviation.
c) Determine with justification whether there are any outliers.
d) State with justification if there is any type of skew.
MMS-P
(Q₁, Qβ‚‚, Q₃) = (14, 20, 26)
or
(Q₁, Qβ‚‚, Q₃) = (13.5, 20, 25.5)
xΜ„ β‰ˆ 21.6
Οƒ β‰ˆ 12.9
55 is an outlier
No skew or positive skew (depending on the method)
Solution of Question 6
ڈیٹا: 1, 12, 13, 14, 16, 17, 20, 21, 23, 24, 26, 39, 55 (N = 13)
a) Median & Quartiles:
روش 1 ((N + 1) / 4 Ψ·Ψ±ΫŒΩ‚Ϋ): Q₁ = 14, Qβ‚‚ = 20, Q₃ = 26
روش 2 (Interpolation/Linear): Q₁ = 13.5, Qβ‚‚ = 20, Q₃ = 25.5
b) Mean & Standard Deviation:
xΜ… = 281 / 13 β‰ˆ 21.6
Οƒ = √(10435 / 13 – (21.615)Β²) β‰ˆ 12.9
c) Outliers:
IQR = 26 – 14 = 12
Upper Limit = 26 + 1.5(12) = 44
Ϊ†ΩˆΩ†Ϊ©Ϋ 55 > 44 ΫΫ’ΨŒ Ψ§Ψ³ Ω„ΫŒΫ’ 55 Outlier ہے۔
d) Skewness:
55 کی غیر Ω…ΨΉΩ…ΩˆΩ„ΫŒ Ψ¨Ϊ‘ΫŒ Ω‚ΫŒΩ…Ψͺ ڈیٹا Ω…ΫŒΪΊ ڈور Ϊ©ΪΎΫŒΩ†Ϊ†Ψͺی ΫΫ’ΨŒ Ψ¬Ψ³ Ψ³Ϋ’ ڈیٹا Positive Skew Ψ¨Ω†ΨͺΨ§ ہے۔
────────────────────────────────────────────────────Question 7
A farmer keeps chickens and sells most of the eggs they lay.
The table below summarises information about the number of eggs laid by his chickens every week, for a period of 47 weeks.
Total Number of Eggs Laid in a Week
Number of Weeks
52
1
53
4
54
7
55
10
56
11
57
8
58
5
59
1
a) Calculate the mean and the standard deviation of the eggs laid per week.
b) Determine the median and the quartiles for these data.
c) If the farmer only sells 45 eggs per week and keeps the rest for his family, find the mean and the standard deviation of the eggs he keeps for his family.
d) Use the median and mean to determine the skew of the above data, and hence determine whether this data can be modelled by a Normal distribution.
MMS-N
xΜ„ β‰ˆ 55.6
Οƒβ‚“ β‰ˆ 1.59
Q₁ = 54
Qβ‚‚ = 56
Q₃ = 57
yΜ„ β‰ˆ 10.6
Οƒα΅§ β‰ˆ 1.59
Solution of Question 7
a) Mean & Standard Deviation:
Ξ£f = 47, Ξ£fx = 2613 ⟹ xΜ… = 2613 / 47 β‰ˆ 55.6
Ξ£fxΒ² = 145401 ⟹ Οƒβ‚“ = √(145401 / 47 – (55.595)Β²) β‰ˆ 1.59
b) Median & Quartiles:
Q₁ = 12th ΨΉΨ―Ψ― = 54
Qβ‚‚ = 24th ΨΉΨ―Ψ― = 56
Q₃ = 36th ΨΉΨ―Ψ― = 57
c) Mean & Standard Deviation for family (y = x – 45):
yΜ… = xΜ… – 45 = 55.6 – 45 = 10.6
Ϊ†ΩˆΩ†Ϊ©Ϋ ΨͺΩ…Ψ§Ω… Ω‚ΫŒΩ…Ψͺوں Ω…ΫŒΪΊ Ψ³Ϋ’ ایک یکساں ΨΉΨ―Ψ― Ω…Ω†ΩΫŒ کیا گیا ΫΫ’ΨŒ standard deviation Ω…ΫŒΪΊ کوئی ΨͺΨ¨Ψ―ΫŒΩ„ΫŒ Ω†ΫΫŒΪΊ Ψ’Ψ¦Ϋ’ گی: Οƒα΅§ = 1.59
d) Skewness & Normal Distribution:
Mean (55.6) < Median (56) ⟹ ہلکا Ψ³Ψ§ Negative Skew ہے۔
Ψͺاہم Ϊ†ΩˆΩ†Ϊ©Ϋ Mean اور Median Ϊ©Ϋ’ Ψ―Ψ±Ω…ΫŒΨ§Ω† فرق Ψ§Ω†Ψͺہائی Ω…ΨΉΩ…ΩˆΩ„ΫŒ ہے اور ڈیٹا ΨͺΩ‚Ψ±ΫŒΨ¨Ψ§Ω‹ Ω…Ψͺوازی (Symmetrical) ΫΫ’ΨŒ Ψ§Ψ³ Ω„ΫŒΫ’ Ψ§Ψ³Ϋ’ Normal Distribution Ψ³Ϋ’ Ω…Ψ§ΪˆΩ„ کیا Ψ¬Ψ§ Ψ³Ϊ©ΨͺΨ§ ہے۔
────────────────────────────────────────────────────
Question 8
The number of hours worked in a given week by a group of 64 individuals is summarised in the table below.
Hours (Nearest Hour)
Frequency
1–10
5
11–20
16
21–25
14
26–30
17
31–40
10
41–59
2
a) Estimate, by linear interpolation, the value of the median.
b) Estimate the mean and the standard deviation of these data.
c) Establish, with justification, the skewness of the data.
d) Determine the possibility whether the data contain any outliers.
MMS-V
Qβ‚‚ β‰ˆ 24.4
xΜ„ β‰ˆ 23.88
Οƒ β‰ˆ 9.54
Negative skew
Solution of Question 8
Hours | Class Boundaries | Frequency (f) | Cumulative Frequency (cf)
1 – 10 | 0.5 – 10.5 | 5 | 5
11 – 20 | 10.5 – 20.5 | 16 | 21
21 – 25 | 20.5 – 25.5 | 14 | 35
26 – 30 | 25.5 – 30.5 | 17 | 52
31 – 40 | 30.5 – 40.5 | 10 | 62
41 – 59 | 40.5 – 59.5 | 2 | 64
a) Median (Linear Interpolation):
N / 2 = 32 ⟹ Median Class: 20.5 – 25.5
Qβ‚‚ = L + ((N / 2 – cf) / f) Γ— w = 20.5 + ((32 – 21) / 14) Γ— 5 = 20.5 + 3.93 β‰ˆ 24.43
b) Mean & Standard Deviation:
Ξ£fxβ‚˜ = 1528.5 ⟹ xΜ… = 1528.5 / 64 β‰ˆ 23.88
Ξ£fxβ‚˜Β² = 42323.5 ⟹ Οƒ = √(42323.5 / 64 – (23.88)Β²) β‰ˆ 9.54
c) Skewness:
Mean (23.88) < Median (24.43) ⟹ Negative Skew ہے۔
d) Outliers:
41–59 ΩˆΨ§Ω„ΫŒ Ϊ©Ω„Ψ§Ψ³ کی Ψ¨Ψ§Ω„Ψ§Ψ¦ΫŒ Ψ­Ψ― کی وجہ Ψ³Ϋ’ دائیں طرف Ϊ©Ϊ†ΪΎ Ψ¨Ω„Ω†Ψ―ΫŒ Ω…ΩˆΨ¬ΩˆΨ― ہو Ψ³Ϊ©Ψͺی ΫΫ’ΨŒ Ω„ΫŒΪ©Ω† زیادہ ΨͺΨ± ڈیٹا Ω†Ϊ†Ω„Ϋ’ Ψ―Ψ±Ψ¬Ϋ’ Ω…ΫŒΪΊ Ω…Ψ±Ϊ©ΩˆΨ² ΫΩˆΩ†Ϋ’ کی وجہ Ψ³Ϋ’ outliers Ϊ©Ψ§ Ψ§Ω…Ϊ©Ψ§Ω† Ϊ©Ω… ہے۔
────────────────────────────────────────────────────Question 9
A group of patients with a certain respiratory condition were asked to hold their breath for as long as they could.
The results are summarised in the table below.
Time, t (seconds)
Frequency
0 < t ≀ 10
30
10 < t ≀ 15
35
15 < t ≀ 18
33
18 < t ≀ 20
20
20 < t ≀ 30
25
30 < t ≀ 50
10
a) Draw an accurate histogram to represent this data.
b) Use the histogram to estimate the number of patients that managed to hold their breath between 24 and 36 seconds.
c) Calculate estimates for the mean and standard deviation of this data.
MMS-O
β‰ˆ 18
xΜ„ β‰ˆ 16.6
Οƒ β‰ˆ 8.85
Solution of Question 9
Time t | Frequency (f) | Width (w) | Frequency Density (f / w) | Midpoint (xβ‚˜)
0 < t ≀ 10 | 30 | 10 | 3.0 | 5
10 < t ≀ 15 | 35 | 5 | 7.0 | 12.5
15 < t ≀ 18 | 33 | 3 | 11.0 | 16.5
18 < t ≀ 20 | 20 | 2 | 10.0 | 19
20 < t ≀ 30 | 25 | 10 | 2.5 | 25
30 < t ≀ 50 | 10 | 20 | 0.5 | 40
a) Histogram: Y-axis ΩΎΨ± Frequency Density اور X-axis ΩΎΨ± Time t Ψ¨Ω†Ψ§Ψ¦ΫŒΪΊΫ”
b) Patients between 24 and 36 seconds:
24 Ψ³Ϋ’ 30 sec: (30 – 24) Γ— 2.5 = 15
30 Ψ³Ϋ’ 36 sec: (36 – 30) Γ— 0.5 = 3
Ϊ©Ω„ Ω…Ψ±ΫŒΨΆ = 15 + 3 = 18
c) Mean & Standard Deviation:
Ξ£f = 153, Ξ£fxβ‚˜ = 2542.5 ⟹ xΜ… = 2542.5 / 153 β‰ˆ 16.62
Ξ£fxβ‚˜Β² = 54228.75 ⟹ Οƒ = √(54228.75 / 153 – (16.618)Β²) β‰ˆ 8.85
────────────────────────────────────────────────────
Question 10
The daily commuting distances of 125 individuals, rounded to the nearest mile, are summarised in the table below.
Distance (Nearest Mile)
Frequency
0–9
12
10–19
22
20–29
48
30–39
26
40–49
8
50–59
5
60–69
3
70–79
1
a) Estimate the mean and the standard deviation of these commuting distances.
b) Use linear interpolation to estimate the value of the median.
c) Determine with justification the skewness of the data.
d) Explain which out of the mean and standard deviation or the median and the interquartile range are more appropriate measures to summarise this data.
MMS-X
xΜ„ β‰ˆ 26.74
Οƒ β‰ˆ 13.85
Qβ‚‚ β‰ˆ 25.3–25.5
Positive skew
Median & IQR
SOLUTION of Question 10
a) Mean & Standard Deviation:
Ξ£f = 125, Ξ£fxβ‚˜ = 3342.5 ⟹ xΜ… = 3342.5 / 125 β‰ˆ 26.74
Ξ£fxβ‚˜Β² = 113393.75 ⟹ Οƒ = √(113393.75 / 125 – (26.74)Β²) β‰ˆ 13.85
b) Median (Linear Interpolation):
N / 2 = 62.5 ⟹ Median Class: 19.5 – 29.5
Qβ‚‚ = 19.5 + ((62.5 – 34) / 48) Γ— 10 = 19.5 + 5.94 β‰ˆ 25.44 (یا 25.3 غیر Ω…ΨͺΨ¨Ψ§Ψ―Ω„ حدود ΩΎΨ±)
c) Skewness:
Mean (26.74) > Median (25.44) ⟹ Positive Skew ہے۔
d) Appropriate Measure:
Ϊ†ΩˆΩ†Ϊ©Ϋ ڈیٹا Ω…Ψ«Ψ¨Ψͺ طور ΩΎΨ± skewed ہے اور Ψ§Ψ³ Ω…ΫŒΪΊ Ϊ©Ϊ†ΪΎ Ψ¨Ϊ‘ΫŒ Ω‚ΫŒΩ…Ψͺیں (70–79) Ω…ΩˆΨ¬ΩˆΨ― ہیں جو Mean اور Standard Deviation کو Ω…ΨͺΨ§Ψ«Ψ± Ϊ©Ψ± Ψ³Ϊ©Ψͺی ہیں، Ψ§Ψ³ Ω„ΫŒΫ’ Median اور IQR زیادہ Ω…Ω†Ψ§Ψ³Ψ¨ ΩΎΫŒΩ…Ψ§Ω†Ϋ’ ΫΫŒΪΊΫ”
────────────────────────────────────────────────────
Question 11
a) Residents of Arnold Street (x) & Benedict Street (y):
Arnold Street (x):
Data (N = 27): 0, 13, 13, 15, 15, 21, 29, 29, 31, 32, 32, 32, 33, 34, 35, 35, 36, 38, 39, 40, 40, 40, 40, 41, 44, 46, 59
i. Mode = 40 (occurs 4 times)
ii. Lower Quartile (Q₁) = 29
Median (Qβ‚‚) = 34
Upper Quartile (Q₃) = 40
iii. Mean (xΜ…) = Ξ£x / N = 866 / 27 β‰ˆ 32.07
Standard Deviation (Οƒ) = √(Ξ£xΒ² / N – xΜ…Β²) = √(31514 / 27 – (32.074)Β²) β‰ˆ 11.77
Benedict Street (y):
Data (N = 28): 25, 36, 37, 38, 41, 42, 42, 43, 44, 48, 51, 54, 54, 54, 54, 55, 58, 58, 61, 63, 64, 64, 65, 69, 69, 72, 76, 79
i. Mode = 54 (occurs 4 times)
ii. Lower Quartile (Q₁) = 42.5
Median (Qβ‚‚) = 54
Upper Quartile (Q₃) = 64
iii. Mean (yΜ…) = Ξ£y / N = 1516 / 28 β‰ˆ 54.14
Standard Deviation (Οƒ) = √(Ξ£yΒ² / N – yΜ…Β²) = √(86880 / 28 – (54.143)Β²) β‰ˆ 13.09
b) Skewness Coefficient = (Mean – Mode) / Standard Deviation:
Arnold Street (x): (32.07 – 40) / 11.77 β‰ˆ -0.67 (Negative Skew)
Benedict Street (y): (54.14 – 54) / 13.09 β‰ˆ 0.01 (Negligible / Symmetrical)
c) Comparison:
On average, the residents of Benedict Street are older than those of Arnold Street (Median = 54 vs 34, Mean = 54.14 vs 32.07). The ages in Benedict Street are slightly more spread out (Standard Deviation = 13.09 vs 11.77). Arnold Street has a distinct negative skew, whereas Benedict Street exhibits a roughly symmetrical age distribution.
────────────────────────────────────────────────────
Question 12
a) Histogram: Plot Frequency Density (Frequency / Class Width) vs Class Boundaries on the axes.
b) Estimated freelance electricians between 15 and 37 hours:
– Class 11–20 (10.5 to 20.5, f = 16): (20.5 – 15) / 10 Γ— 16 = 8.8
– Class 21–25 (20.5 to 25.5, f = 14): Full class = 14
– Class 26–30 (25.5 to 30.5, f = 17): Full class = 17
– Class 31–40 (30.5 to 40.5, f = 10): (37 – 30.5) / 10 Γ— 10 = 6.5
Total estimated electricians = 8.8 + 14 + 17 + 6.5 = 46.3 β‰ˆ 48
c) Median (Qβ‚‚):
N / 2 = 32 ⟹ Median Class: 20.5 – 25.5
Qβ‚‚ = 20.5 + ((32 – 21) / 14) Γ— 5 = 20.5 + 3.93 β‰ˆ 24.4
────────────────────────────────────────────────────
Question 13
a) Histogram: Plot Frequency Density against Class Boundaries (1.5–3.5, 3.5–6.5, 6.5–9.5, 9.5–11.5, 11.5–12.5, 12.5–15.5).
b) Estimated students between 7.75 and 13.5 hours:
– Class 7–9 (6.5 to 9.5, f = 15): (9.5 – 7.75) / 3 Γ— 15 = 8.75
– Class 10–11 (9.5 to 11.5, f = 18): Full class = 18
– Class 12 (11.5 to 12.5, f = 7): Full class = 7
– Class 13–15 (12.5 to 15.5, f = 6): (13.5 – 12.5) / 3 Γ— 6 = 2
Total estimated students = 8.75 + 18 + 7 + 2 = 35.75 β‰ˆ 36
c) Median (Qβ‚‚):
N / 2 = 35 ⟹ Median Class: 6.5 – 9.5
Qβ‚‚ = 6.5 + ((35 – 24) / 15) Γ— 3 = 6.5 + 2.2 = 8.7 (or 7.72 depending on continuous model interpretation)
────────────────────────────────────────────────────
Question 14
a) Mean & Standard Deviation:
Midpoints (xβ‚˜): 12.5, 16, 18.5, 20, 22, 26
Ξ£f = 114, Ξ£fxβ‚˜ = 2108.5 ⟹ xΜ… = 2108.5 / 114 β‰ˆ 18.50
Ξ£fxβ‚˜Β² = 41132.25 ⟹ Οƒ = √(41132.25 / 114 – (18.495)Β²) β‰ˆ 4.33
b) Median (Qβ‚‚):
N / 2 = 57 ⟹ Median Class: 17.5 – 19.5
Qβ‚‚ = 17.5 + ((57 – 48) / 19) Γ— 2 = 17.5 + 0.95 β‰ˆ 18.4
c) Histogram: Plot Frequency Density against Class Boundaries (10.5–14.5, 14.5–17.5, 17.5–19.5, 19.5–20.5, 20.5–23.5, 23.5–28.5).
d) Proportion within 3 Standard Deviations (xΜ… Β± 3Οƒ):
Range = 18.5 Β± 3(4.33) = [5.51, 31.49]
All data values fall within 10.5 to 28.5, so 100% of the data lies within 3 standard deviations.
e) Normal Distribution Suitability:
Yes, the distribution is roughly symmetrical with Mean (18.5) β‰ˆ Median (18.4), and 100% of the data lies within 3 standard deviations, which aligns well with Normal distribution properties.
────────────────────────────────────────────────────
Question 15
Coding: y = (x – 3325) / 50 ⟹ x = 50y + 3325
Class Midpoints (x): 3275, 3325, 3375, 3425, 3475
Coded values (y): -1, 0, 1, 2, 3
Frequencies (f): 19, 45, 16, 5, 2
Ξ£f = 87
Ξ£fy = 19(-1) + 45(0) + 16(1) + 5(2) + 2(3) = 13
Ξ£fyΒ² = 19(1) + 45(0) + 16(1) + 5(4) + 2(9) = 73
yΜ… = 13 / 87 β‰ˆ 0.1494
xΜ… = 50(0.1494) + 3325 β‰ˆ 3332.47 β‰ˆ 3332
Οƒα΅§ = √(73 / 87 – (0.1494)Β²) = √(0.8391 – 0.0223) β‰ˆ 0.9038
Οƒβ‚“ = 50 Γ— Οƒα΅§ = 50(0.9038) β‰ˆ 45.19 β‰ˆ 45.2
────────────────────────────────────────────────────
Question 16
Heights between 4.5 and 8.5 cm (width = 4) have frequency = 18.
Area = width Γ— height ⟹ 4 Γ— h = 18 ⟹ frequency density multiplier = 1.
a) Median:
Total Frequency = Area under histogram
– [0, 4.5]: 4.5 Γ— 2 = 9
– [4.5, 8.5]: 4 Γ— 4.5 = 18
– [8.5, 12.5]: 4 Γ— 6 = 24
– [12.5, 14.5]: 2 Γ— 8 = 16
– [14.5, 16.5]: 2 Γ— 3.5 = 7
Total N = 9 + 18 + 24 + 16 + 7 = 74
N / 2 = 37 ⟹ Median lies in class [12.5, 14.5]
Median β‰ˆ 13.6
b) Mean & Standard Deviation:
Class midpoints: 2.25, 6.5, 10.5, 13.5, 15.5
Mean (xΜ…) β‰ˆ 13.48
Standard Deviation (Οƒ) β‰ˆ 3.45
────────────────────────────────────────────────────
Question 17
For class 47–50 (boundaries 46.5 to 50.5):
Width = 4 minutes ⟹ represented by base = 6 cm.
Scale for width = 6 cm / 4 minutes = 1.5 cm per minute.
Area of rectangle = 6 cm Γ— 3.6 cm = 21.6 cmΒ²
Frequency = 48 ⟹ Scale for frequency = 21.6 cm² / 48 = 0.45 cm² per unit frequency.
For class 51–55 (boundaries 50.5 to 55.5):
Width = 5 minutes ⟹ Base = 5 Γ— 1.5 cm = 7.5 cm.
Frequency = 30 ⟹ Area = 30 Γ— 0.45 cmΒ² = 13.5 cmΒ².
Height = Area / Base = 13.5 / 7.5 = 1.8 cm.
Base = 7.5 cm, Height = 1.8 cm
────────────────────────────────────────────────────
Question 18
Coding: y = 50(x – 0.09) ⟹ x = y / 50 + 0.09
Midpoints (x): 0.03, 0.05, 0.07, 0.09, 0.11
Coded values (y): -3, -2, -1, 0, 1
Frequencies (f): 25, 76, 111, 255, 33
a) Mean & Standard Deviation:
Ξ£f = 500
Ξ£fy = 25(-3) + 76(-2) + 111(-1) + 255(0) + 33(1) = -305
Ξ£fyΒ² = 25(9) + 76(4) + 111(1) + 255(0) + 33(1) = 674
yΜ… = -305 / 500 = -0.61
xΜ… = -0.61 / 50 + 0.09 = -0.0122 + 0.09 = 0.0778
Οƒα΅§ = √(674 / 500 – (-0.61)Β²) = √(1.348 – 0.3721) β‰ˆ 0.9879
Οƒβ‚“ = Οƒα΅§ / 50 = 0.9879 / 50 β‰ˆ 0.01978 β‰ˆ 0.0197
b) Median (Qβ‚‚):
N / 2 = 250 ⟹ Median Class: 0.08 < d ≀ 0.10 (Cumulative f before = 212)
Qβ‚‚ = 0.08 + ((250 – 212) / 255) Γ— 0.02 = 0.08 + 0.00298 β‰ˆ 0.08
c) Skewness:
Mean (0.0778) < Median (0.08) ⟹ Negative Skew.
────────────────────────────────────────────────────
Question 19
For class 125 ≀ W < 130:
Width = 5 g ⟹ represented by base = 1.8 cm.
Scale for width = 1.8 cm / 5 g = 0.36 cm per gram.
Area = 1.8 cm Γ— 12 cm = 21.6 cmΒ²
Frequency = 75 ⟹ Scale for area = 21.6 cm² / 75 = 0.288 cm² per unit frequency.
For class 150 ≀ W < 170:
Width = 20 g ⟹ Base = 20 Γ— 0.36 cm = 7.2 cm.
Frequency = 40 ⟹ Area = 40 Γ— 0.288 cmΒ² = 11.52 cmΒ².
Height = Area / Base = 11.52 / 7.2 = 1.6 cm.
Base = 7.2 cm, Height = 1.6 cm
────────────────────────────────────────────────────
Question 20
Let y = x – 50 ⟹ x = y + 50
n = 40
Ξ£y = 140
Ξ£yΒ² = 4490
Mean of y (yΜ…) = 140 / 40 = 3.5
Mean of x (xΜ…) = yΜ… + 50 = 3.5 + 50 = 53.5 kg
Variance of y (Οƒα΅§Β²) = Ξ£yΒ² / n – yΜ…Β² = 4490 / 40 – (3.5)Β² = 112.25 – 12.25 = 100
Standard Deviation (Οƒ) = √100 = 10 kg
────────────────────────────────────────────────────
Question 21
Class 146–150 (Boundaries 145.5 to 150.5):
Class width = 5
Rectangle base = 2.8 cm, height = 7.5 cm ⟹ Area = 2.8 Γ— 7.5 = 21 cmΒ²
Frequency = 75
For any rectangle in a histogram:
Frequency ∝ Area
Area of given class = 2.8 Γ— 7.5 = 21 cmΒ²
Area of new class = 5.6 Γ— 10.5 = 58.8 cmΒ²
Frequency of new class = 75 Γ— (58.8 / 21) = 75 Γ— 2.8 = 210
f = 210
────────────────────────────────────────────────────
Question 22
Let y = (x – 255) / 2 ⟹ x = 2y + 255
n = 5
Ξ£y = 50
Ξ£yΒ² = 1650
Mean of y (yΜ…) = 50 / 5 = 10
Mean of x (xΜ…) = 2(10) + 255 = 275
Variance of y (Οƒα΅§Β²) = Ξ£yΒ² / n – yΜ…Β² = 1650 / 5 – 10Β² = 330 – 100 = 230
Standard deviation of y (Οƒα΅§) = √230
Standard deviation of x (Οƒβ‚“) = 2 Γ— Οƒα΅§ = 2√230 β‰ˆ 30.33
xΜ… = 275, Οƒ β‰ˆ 30.3 (or 2√230)
────────────────────────────────────────────────────
Question 23
Class 120 ≀ h < 130:
Width = 10 cm
Rectangle base = 4.2 cm, height = 9 cm ⟹ Area = 4.2 Γ— 9 = 37.8 cmΒ²
Frequency = 72
New class:
Rectangle base = 2.1 cm, height = 8 cm ⟹ Area = 2.1 Γ— 8 = 16.8 cmΒ²
Frequency = 72 Γ— (16.8 / 37.8) = 72 Γ— (4/9) = 32
f = 32
────────────────────────────────────────────────────
Question 24
Class Midpoints (xβ‚˜) and Boundaries:
– 2–6 (1.5 to 6.5, width 5): Midpoint = 4, f = 6
– 7–11 (6.5 to 11.5, width 5): Midpoint = 9, f = 15
– 12–16 (11.5 to 16.5, width 5): Midpoint = 14, f = k
– 17–31 (16.5 to 31.5, width 15): Midpoint = 24, f = 24
– 32–36 (31.5 to 36.5, width 5): Midpoint = 34, f = 12
Mean xΜ… = Ξ£fxβ‚˜ / Ξ£f = 18.6
Ξ£f = 6 + 15 + k + 24 + 12 = 57 + k
Ξ£fxβ‚˜ = 6(4) + 15(9) + k(14) + 24(24) + 12(34) = 24 + 135 + 14k + 576 + 408 = 1143 + 14k
Equation:
(1143 + 14k) / (57 + k) = 18.6
1143 + 14k = 18.6(57 + k)
1143 + 14k = 1060.2 + 18.6k
82.8 = 4.6k ⟹ k = 18
Total N = 57 + 18 = 75
Ξ£fxβ‚˜ = 1143 + 14(18) = 1395
Ξ£fxβ‚˜Β² = 6(4Β²) + 15(9Β²) + 18(14Β²) + 24(24Β²) + 12(34Β²)
= 96 + 1215 + 3528 + 13824 + 13872 = 32535
Variance (σ²) = Ξ£fxβ‚˜Β² / N – xΜ…Β² = 32535 / 75 – (18.6)Β² = 433.8 – 345.96 = 87.84
Standard Deviation (Οƒ) = √87.84 β‰ˆ 9.37
Οƒ β‰ˆ 9.37
────────────────────────────────────────────────────
Question 25
Class 24–30 (Boundaries 23.5 to 30.5):
Width = 7, Frequency = 63
Base = 2.8 cm, Height = 6 cm
Scale for Base: 2.8 cm / 7 units = 0.4 cm per unit width.
Area = 2.8 Γ— 6 = 16.8 cmΒ²
Scale for Area: 16.8 cmΒ² / 63 = 4/15 cmΒ² per unit frequency.
Class 31–35 (Boundaries 30.5 to 35.5):
Width = 5, Frequency = 60
Base = 5 Γ— 0.4 cm = 2 cm
Area = 60 Γ— (4/15) = 16 cmΒ²
Height = Area / Base = 16 / 2 = 8 cm
Base = 2 cm, Height = 8 cm
────────────────────────────────────────────────────
Question 26
Class Midpoints (xβ‚˜) and Boundaries:
– 3–5 (2.5 to 5.5, width 3): Midpoint = 4, f = 12
– 6–7 (5.5 to 7.5, width 2): Midpoint = 6.5, f = 14
– 8 (7.5 to 8.5, width 1): Midpoint = 8, f = 19
– 9–11 (8.5 to 11.5, width 3): Midpoint = 10, f = 13
– 12–17 (11.5 to 17.5, width 6): Midpoint = 14.5, f = 6
N = 64
Ξ£fxβ‚˜ = 12(4) + 14(6.5) + 19(8) + 13(10) + 6(14.5) = 48 + 91 + 152 + 130 + 87 = 508
Ξ£fxβ‚˜Β² = 12(16) + 14(42.25) + 19(64) + 13(100) + 6(210.25) = 192 + 591.5 + 1216 + 1300 + 1261.5 = 4561
a) Mean (xΜ…) = 508 / 64 β‰ˆ 7.94
Variance (σ²) = 4561 / 64 – (7.9375)Β² = 71.2656 – 63.0049 = 8.2607
Standard Deviation (Οƒ) = √8.2607 β‰ˆ 2.87
b) Median (Qβ‚‚): N / 2 = 32
Cumulative frequencies: 12, 26, 45…
Median class: 7.5 to 8.5 (f = 19, cf before = 26)
Qβ‚‚ = 7.5 + ((32 – 26) / 19) Γ— 1 = 7.5 + 0.316 β‰ˆ 7.82
c) Rectangle for class 3–5: Width = 3 ⟹ Base = 1.2 cm (scale = 0.4 cm/unit)
Area = 1.2 Γ— 5 = 6 cmΒ² ⟹ scale = 6/12 = 0.5 cmΒ² per unit f.
For class 12–17 (width = 6, f = 6):
Base = 6 Γ— 0.4 = 2.4 cm
Area = 6 Γ— 0.5 = 3 cmΒ²
Height = 3 / 2.4 = 1.25 cm
d) Outlier Boundaries:
IQR = Q₃ – Q₁ = 9.19 – 6.07 = 3.12
Upper Boundary = Q₃ + 1.5(IQR) = 9.19 + 1.5(3.12) = 13.87
Lower Boundary = Q₁ – 1.5(IQR) = 6.07 – 1.5(3.12) = 1.39
Since max distance is 17.5 > 13.87, there are potential upper outliers.
e) Skewness: Mean (7.94) > Median (7.82), indicating slight positive skew. Normal distribution requires symmetry, so it may not be ideal.
────────────────────────────────────────────────────
Question 27
Class Midpoints (xβ‚˜) and Widths:
– 1–3 (width 2): Midpoint = 2, f = 15
– 3–5 (width 2): Midpoint = 4, f = 31
– 5–6 (width 1): Midpoint = 5.5, f = 45
– 6–6.5 (width 0.5): Midpoint = 6.25, f = 37
– 6.5–7 (width 0.5): Midpoint = 6.75, f = 21
– 7–10 (width 3): Midpoint = 8.5, f = 15
N = 164
Ξ£fxβ‚˜ = 15(2) + 31(4) + 45(5.5) + 37(6.25) + 21(6.75) + 15(8.5) = 30 + 124 + 247.5 + 231.25 + 141.75 + 127.5 = 902
Ξ£fxβ‚˜Β² = 15(4) + 31(16) + 45(30.25) + 37(39.0625) + 21(45.5625) + 15(72.25) = 60 + 496 + 1361.25 + 1445.3125 + 956.8125 + 1083.75 = 5403.125
a) Mean (xΜ…) = 902 / 164 = 5.5 kg
Variance = 5403.125 / 164 – 5.5Β² = 32.94588 – 30.25 = 2.69588
Standard Deviation (Οƒ) = √2.69588 β‰ˆ 1.64 kg
b) Median (Qβ‚‚): N / 2 = 82
Cumulative frequencies: 15, 46, 91…
Median class: 5 to 6 (f = 45, cf before = 46)
Qβ‚‚ = 5 + ((82 – 46) / 45) Γ— 1 = 5 + 36/45 = 5.80 kg
Since Mean (5.5) < Median (5.8), the distribution is negatively skewed.
c) Class 1 ≀ w < 3: Width = 2, Base = 2.4 cm ⟹ scale = 1.2 cm per unit width.
Area = 2.4 Γ— 2.5 = 6 cmΒ² ⟹ scale = 6 / 15 = 0.4 cmΒ² per unit f.
For class 6.5 ≀ w < 7: Width = 0.5, f = 21
Base = 0.5 Γ— 1.2 = 0.6 cm
Area = 21 Γ— 0.4 = 8.4 cmΒ²
Height = 8.4 / 0.6 = 14 cm
d) Outlier limits:
IQR = 6.43 – 4.68 = 1.75
Lower Limit = 4.68 – 1.5(1.75) = 2.055
Upper Limit = 6.43 + 1.5(1.75) = 9.055
Values below 2.055 or above 9.055 are potential outliers.
e) Due to negative skewness and presence of potential outliers, a Normal distribution is not appropriate.
────────────────────────────────────────────────────
Question 28
Coding: y = (x – 662.5) / 25 ⟹ x = 25y + 662.5
Classes, Midpoints (x), Coded values (y), and Frequencies (f):
– 600–625: x = 612.5, y = -2, f = 11
– 625–650: x = 637.5, y = -1, f = 14
– 650–675: x = 662.5, y = 0, f = 28
– 675–700: x = 687.5, y = 1, f = 7
– 700–725: x = 712.5, y = 2, f = 5
– 725–750: x = 737.5, y = 3, f = 2
– 750–775: x = 762.5, y = 4, f = 1
N = 68
Ξ£fy = 11(-2) + 14(-1) + 28(0) + 7(1) + 5(2) + 2(3) + 1(4) = -22 – 14 + 0 + 7 + 10 + 6 + 4 = -9
Ξ£fyΒ² = 11(4) + 14(1) + 28(0) + 7(1) + 5(4) + 2(9) + 1(16) = 44 + 14 + 0 + 7 + 20 + 18 + 16 = 119
a) yΜ… = -9 / 68 β‰ˆ -0.13235
xΜ… = 25(-0.13235) + 662.5 β‰ˆ 659.19
Οƒα΅§Β² = 119 / 68 – (-0.13235)Β² = 1.75 – 0.01751 = 1.73249
Οƒα΅§ = √1.73249 β‰ˆ 1.3162
Οƒβ‚“ = 25 Γ— 1.3162 β‰ˆ 32.91
b) Median (Qβ‚‚): N / 2 = 34
Cumulative frequencies: 11, 25, 53…
Median class: 650 to 675 (f = 28, cf before = 25)
Qβ‚‚ = 650 + ((34 – 25) / 28) Γ— 25 = 650 + 8.036 = 658.0
────────────────────────────────────────────────────
Question 29
a) From histogram frequency densities:
Students scoring between 52 and 74 = 60
b) Total frequency N = 250
Median position = 125th student
Median (Qβ‚‚) β‰ˆ 49
c) Estimated Mean (xΜ…) β‰ˆ 51.8
Estimated Standard Deviation (Οƒ) β‰ˆ 22.22
Question 30
Group 1: n₁ = 20, x̅₁ = 18.5, σ₁ = 6.5
Group 2: nβ‚‚ = 12, xΜ…β‚‚ = 25, Οƒβ‚‚ = 7.5
Combined N = 32
Combined Mean (xΜ…):
xΜ… = (n₁x̅₁ + nβ‚‚xΜ…β‚‚) / N = (20 Γ— 18.5 + 12 Γ— 25) / 32 = (370 + 300) / 32 = 670 / 32 β‰ˆ 20.94
Combined Variance (σ²):
Ξ£x₁² = n₁(σ₁² + x̅₁²) = 20(6.5Β² + 18.5Β²) = 20(42.25 + 342.25) = 20(384.5) = 7690
Ξ£xβ‚‚Β² = nβ‚‚(Οƒβ‚‚Β² + xΜ…β‚‚Β²) = 12(7.5Β² + 25Β²) = 12(56.25 + 625) = 12(681.25) = 8175
Combined Ξ£xΒ² = 7690 + 8175 = 15865
Combined Variance = 15865 / 32 – (20.9375)Β² = 495.78125 – 438.3789 = 57.40235
Combined Standard Deviation = √57.40235 β‰ˆ 7.58
────────────────────────────────────────────────────
Question 31
Total: N = 40, xΜ… = 65, Οƒ = 18
Boys: n₁ = 24, x̅₁ = 72, σ₁ = 20
Girls: nβ‚‚ = 16, xΜ…β‚‚ = ?, Οƒβ‚‚ = ?
Mean for girls (xΜ…β‚‚):
n₁x̅₁ + nβ‚‚xΜ…β‚‚ = N xΜ…
24(72) + 16(xΜ…β‚‚) = 40(65)
1728 + 16xΜ…β‚‚ = 2600
16xΜ…β‚‚ = 872 ⟹ xΜ…β‚‚ = 54.5
Standard deviation for girls (Οƒβ‚‚):
Total Ξ£xΒ² = 40(18Β² + 65Β²) = 40(324 + 4225) = 40(4549) = 181960
Boys Ξ£x₁² = 24(20Β² + 72Β²) = 24(400 + 5184) = 24(5584) = 134016
Girls Ξ£xβ‚‚Β² = 181960 – 134016 = 47944
Girls Variance (Οƒβ‚‚Β²) = 47944 / 16 – (54.5)Β² = 2996.5 – 2970.25 = 26.25
Girls Standard Deviation (Οƒβ‚‚) = √26.25 β‰ˆ 5.12
────────────────────────────────────────────────────
Question 32
Given:
n = 40
Ξ£(x – 50) = 150
Ξ£(x – 50)Β² = 4650
Expand Ξ£(x – 50)Β²:
Ξ£(xΒ² – 100x + 2500) = 4650
Ξ£xΒ² – 100Ξ£x + 2500(40) = 4650
Ξ£xΒ² – 100Ξ£x + 100000 = 4650 — (Equation 1)
From Ξ£(x – 50) = 150:
Ξ£x – 50(40) = 150
Ξ£x – 2000 = 150 ⟹ Ξ£x = 2150
Substitute Ξ£x into Equation 1:
Ξ£xΒ² – 100(2150) + 100000 = 4650
Ξ£xΒ² – 215000 + 100000 = 4650
Ξ£xΒ² – 115000 = 4650
Ξ£xΒ² = 119650
────────────────────────────────────────────────────
Question 33
Given:
xΜ… = (1/n) Ξ£xα΅£ ⟹ Ξ£xα΅£ = n xΜ…
Variance σ² = (1/n) Ξ£xα΅£Β² – xΜ…Β² = 3 ⟹ Ξ£xα΅£Β² = n(3 + xΜ…Β²)
We want to find Ξ£(xα΅£ + 1)Β²:
Ξ£(xα΅£ + 1)Β² = Ξ£(xα΅£Β² + 2xα΅£ + 1)
= Ξ£xα΅£Β² + 2Ξ£xα΅£ + Ξ£1
= n(3 + xΜ…Β²) + 2(n xΜ…) + n
= 3n + n xΜ…Β² + 2n xΜ… + n
= n xΜ…Β² + 2n xΜ… + 4n
= n(xΜ…Β² + 2xΜ… + 4)
Given xΜ… = 3:
Ξ£(xα΅£ + 1)Β² = n(3Β² + 2(3) + 4) = n(9 + 6 + 4) = 19n [or n(1 + 18) = 19n]
────────────────────────────────────────────────────
Question 34
Given:
n = 20
Ξ£(x – 10) = 220
Ξ£(x – 10)Β² = 2720
a) Show Ξ£xΒ² = 9120:
Ξ£(x – 10) = 220 ⟹ Ξ£x – 200 = 220 ⟹ Ξ£x = 420
Expand Ξ£(x – 10)Β²:
Ξ£(xΒ² – 20x + 100) = 2720
Ξ£xΒ² – 20Ξ£x + 2000 = 2720
Ξ£xΒ² – 20(420) + 2000 = 2720
Ξ£xΒ² – 8400 + 2000 = 2720
Ξ£xΒ² – 6400 = 2720
Ξ£xΒ² = 9120 (Proved)
b) Mean (xΜ…) & Standard Deviation (Οƒ):
xΜ… = Ξ£x / n = 420 / 20 = 21
σ² = Ξ£xΒ² / n – xΜ…Β² = 9120 / 20 – 21Β² = 456 – 441 = 15
Οƒ = √15 β‰ˆ 3.87
────────────────────────────────────────────────────
Q1: Which one of the following data types could best be described as Personally Identifiable Information (PII)?
Answer: C Shipping addresses for customers’ most recent orders.
Q2: You are currently gathering data relating to coastal erosion, in order to predict the measurement of the erosion in another 20 years’ time. It includes a series of measurements that have been recorded over a period of 50 years. Which one of the following types of data would you most likely use in your analysis?
Answer: A Continuous.
Q3: Which one of the following data structures includes the use of a parent node?
Answer: C Tree.
Q4: You want to gain insight into the influence your customers have on brand visibility. You have a structured data source in the form of a customer relationship management (CRM) system, as well as unstructured data from social media feeds. What would be the main benefit of using this unstructured data alongside the structured data source?
Answer: B You could identify how many times your brand’s name is mentioned.
Q5: You have been tasked to produce a monthly report on sales from the previous month. What sort of analytics would you use?
Answer: B Descriptive analytics.
Q6: You want to create a data model that describes the technical requirements of a data analysis project. The intended audience will be non-technical company directors. Select the most appropriate data model from the following options.
Answer: A Conceptual data model.
Q7: Which one of the following should define what data is collected and stored in an organisation?
Answer: A Architectural policies.
Q8: Which one of the following data architecture functions would support business intelligence activities on historical data?
Answer: C Data warehousing.
Q9: The rate at which data is generated, collected, processed and analysed describes which challenge associated with big data?
Answer: B Velocity.
Q10: When talking about big data, which one of the following could you reasonably expect to see in your original datasets?
Answer: C Unstructured content.
Q11: Which one of following would you expect to feature in an entity relationship diagram (ERD)?
Answer: A Attributes.
Q12: Which one of the following is an advantage of a relational database over a NoSQL database?
Answer: B Speed of transactions for low volumes of data.
Q13: Which stage of ETL would cleanse data?
Answer: B Transform.
Q14: Which one of the following operations would you expect to happen in the Extract phase of ETL?
Answer: D Staging data from source systems.
Q15: Which one of the following has a key purpose of being able to create “a highly intuitive, drag-and-drop interface for building visualisations, reports, and dashboards”?
Answer: C PowerBI.
Q16: Which one of the following is a drawback when using the mean average of a set of data?
Answer: A Extreme values can heavily influence the result.
Q17: Which one of the following is a drawback when using the mode average of a set of data?
Answer: B It may not provide a single value as the answer.
Q18: Which one of the following is a drawback when using the median average of a set of data?
Answer: A The result may not appear in the original data set.
Q19: When considering strategies to improve data analysis modelling, which one of the following methods is most effective?
Answer: D Adding more data to the training set.
Q20: When visualising data for stakeholders, which one of the following is the key factor to consider?
Answer: B The accessibility of the presented data and ease of understanding.
Scenario 1: Database design and SQL
Q21: More information is required on employees to help escalate issues with IT asset conditions. It has been decided to add another table storing information on an employee’s manager. What new field would be a suitable primary key for the β€˜Managers’ table?
Answer: A employee_id.
Q22: You have been tasked with creating a data dictionary for this database. Which one of the following should be listed in a data dictionary?
Answer: D field size.
Q23: A key report measure on the scorecard is the number of assets returned in a month, and their condition when returned. Which two of the following aggregations should be used to query the ‘Employee_Assets’ table for this measure?
Answer: A Group By & B Count.
Q24: If an asset_id is removed from the β€˜IT_assets’ table, all related information for that asset should also be removed from the other tables. What database functionality ensures that this happens?
Answer: C Cascading delete.
Q25: Fill in the blanks to complete the SQL query shown below:
SELECT ___________(employee_id), department from ___________where first_name = "" group by ___________ order by ___________
Answer: SELECT COUNT(employee_id), department from EMPLOYEES where first_name = "" group by department order by 2
Scenario 2: Data Preparation and Integration
Q26: Order the following Python commands into a logical flow for importing the data contained in the file called “IToutages.csv”. After importing the entire file, you should print the first ten characters of data to screen.
Answer:
  1. f = open("IToutages.csv")
  2. data=f.read()
  3. FirstTenChars = data[:10]
  4. print(FirstTenChars)
  5. f.close
Q27: Which one of the following is the missing line of code in the Python programme below to find the mean of 2,3,4,5,6?
Answer: C Mean=Total/5.
Q28: Which one of the following R commands will correctly show different averages and quartiles of a dataset?
Answer: D summary(dataset)
Q29: Which one of the following R commands would be used to see a snapshot of the first six rows in a dataset?
Answer: A head(dataset)
Q30: Order the following lines of code into a logical flow to read in the ‘ITemployees.txt’ file and plot the data as an annual time series, then change the timeseries to be monthly and plot a second graph that starts in 1999.
Answer:
  1. EmployeesByYear <- read.csv("ITemployees.txt")
  2. employeetimeseries <- ts(EmployeesByYear)
  3. plot.ts(employeetimeseries)
  4. employeetimeseries <- ts(EmployeesByYear, frequency = 12, start = 1999)
  5. monthplot(employeetimeseries)
Scenario 3: Normalization and Analytics
Q31: Which two of the following outcomes would you expect when normalising this data to first normal form?
Answer: B The Salesperson information would be repeated, creating 12 separate rows of data & D Two or more separate tables would be created.
Q32: Which three of the following tables could be created in a normalised form of the data?
Answer: B Purchase table, C Customer table & E Salesperson table.
Q33: If you created a table storing information about salespeople, assuming the salesperson number is unique, what type of key would this make?
Answer: B Primary key.
Q34: When producing models as part of your database design, creating the normalised form of this data would be categorised as which form of model?
Answer: C Conceptual model.
Q35: Which one of the following statements would be an appropriate null hypothesis?
Answer: B The purchase date has no impact on the number of sales.
Scenario 4: Data Modelling
Q36: If the first column (col1) contained “1” and the second column (col2) contained “0”, what would be returned from col1 AND col2?
Answer: A 0.
Q37: Which one of the following statements would form a suitable null hypothesis for this model?
Answer: C H0 – the amount of daily sunshine does not impact the total daily sales.
Q38: What would be an appropriate size subset of your data to use for this training?
Answer: D 70%.
Q39: Your model shows a correlation coefficient of 0.4. How would you interpret this coefficient result?
Answer: D Sales are moderately correlated to the number of people in the town centre.
Q40: Which visualisation would be most appropriate for a linear regression forecast?
Answer: D Scatter chart

Leave a Reply

Your email address will not be published. Required fields are marked *