京都大学 情報学研究科 知能情報学専攻 2024年8月実施 専門科目 S-1
Author
itsuitsuki
Description
大学公表の原題
Q.1
Suppose the probability density function f ( x ) f(x) f ( x ) of a random variable X X X is as follows.
f ( x ) = { 0 ( x < 0 ) c x ( 3 − x ) ( 0 ≤ x < 3 ) 0 ( 3 ≤ x ) f(x) = \begin{cases} 0 & (x < 0) \\ cx(3-x) & (0 \leq x < 3) \\ 0 & (3 \leq x) \end{cases} f ( x ) = ⎩ ⎨ ⎧ 0 c x ( 3 − x ) 0 ( x < 0 ) ( 0 ≤ x < 3 ) ( 3 ≤ x )
c c c is a positive constant (c > 0 c > 0 c > 0 ).
(1) Compute the value of the constant c c c .
(2) Compute the mean and variance of the random variable X X X .
Q.2
Let X X X and Y Y Y be independent random variables following the binomial distributions B ( m , p ) \mathrm{B}(m, p) B ( m , p ) and B ( n , p ) \mathrm{B}(n, p) B ( n , p ) , respectively. Derive the distribution of Z = X + Y Z = X + Y Z = X + Y .
Q.3
Consider normal populations A and B with a common population variance. One sample of size 18 is selected from the normal population A, denoted as ( x 1 , x 2 , … , x 18 ) (x_1, x_2, \dots, x_{18}) ( x 1 , x 2 , … , x 18 ) , and the other one of size 18 from the normal population B, denoted as ( y 1 , y 2 , … , y 18 ) (y_1, y_2, \dots, y_{18}) ( y 1 , y 2 , … , y 18 ) . The statistics derived from the samples are as follows.
x ˉ = 1 18 ∑ i = 1 18 x i s x 2 = 1 17 ∑ i = 1 18 ( x i − x ˉ ) 2 y ˉ = 1 18 ∑ i = 1 18 y i s y 2 = 1 17 ∑ i = 1 18 ( y i − y ˉ ) 2 s x y = 1 17 ∑ i = 1 18 ( x i − x ˉ ) ( y i − y ˉ ) \bar{x} = \frac{1}{18} \sum_{i=1}^{18} x_i \qquad s_x^2 = \frac{1}{17} \sum_{i=1}^{18} (x_i - \bar{x})^2 \\
\bar{y} = \frac{1}{18} \sum_{i=1}^{18} y_i \qquad s_y^2 = \frac{1}{17} \sum_{i=1}^{18} (y_i - \bar{y})^2 \qquad s_{xy} = \frac{1}{17} \sum_{i=1}^{18} (x_i - \bar{x})(y_i - \bar{y}) x ˉ = 18 1 i = 1 ∑ 18 x i s x 2 = 17 1 i = 1 ∑ 18 ( x i − x ˉ ) 2 y ˉ = 18 1 i = 1 ∑ 18 y i s y 2 = 17 1 i = 1 ∑ 18 ( y i − y ˉ ) 2 s x y = 17 1 i = 1 ∑ 18 ( x i − x ˉ ) ( y i − y ˉ )
The values 2.110 and 2.032 may be used for the upper 2.5% point of the t t t -distribution with 17 and 34 degrees of freedom, respectively.
(1) Assuming that the samples are paired as ( x i , y i ) (x_i, y_i) ( x i , y i ) , i = 1 , 2 , … , 18 i = 1, 2, \dots, 18 i = 1 , 2 , … , 18 and randomly selected, compute the 95% confidence interval for the difference between the two population means using necessary statistics among those mentioned above.
(2) Assuming that the samples are unpaired and randomly selected from each population, compute the 95% confidence interval for the difference between the two population means using necessary statistics among those mentioned above.
Q.4
Consider a linear regression model Y i = α + β x i + ϵ i ( i = 1 , 2 , … , 16 ) Y_i = \alpha + \beta x_i + \epsilon_i \ (i = 1, 2, \dots, 16) Y i = α + β x i + ϵ i ( i = 1 , 2 , … , 16 ) where the variation in random variables Y i ( i = 1 , 2 , … , 16 ) Y_i \ (i = 1, 2, \dots, 16) Y i ( i = 1 , 2 , … , 16 ) are explained by the corresponding constants x i ( i = 1 , 2 , … , 16 ) x_i \ (i = 1, 2, \dots, 16) x i ( i = 1 , 2 , … , 16 ) with the regression coefficients of α \alpha α and β \beta β . Assume that ϵ i ( i = 1 , 2 , … , 16 ) \epsilon_i \ (i = 1, 2, \dots, 16) ϵ i ( i = 1 , 2 , … , 16 ) are independent and follow a normal distribution N ( 0 , σ 2 ) \mathrm{N}(0, \sigma^2) N ( 0 , σ 2 ) with mean 0 and variance σ 2 \sigma^2 σ 2 . Let α ^ \hat{\alpha} α ^ and β ^ \hat{\beta} β ^ be the least squares estimators of α \alpha α and β \beta β , respectively.
(1) Compute the standard deviation of β ^ \hat{\beta} β ^ .
(2) Given a new constant x 17 x_{17} x 17 , show the 95% prediction interval of Y 17 = α + β x 17 + ϵ 17 Y_{17} = \alpha + \beta x_{17} + \epsilon_{17} Y 17 = α + β x 17 + ϵ 17 using α ^ , β ^ , σ \hat{\alpha}, \hat{\beta}, \sigma α ^ , β ^ , σ , and x 1 , x 2 , … , x 17 x_1, x_2, \dots, x_{17} x 1 , x 2 , … , x 17 where ϵ 17 \epsilon_{17} ϵ 17 is independent of ϵ i ( i = 1 , 2 , … , 16 ) \epsilon_i \ (i = 1, 2, \dots, 16) ϵ i ( i = 1 , 2 , … , 16 ) and follows a normal distribution N ( 0 , σ 2 ) \mathrm{N}(0, \sigma^2) N ( 0 , σ 2 ) . The value 2.145 for the upper 2.5% point of the t t t -distribution with 14 degrees of freedom may be used.
题目描述
随机变量 X X X 的概率密度函数为
f ( x ) = { 0 ( x < 0 ) , c x ( 3 − x ) ( 0 ≤ x < 3 ) , 0 ( x ≥ 3 ) , f(x)=
\begin{cases}
0 & (x<0),\\
cx(3-x) & (0\leq x<3),\\
0 & (x\geq 3),
\end{cases} f ( x ) = ⎩ ⎨ ⎧ 0 c x ( 3 − x ) 0 ( x < 0 ) , ( 0 ≤ x < 3 ) , ( x ≥ 3 ) ,
其中 c > 0 c>0 c > 0 。(1)求常数 c c c ;(2)求 X X X 的均值与方差。
设 X X X 与 Y Y Y 相互独立,且分别服从二项分布 B ( m , p ) \mathrm{B}(m,p) B ( m , p ) 与 B ( n , p ) \mathrm{B}(n,p) B ( n , p ) 。推导 Z = X + Y Z=X+Y Z = X + Y 的概率分布。
从方差相同的两个正态总体 A、B 中分别随机抽取容量为 18 的样本 x 1 , … , x 18 x_1,\ldots,x_{18} x 1 , … , x 18 与 y 1 , … , y 18 y_1,\ldots,y_{18} y 1 , … , y 18 。定义
x ˉ = 1 18 ∑ i = 1 18 x i , s x 2 = 1 17 ∑ i = 1 18 ( x i − x ˉ ) 2 , \bar{x}=\frac{1}{18}\sum_{i=1}^{18}x_i,\qquad
s_x^2=\frac{1}{17}\sum_{i=1}^{18}(x_i-\bar{x})^2, x ˉ = 18 1 i = 1 ∑ 18 x i , s x 2 = 17 1 i = 1 ∑ 18 ( x i − x ˉ ) 2 ,
y ˉ = 1 18 ∑ i = 1 18 y i , s y 2 = 1 17 ∑ i = 1 18 ( y i − y ˉ ) 2 , \bar{y}=\frac{1}{18}\sum_{i=1}^{18}y_i,\qquad
s_y^2=\frac{1}{17}\sum_{i=1}^{18}(y_i-\bar{y})^2, y ˉ = 18 1 i = 1 ∑ 18 y i , s y 2 = 17 1 i = 1 ∑ 18 ( y i − y ˉ ) 2 ,
以及
s x y = 1 17 ∑ i = 1 18 ( x i − x ˉ ) ( y i − y ˉ ) . s_{xy}=\frac{1}{17}\sum_{i=1}^{18}(x_i-\bar{x})(y_i-\bar{y}). s x y = 17 1 i = 1 ∑ 18 ( x i − x ˉ ) ( y i − y ˉ ) .
自由度为 17 和 34 的 t t t 分布上侧 2.5 % 2.5\% 2.5% 分位点可分别取 2.110 和 2.032。(1)若两组样本按 ( x i , y i ) (x_i,y_i) ( x i , y i ) 配对且为随机抽样,请使用上述必要统计量求两个总体均值之差的 95 % 95\% 95% 置信区间;(2)若两组样本不配对、分别从两个总体随机抽取,请求该均值差的 95 % 95\% 95% 置信区间。
考虑线性回归模型
Y i = α + β x i + ϵ i ( i = 1 , 2 , … , 16 ) , Y_i=\alpha+\beta x_i+\epsilon_i\qquad(i=1,2,\ldots,16), Y i = α + β x i + ϵ i ( i = 1 , 2 , … , 16 ) ,
其中 x i x_i x i 为常数,ϵ i \epsilon_i ϵ i 相互独立且服从 N ( 0 , σ 2 ) \mathrm{N}(0,\sigma^2) N ( 0 , σ 2 ) ,α ^ \hat{\alpha} α ^ 与 β ^ \hat{\beta} β ^ 分别为 α \alpha α 与 β \beta β 的最小二乘估计量。(1)求 β ^ \hat{\beta} β ^ 的标准差;(2)给定新的常数 x 17 x_{17} x 17 ,设
Y 17 = α + β x 17 + ϵ 17 , Y_{17}=\alpha+\beta x_{17}+\epsilon_{17}, Y 17 = α + β x 17 + ϵ 17 ,
其中 ϵ 17 \epsilon_{17} ϵ 17 与前 16 个误差项独立且同样服从 N ( 0 , σ 2 ) \mathrm{N}(0,\sigma^2) N ( 0 , σ 2 ) 。使用 α ^ , β ^ , σ \hat{\alpha},\hat{\beta},\sigma α ^ , β ^ , σ 以及 x 1 , … , x 17 x_1,\ldots,x_{17} x 1 , … , x 17 写出 Y 17 Y_{17} Y 17 的 95 % 95\% 95% 预测区间。自由度为 14 的 t t t 分布上侧 2.5 % 2.5\% 2.5% 分位点可取 2.145。
Kai
Q.1
(1)
Normalization gives
1 = c ∫ 0 3 x ( 3 − x ) d x = 9 2 c , c = 2 9 . 1=c\int_0^3x(3-x)\,dx=\frac92c,
\qquad c=\frac29. 1 = c ∫ 0 3 x ( 3 − x ) d x = 2 9 c , c = 9 2 .
(2)
E [ X ] = 2 9 ∫ 0 3 x 2 ( 3 − x ) d x = 3 2 , E [ X 2 ] = 2 9 ∫ 0 3 x 3 ( 3 − x ) d x = 27 10 . E[X]=\frac29\int_0^3x^2(3-x)\,dx=\frac32,
\qquad
E[X^2]=\frac29\int_0^3x^3(3-x)\,dx=\frac{27}{10}. E [ X ] = 9 2 ∫ 0 3 x 2 ( 3 − x ) d x = 2 3 , E [ X 2 ] = 9 2 ∫ 0 3 x 3 ( 3 − x ) d x = 10 27 .
Hence Var ( X ) = 27 / 10 − 9 / 4 = 9 / 20 \operatorname{Var}(X)=27/10-9/4=9/20 Var ( X ) = 27/10 − 9/4 = 9/20 .
Q.2
Independence gives the probability-generating function
E [ t X + Y ] = E [ t X ] E [ t Y ] = ( 1 − p + p t ) m ( 1 − p + p t ) n = ( 1 − p + p t ) m + n . E[t^{X+Y}]=E[t^X]E[t^Y]
=(1-p+pt)^m(1-p+pt)^n=(1-p+pt)^{m+n}. E [ t X + Y ] = E [ t X ] E [ t Y ] = ( 1 − p + pt ) m ( 1 − p + pt ) n = ( 1 − p + pt ) m + n .
Consequently Z ∼ B ( m + n , p ) Z\sim\mathrm B(m+n,p) Z ∼ B ( m + n , p ) , including p = 0 , 1 p=0,1 p = 0 , 1 .
Q.3
(1)
Let d i = x i − y i d_i=x_i-y_i d i = x i − y i . Then
d ˉ = x ˉ − y ˉ , s d 2 = s x 2 + s y 2 − 2 s x y . \bar d=\bar x-\bar y,\qquad s_d^2=s_x^2+s_y^2-2s_{xy}. d ˉ = x ˉ − y ˉ , s d 2 = s x 2 + s y 2 − 2 s x y .
For independent, normally distributed pair differences with positive variance, the Student pivot has 17 17 17 degrees of freedom. The interval is
x ˉ − y ˉ ± 2.110 s x 2 + s y 2 − 2 s x y 18 . \boxed{\bar x-\bar y\ \pm\ 2.110\sqrt{\frac{s_x^2+s_y^2-2s_{xy}}{18}}}. x ˉ − y ˉ ± 2.110 18 s x 2 + s y 2 − 2 s x y .
Joint normality of each pair suffices for normality of the differences. Normal marginal distributions alone do not imply this condition; without it the displayed Student interval is an approximation rather than an exact finite-sample interval.
(2)
For independent unpaired samples, the pooled estimator is
s p 2 = 17 s x 2 + 17 s y 2 34 = s x 2 + s y 2 2 . s_p^2=\frac{17s_x^2+17s_y^2}{34}=\frac{s_x^2+s_y^2}{2}. s p 2 = 34 17 s x 2 + 17 s y 2 = 2 s x 2 + s y 2 .
The Student pivot has 34 34 34 degrees of freedom, giving
x ˉ − y ˉ ± 2.032 s x 2 + s y 2 18 . \boxed{\bar x-\bar y\ \pm\ 2.032\sqrt{\frac{s_x^2+s_y^2}{18}}}. x ˉ − y ˉ ± 2.032 18 s x 2 + s y 2 .
Q.4
Write x ˉ = ∑ i = 1 16 x i / 16 \bar x=\sum_{i=1}^{16}x_i/16 x ˉ = ∑ i = 1 16 x i /16 and S x x = ∑ i = 1 16 ( x i − x ˉ ) 2 > 0 S_{xx}=\sum_{i=1}^{16}(x_i-\bar x)^2>0 S xx = ∑ i = 1 16 ( x i − x ˉ ) 2 > 0 .
(1)
Since
β ^ − β = ∑ i ( x i − x ˉ ) ϵ i S x x , \hat\beta-\beta=\frac{\sum_i(x_i-\bar x)\epsilon_i}{S_{xx}}, β ^ − β = S xx ∑ i ( x i − x ˉ ) ϵ i ,
independence yields Var ( β ^ ) = σ 2 / S x x \operatorname{Var}(\hat\beta)=\sigma^2/S_{xx} Var ( β ^ ) = σ 2 / S xx , so its standard deviation is σ / S x x \sigma/\sqrt{S_{xx}} σ / S xx .
(2)
The fitted value is Y ^ 17 = α ^ + β ^ x 17 \hat Y_{17}=\hat\alpha+\hat\beta x_{17} Y ^ 17 = α ^ + β ^ x 17 . Using Cov ( ϵ ˉ , β ^ ) = 0 \operatorname{Cov}(\bar\epsilon,\hat\beta)=0 Cov ( ϵ ˉ , β ^ ) = 0 and independence of the new error,
Var ( Y 17 − Y ^ 17 ) = σ 2 ( 1 + 1 16 + ( x 17 − x ˉ ) 2 S x x ) . \operatorname{Var}(Y_{17}-\hat Y_{17})
=\sigma^2\left(1+\frac1{16}+\frac{(x_{17}-\bar x)^2}{S_{xx}}\right). Var ( Y 17 − Y ^ 17 ) = σ 2 ( 1 + 16 1 + S xx ( x 17 − x ˉ ) 2 ) .
With the population standard deviation σ \sigma σ specified, the prediction interval is
α ^ + β ^ x 17 ± 1.960 σ 1 + 1 16 + ( x 17 − x ˉ ) 2 S x x . \boxed{\hat\alpha+\hat\beta x_{17}\ \pm\ 1.960\,\sigma
\sqrt{1+\frac1{16}+\frac{(x_{17}-\bar x)^2}{S_{xx}}}}. α ^ + β ^ x 17 ± 1.960 σ 1 + 16 1 + S xx ( x 17 − x ˉ ) 2 .
If σ \sigma σ is instead estimated from residuals, set
s 2 = 1 14 ∑ i = 1 16 ( Y i − α ^ − β ^ x i ) 2 . s^2=\frac1{14}\sum_{i=1}^{16}(Y_i-\hat\alpha-\hat\beta x_i)^2. s 2 = 14 1 i = 1 ∑ 16 ( Y i − α ^ − β ^ x i ) 2 .
The independent residual variance estimate gives the Student version
α ^ + β ^ x 17 ± 2.145 s 1 + 1 16 + ( x 17 − x ˉ ) 2 S x x . \hat\alpha+\hat\beta x_{17}\ \pm\ 2.145\,s
\sqrt{1+\frac1{16}+\frac{(x_{17}-\bar x)^2}{S_{xx}}}. α ^ + β ^ x 17 ± 2.145 s 1 + 16 1 + S xx ( x 17 − x ˉ ) 2 .