東京大学 情報理工学研究科 2023年8月実施 数学 第1問
Author
zephyr , 祭音Myyura (assisted by ChatGPT 5.4 Thinking)
Description
Let R 3 \mathbb{R}^3 R 3 be the set of the three-dimensional real column vectors and R 3 × 3 \mathbb{R}^{3 \times 3} R 3 × 3 be the set of the three-by-three real matrices. Let n 1 \mathbf{n}_1 n 1 , n 2 \mathbf{n}_2 n 2 , and n 3 ∈ R 3 \mathbf{n}_3 \in \mathbb{R}^3 n 3 ∈ R 3 be linearly independent unit-length vectors and n 4 ∈ R 3 \mathbf{n}_4 \in \mathbb{R}^3 n 4 ∈ R 3 be a unit-length vector not parallel to n 1 \mathbf{n}_1 n 1 , n 2 \mathbf{n}_2 n 2 , or n 3 \mathbf{n}_3 n 3 . Let A \mathbf{A} A and B \mathbf{B} B be square matrices defined as
A = ( n 1 T − n 2 T n 2 T − n 3 T n 3 T − n 4 T ) , B = ∑ i = 1 4 n i n i T . \mathbf{A} = \begin{pmatrix}
\mathbf{n}_1^\mathrm{T} - \mathbf{n}_2^\mathrm{T} \\
\mathbf{n}_2^\mathrm{T} - \mathbf{n}_3^\mathrm{T} \\
\mathbf{n}_3^\mathrm{T} - \mathbf{n}_4^\mathrm{T}
\end{pmatrix},
\quad
\mathbf{B} = \sum_{i=1}^{4} \mathbf{n}_i \mathbf{n}_i^\mathrm{T}. A = n 1 T − n 2 T n 2 T − n 3 T n 3 T − n 4 T , B = i = 1 ∑ 4 n i n i T .
Here, X T \mathbf{X}^\mathrm{T} X T and x T \mathbf{x}^\mathrm{T} x T denote the transpose of a matrix X \mathbf{X} X and a vector x \mathbf{x} x , respectively. Answer the following questions.
(1) Find the condition for n 4 \mathbf{n}_4 n 4 such that the rank of A \mathbf{A} A is three.
(2) In the three-dimensional Euclidean space R 3 \mathbb{R}^3 R 3 , consider four planes Π i = { x ∈ R 3 ∣ n i T x − d i = 0 } \Pi_i = \{\mathbf{x} \in \mathbb{R}^3 \mid \mathbf{n}_i^\mathrm{T} \mathbf{x} - d_i = 0\} Π i = { x ∈ R 3 ∣ n i T x − d i = 0 } ( d i (d_i ( d i is a real number, and i = 1 , 2 , 3 , 4 ) i = 1, 2, 3, 4) i = 1 , 2 , 3 , 4 ) that satisfy the following three conditions: (i) the rank of A \mathbf{A} A is three, (ii) Ω = { x ∈ R 3 ∣ n i T x − d i ≥ 0 , i = 1 , 2 , 3 , 4 } \Omega = \{\mathbf{x} \in \mathbb{R}^3 \mid \mathbf{n}_i^\mathrm{T} \mathbf{x} - d_i \geq 0, \, i = 1, 2, 3, 4\} Ω = { x ∈ R 3 ∣ n i T x − d i ≥ 0 , i = 1 , 2 , 3 , 4 } is not the empty set, and (iii) there exists a sphere C ( C ⊂ Ω ) \mathbf{C} (\mathbf{C} \subset \Omega) C ( C ⊂ Ω ) to which Π i \Pi_i Π i ( i = 1 , 2 , 3 , 4 ) (i = 1, 2, 3, 4) ( i = 1 , 2 , 3 , 4 ) are tangent. The position vector of the center of C \mathbf{C} C is represented by A − 1 u \mathbf{A}^{-1} \mathbf{u} A − 1 u using a vector u ∈ R 3 \mathbf{u} \in \mathbb{R}^3 u ∈ R 3 . Express u \mathbf{u} u using d i ( i = 1 , 2 , 3 , 4 ) d_i \, (i = 1, 2, 3, 4) d i ( i = 1 , 2 , 3 , 4 ) .
(3) Show that B \mathbf{B} B is a positive definite symmetric matrix.
(4) Consider the point P \mathbf{P} P from which the sum of squared distances to four planes { x ∈ R 3 ∣ n i T x − d i = 0 } \{\mathbf{x} \in \mathbb{R}^3 \mid \mathbf{n}_i^\mathrm{T} \mathbf{x} - d_i = 0\} { x ∈ R 3 ∣ n i T x − d i = 0 } ( d i (d_i ( d i is a real number, and i = 1 , 2 , 3 , 4 ) i = 1, 2, 3, 4) i = 1 , 2 , 3 , 4 ) is minimized. The position vector of P \mathbf{P} P is represented by B − 1 v \mathbf{B}^{-1} \mathbf{v} B − 1 v using a vector v ∈ R 3 \mathbf{v} \in \mathbb{R}^3 v ∈ R 3 . Express v \mathbf{v} v using n i \mathbf{n}_i n i and d i ( i = 1 , 2 , 3 , 4 ) d_i \, (i = 1, 2, 3, 4) d i ( i = 1 , 2 , 3 , 4 ) .
(5) Let l i l_i l i be a straight line through a point Q i Q_i Q i , the position vector of which is x i ∈ R 3 \mathbf{x}_i \in \mathbb{R}^3 x i ∈ R 3 , parallel to n i \mathbf{n}_i n i ( i = 1 , 2 , 3 ) (i = 1, 2, 3) ( i = 1 , 2 , 3 ) in R 3 \mathbb{R}^3 R 3 . Let R i \mathbf{R}_i R i be the orthogonal projection of an arbitrary point P \mathbf{P} P , the position vector of which is y ∈ R 3 \mathbf{y} \in \mathbb{R}^3 y ∈ R 3 , onto l i l_i l i . The position vector of R i \mathbf{R}_i R i is represented by y − W i ( y − x i ) \mathbf{y} - \mathbf{W}_i(\mathbf{y} - \mathbf{x}_i) y − W i ( y − x i ) using a matrix W i ∈ R 3 × 3 \mathbf{W}_i \in \mathbb{R}^{3 \times 3} W i ∈ R 3 × 3 . The identity matrix is denoted by I ∈ R 3 × 3 \mathbf{I} \in \mathbb{R}^{3 \times 3} I ∈ R 3 × 3 .
(a) Express W i \mathbf{W}_i W i using n i \mathbf{n}_i n i and I \mathbf{I} I .
(b) Show that W i T W i = W i \mathbf{W}_i^\mathrm{T} \mathbf{W}_i = \mathbf{W}_i W i T W i = W i .
(c c c ) Consider a plane Σ = { x ∈ R 3 ∣ a T x = b } \Sigma = \{\mathbf{x} \in \mathbb{R}^3 \mid \mathbf{a}^\mathrm{T} \mathbf{x} = b\} Σ = { x ∈ R 3 ∣ a T x = b } ( a ∈ R 3 (\mathbf{a} \in \mathbb{R}^3 ( a ∈ R 3 is a non-zero vector, and b b b is a real number). Let S ∈ Σ \mathbf{S} \in \Sigma S ∈ Σ be the point from which the sum of squared distances to l 1 l_1 l 1 , l 2 l_2 l 2 , and l 3 l_3 l 3 is minimized. When n 1 \mathbf{n}_1 n 1 , n 2 \mathbf{n}_2 n 2 , and n 3 \mathbf{n}_3 n 3 are orthogonal to each other, the position vector of S \mathbf{S} S is represented by ( I − a a T a T a ) w + a b a T a \left( \mathbf{I} - \frac{\mathbf{a}\mathbf{a}^\mathrm{T}}{\mathbf{a}^\mathrm{T}\mathbf{a}} \right) \mathbf{w} + \frac{\mathbf{a}b}{\mathbf{a}^\mathrm{T}\mathbf{a}} ( I − a T a a a T ) w + a T a a b . using a vector w ∈ R 3 \mathbf{w} \in \mathbb{R}^3 w ∈ R 3 which is independent of a \mathbf{a} a and b b b . Express w \mathbf{w} w using W i \mathbf{W}_i W i and x i ( i = 1 , 2 , 3 ) \mathbf{x}_i \, (i = 1, 2, 3) x i ( i = 1 , 2 , 3 ) .
Kai
(1)
Given the matrix A \mathbf{A} A :
A = ( n 1 T − n 2 T n 2 T − n 3 T n 3 T − n 4 T ) , \mathbf{A} = \begin{pmatrix}
\mathbf{n}_1^\mathrm{T} - \mathbf{n}_2^\mathrm{T} \\
\mathbf{n}_2^\mathrm{T} - \mathbf{n}_3^\mathrm{T} \\
\mathbf{n}_3^\mathrm{T} - \mathbf{n}_4^\mathrm{T}
\end{pmatrix}, A = n 1 T − n 2 T n 2 T − n 3 T n 3 T − n 4 T ,
we need to determine the conditions on n 4 \mathbf{n}_4 n 4 that ensure A \mathbf{A} A has a rank of three.
Let's assume that the third row of matrix A \mathbf{A} A can be written as a linear combination of the first two rows. Thus, we assume:
n 3 T − n 4 T = α ( n 1 T − n 2 T ) + β ( n 2 T − n 3 T ) , \mathbf{n}_3^\mathrm{T} - \mathbf{n}_4^\mathrm{T} = \alpha (\mathbf{n}_1^\mathrm{T} - \mathbf{n}_2^\mathrm{T}) + \beta (\mathbf{n}_2^\mathrm{T} - \mathbf{n}_3^\mathrm{T}), n 3 T − n 4 T = α ( n 1 T − n 2 T ) + β ( n 2 T − n 3 T ) ,
where α \alpha α and β \beta β are some scalars. Substituting n 4 \mathbf{n}_4 n 4 as a linear combination of n 1 \mathbf{n}_1 n 1 , n 2 \mathbf{n}_2 n 2 , and n 3 \mathbf{n}_3 n 3 , we have:
n 3 T − ( c 1 n 1 T + c 2 n 2 T + c 3 n 3 T ) = α ( n 1 T − n 2 T ) + β ( n 2 T − n 3 T ) . \mathbf{n}_3^\mathrm{T} - (c_1 \mathbf{n}_1^\mathrm{T} + c_2 \mathbf{n}_2^\mathrm{T} + c_3 \mathbf{n}_3^\mathrm{T}) = \alpha (\mathbf{n}_1^\mathrm{T} - \mathbf{n}_2^\mathrm{T}) + \beta (\mathbf{n}_2^\mathrm{T} - \mathbf{n}_3^\mathrm{T}). n 3 T − ( c 1 n 1 T + c 2 n 2 T + c 3 n 3 T ) = α ( n 1 T − n 2 T ) + β ( n 2 T − n 3 T ) .
Expanding and rearranging the equation, we get:
n 3 T − c 1 n 1 T − c 2 n 2 T − c 3 n 3 T = α n 1 T − α n 2 T + β n 2 T − β n 3 T . \mathbf{n}_3^\mathrm{T} - c_1 \mathbf{n}_1^\mathrm{T} - c_2 \mathbf{n}_2^\mathrm{T} - c_3 \mathbf{n}_3^\mathrm{T} = \alpha \mathbf{n}_1^\mathrm{T} - \alpha \mathbf{n}_2^\mathrm{T} + \beta \mathbf{n}_2^\mathrm{T} - \beta \mathbf{n}_3^\mathrm{T}. n 3 T − c 1 n 1 T − c 2 n 2 T − c 3 n 3 T = α n 1 T − α n 2 T + β n 2 T − β n 3 T .
Grouping like terms:
( 1 − c 3 + β ) n 3 T − c 1 n 1 T − c 2 n 2 T = α n 1 T + ( β − α ) n 2 T . (1 - c_3 + \beta) \mathbf{n}_3^\mathrm{T} - c_1 \mathbf{n}_1^\mathrm{T} - c_2 \mathbf{n}_2^\mathrm{T} = \alpha \mathbf{n}_1^\mathrm{T} + (\beta - \alpha) \mathbf{n}_2^\mathrm{T}. ( 1 − c 3 + β ) n 3 T − c 1 n 1 T − c 2 n 2 T = α n 1 T + ( β − α ) n 2 T .
For this equation to hold for arbitrary vectors n 1 \mathbf{n}_1 n 1 , n 2 \mathbf{n}_2 n 2 , and n 3 \mathbf{n}_3 n 3 , the coefficients of each vector must match:
For n 1 \mathbf{n}_1 n 1 :
− c 1 = α . -c_1 = \alpha. − c 1 = α .
For n 2 \mathbf{n}_2 n 2 :
− c 2 = β − α . -c_2 = \beta - \alpha. − c 2 = β − α .
For n 3 \mathbf{n}_3 n 3 :
1 − c 3 + β = 0. 1 - c_3 + \beta = 0. 1 − c 3 + β = 0.
Thus, we have the following system of equations:
α = − c 1 , \alpha = -c_1, α = − c 1 ,
β = α − c 2 = − c 1 − c 2 , \beta = \alpha - c_2 = -c_1 - c_2, β = α − c 2 = − c 1 − c 2 ,
1 − c 3 + β = 0 ⇒ 1 − c 3 = − β . 1 - c_3 + \beta = 0 \Rightarrow 1 - c_3 = -\beta. 1 − c 3 + β = 0 ⇒ 1 − c 3 = − β .
Substituting β = − c 1 − c 2 \beta = -c_1 - c_2 β = − c 1 − c 2 into the last equation:
1 − c 3 = c 1 + c 2 . 1 - c_3 = c_1 + c_2. 1 − c 3 = c 1 + c 2 .
So, the conditions under which the third row of A \mathbf{A} A can be written as a linear combination of the first two rows (i.e., the matrix would not have full rank) are:
c 1 + c 2 + c 3 = 1. c_1 + c_2 + c_3 = 1. c 1 + c 2 + c 3 = 1.
Conclusion for Full Rank
For the matrix A \mathbf{A} A to have full rank (rank 3), n 4 \mathbf{n}_4 n 4 must be such that the above condition does not hold. Therefore, the condition for the rank of A \mathbf{A} A to be three is:
c 1 + c 2 + c 3 ≠ 1. c_1 + c_2 + c_3 \neq 1. c 1 + c 2 + c 3 = 1.
(2)
Problem Setup
We are given four planes in R 3 \mathbb{R}^3 R 3 :
Π i = { x ∈ R 3 ∣ n i T x − d i = 0 } , i = 1 , 2 , 3 , 4 , \Pi_i = \{\mathbf{x} \in \mathbb{R}^3 \mid \mathbf{n}_i^\mathrm{T} \mathbf{x} - d_i = 0\}, \quad i = 1, 2, 3, 4, Π i = { x ∈ R 3 ∣ n i T x − d i = 0 } , i = 1 , 2 , 3 , 4 ,
where n 1 , n 2 , n 3 \mathbf{n}_1, \mathbf{n}_2, \mathbf{n}_3 n 1 , n 2 , n 3 , and n 4 \mathbf{n}_4 n 4 are unit vectors, and d 1 , d 2 , d 3 , d_1, d_2, d_3, d 1 , d 2 , d 3 , and d 4 d_4 d 4 are real numbers. The center of a sphere tangent to all four planes is represented as A − 1 u \mathbf{A}^{-1} \mathbf{u} A − 1 u , where A \mathbf{A} A is a 3x3 matrix, and u ∈ R 3 \mathbf{u} \in \mathbb{R}^3 u ∈ R 3 is what we need to find.
Conditions
For the sphere to be tangent to each plane, the distance from the center of the sphere A − 1 u \mathbf{A}^{-1} \mathbf{u} A − 1 u to each plane must satisfy:
n i T A − 1 u = d i + r , i = 1 , 2 , 3 , 4. \mathbf{n}_i^\mathrm{T} \mathbf{A}^{-1} \mathbf{u} = d_i + r, \quad i = 1, 2, 3, 4. n i T A − 1 u = d i + r , i = 1 , 2 , 3 , 4.
We subtract the equations pairwise to eliminate r r r , yielding:
n 2 T A − 1 u − n 1 T A − 1 u = d 2 − d 1 , \mathbf{n}_2^\mathrm{T} \mathbf{A}^{-1} \mathbf{u} - \mathbf{n}_1^\mathrm{T} \mathbf{A}^{-1} \mathbf{u} = d_2 - d_1, n 2 T A − 1 u − n 1 T A − 1 u = d 2 − d 1 ,
n 3 T A − 1 u − n 2 T A − 1 u = d 3 − d 2 , \mathbf{n}_3^\mathrm{T} \mathbf{A}^{-1} \mathbf{u} - \mathbf{n}_2^\mathrm{T} \mathbf{A}^{-1} \mathbf{u} = d_3 - d_2, n 3 T A − 1 u − n 2 T A − 1 u = d 3 − d 2 ,
n 4 T A − 1 u − n 3 T A − 1 u = d 4 − d 3 . \mathbf{n}_4^\mathrm{T} \mathbf{A}^{-1} \mathbf{u} - \mathbf{n}_3^\mathrm{T} \mathbf{A}^{-1} \mathbf{u} = d_4 - d_3. n 4 T A − 1 u − n 3 T A − 1 u = d 4 − d 3 .
These can be rewritten as:
( n 2 T − n 1 T ) A − 1 u = d 2 − d 1 , (\mathbf{n}_2^\mathrm{T} - \mathbf{n}_1^\mathrm{T}) \mathbf{A}^{-1} \mathbf{u} = d_2 - d_1, ( n 2 T − n 1 T ) A − 1 u = d 2 − d 1 ,
( n 3 T − n 2 T ) A − 1 u = d 3 − d 2 , (\mathbf{n}_3^\mathrm{T} - \mathbf{n}_2^\mathrm{T}) \mathbf{A}^{-1} \mathbf{u} = d_3 - d_2, ( n 3 T − n 2 T ) A − 1 u = d 3 − d 2 ,
( n 4 T − n 3 T ) A − 1 u = d 4 − d 3 . (\mathbf{n}_4^\mathrm{T} - \mathbf{n}_3^\mathrm{T}) \mathbf{A}^{-1} \mathbf{u} = d_4 - d_3. ( n 4 T − n 3 T ) A − 1 u = d 4 − d 3 .
Matrix Representation
The matrix A \mathbf{A} A is defined by the differences between the normals:
A = ( n 1 T − n 2 T n 2 T − n 3 T n 3 T − n 4 T ) . \mathbf{A} = \begin{pmatrix}
\mathbf{n}_1^\mathrm{T} - \mathbf{n}_2^\mathrm{T} \\
\mathbf{n}_2^\mathrm{T} - \mathbf{n}_3^\mathrm{T} \\
\mathbf{n}_3^\mathrm{T} - \mathbf{n}_4^\mathrm{T}
\end{pmatrix}. A = n 1 T − n 2 T n 2 T − n 3 T n 3 T − n 4 T .
Thus, we can write the system of equations in matrix form as:
A A − 1 u = ( d 2 − d 1 d 3 − d 2 d 4 − d 3 ) . \mathbf{A} \mathbf{A}^{-1} \mathbf{u} = \begin{pmatrix}
d_2 - d_1 \\
d_3 - d_2 \\
d_4 - d_3
\end{pmatrix}. A A − 1 u = d 2 − d 1 d 3 − d 2 d 4 − d 3 .
Simplifying, we find:
u = ( d 2 − d 1 d 3 − d 2 d 4 − d 3 ) . \mathbf{u} = \begin{pmatrix}
d_2 - d_1 \\
d_3 - d_2 \\
d_4 - d_3
\end{pmatrix}. u = d 2 − d 1 d 3 − d 2 d 4 − d 3 .
(3)
The matrix B \mathbf{B} B is defined as:
B = ∑ i = 1 4 n i n i T . \mathbf{B} = \sum_{i=1}^{4} \mathbf{n}_i \mathbf{n}_i^\mathrm{T}. B = i = 1 ∑ 4 n i n i T .
First, we show that B \mathbf{B} B is symmetric. Since each term n i n i T \mathbf{n}_i \mathbf{n}_i^\mathrm{T} n i n i T is symmetric (as the outer product of a vector with itself is symmetric), their sum B \mathbf{B} B is also symmetric.
Next, to prove that B \mathbf{B} B is positive definite, we need to show that for any non-zero vector x ∈ R 3 \mathbf{x} \in \mathbb{R}^3 x ∈ R 3 , the quadratic form x T B x > 0 \mathbf{x}^\mathrm{T} \mathbf{B} \mathbf{x} > 0 x T Bx > 0 .
x T B x = x T ( ∑ i = 1 4 n i n i T ) x = ∑ i = 1 4 ( n i T x ) 2 . \mathbf{x}^\mathrm{T} \mathbf{B} \mathbf{x} = \mathbf{x}^\mathrm{T} \left( \sum_{i=1}^{4} \mathbf{n}_i \mathbf{n}_i^\mathrm{T} \right) \mathbf{x} = \sum_{i=1}^{4} (\mathbf{n}_i^\mathrm{T} \mathbf{x})^2. x T Bx = x T ( i = 1 ∑ 4 n i n i T ) x = i = 1 ∑ 4 ( n i T x ) 2 .
Since each term is nonnegative, we have
x T B x ≥ 0. \mathbf{x}^{\mathrm T}\mathbf{B}\mathbf{x}\ge 0. x T Bx ≥ 0.
Now suppose
x T B x = 0. \mathbf{x}^{\mathrm T}\mathbf{B}\mathbf{x}=0. x T Bx = 0.
Then
( n i T x ) 2 = 0 ( i = 1 , 2 , 3 , 4 ) , (\mathbf{n}_i^{\mathrm T}\mathbf{x})^2=0
\quad (i=1,2,3,4), ( n i T x ) 2 = 0 ( i = 1 , 2 , 3 , 4 ) ,
so in particular
n 1 T x = n 2 T x = n 3 T x = 0. \mathbf{n}_1^{\mathrm T}\mathbf{x}
=
\mathbf{n}_2^{\mathrm T}\mathbf{x}
=
\mathbf{n}_3^{\mathrm T}\mathbf{x}
=0. n 1 T x = n 2 T x = n 3 T x = 0.
Because n 1 , n 2 , n 3 \mathbf{n}_1,\mathbf{n}_2,\mathbf{n}_3 n 1 , n 2 , n 3 are linearly independent in R 3 \mathbb{R}^3 R 3 , they form a basis of R 3 \mathbb{R}^3 R 3 . Therefore the only vector orthogonal to all three is the zero vector, so x = 0 \mathbf{x}=0 x = 0 , contradicting the assumption that x ≠ 0 \mathbf{x}\neq 0 x = 0 .
Thus,
x T B x > 0 ( x ≠ 0 ) , \mathbf{x}^{\mathrm T}\mathbf{B}\mathbf{x}>0
\qquad (\mathbf{x}\neq 0), x T Bx > 0 ( x = 0 ) ,
and therefore B \mathbf{B} B is positive definite.
Hence B \mathbf{B} B is a positive definite symmetric matrix .
(4)
The sum of squared distances from a point P \mathbf{P} P to the four planes is minimized when P \mathbf{P} P is the point of orthogonal projection of the origin onto these planes. The squared distance from a point P \mathbf{P} P with position vector x \mathbf{x} x to the plane Π i \Pi_i Π i is given by:
Distance 2 = ( n i T x − d i ∥ n i ∥ ) 2 . \text{Distance}^2 = \left(\frac{\mathbf{n}_i^\mathrm{T} \mathbf{x} - d_i}{\|\mathbf{n}_i\|}\right)^2. Distance 2 = ( ∥ n i ∥ n i T x − d i ) 2 .
Since n i \mathbf{n}_i n i are unit vectors (∥ n i ∥ = 1 \|\mathbf{n}_i\| = 1 ∥ n i ∥ = 1 ), this simplifies to:
Distance 2 = ( n i T x − d i ) 2 . \text{Distance}^2 = (\mathbf{n}_i^\mathrm{T} \mathbf{x} - d_i)^2. Distance 2 = ( n i T x − d i ) 2 .
The sum of squared distances to all four planes is:
S ( x ) = ∑ i = 1 4 ( n i T x − d i ) 2 . S(\mathbf{x}) = \sum_{i=1}^{4} (\mathbf{n}_i^\mathrm{T} \mathbf{x} - d_i)^2. S ( x ) = i = 1 ∑ 4 ( n i T x − d i ) 2 .
To minimize S ( x ) S(\mathbf{x}) S ( x ) , we take the gradient with respect to x \mathbf{x} x and set it equal to zero:
∇ S ( x ) = 2 ∑ i = 1 4 ( n i T x − d i ) n i = 0. \nabla S(\mathbf{x}) = 2 \sum_{i=1}^{4} (\mathbf{n}_i^\mathrm{T} \mathbf{x} - d_i) \mathbf{n}_i = 0. ∇ S ( x ) = 2 i = 1 ∑ 4 ( n i T x − d i ) n i = 0.
This equation can be rearranged into the form:
( ∑ i = 1 4 n i n i T ) x = ∑ i = 1 4 d i n i . \left(\sum_{i=1}^{4} \mathbf{n}_i \mathbf{n}_i^\mathrm{T}\right) \mathbf{x} = \sum_{i=1}^{4} d_i \mathbf{n}_i. ( i = 1 ∑ 4 n i n i T ) x = i = 1 ∑ 4 d i n i .
The matrix B \mathbf{B} B is defined as:
B = ∑ i = 1 4 n i n i T , \mathbf{B} = \sum_{i=1}^{4} \mathbf{n}_i \mathbf{n}_i^\mathrm{T}, B = i = 1 ∑ 4 n i n i T ,
which is a 3 × 3 3 \times 3 3 × 3 matrix. Therefore, the position vector x \mathbf{x} x that minimizes the sum of squared distances can be expressed as:
x = B − 1 ∑ i = 1 4 d i n i . \mathbf{x} = \mathbf{B}^{-1} \sum_{i=1}^{4} d_i \mathbf{n}_i. x = B − 1 i = 1 ∑ 4 d i n i .
Given that P \mathbf{P} P is the point minimizing the sum of squared distances, its position vector is B − 1 v \mathbf{B}^{-1} \mathbf{v} B − 1 v , where v \mathbf{v} v is defined by:
v = ∑ i = 1 4 d i n i . \mathbf{v} = \sum_{i=1}^{4} d_i \mathbf{n}_i. v = i = 1 ∑ 4 d i n i .
(5)
Let l i l_i l i be the line through Q i Q_i Q i with direction n i \mathbf{n}_i n i , where the position vector of Q i Q_i Q i is x i \mathbf{x}_i x i , and let y ∈ R 3 \mathbf{y}\in\mathbb{R}^3 y ∈ R 3 be the position vector of an arbitrary point P P P .
(a) Express W i \mathbf{W}_i W i using n i \mathbf{n}_i n i and I \mathbf{I} I .
The orthogonal projection of y \mathbf{y} y onto the line l i l_i l i is
R i = x i + ( n i T ( y − x i ) ) n i . \mathbf{R}_i
=
\mathbf{x}_i+\bigl(\mathbf{n}_i^{\mathrm T}(\mathbf{y}-\mathbf{x}_i)\bigr)\mathbf{n}_i. R i = x i + ( n i T ( y − x i ) ) n i .
This can be rewritten as
R i = y − ( I − n i n i T ) ( y − x i ) . \mathbf{R}_i
=
\mathbf{y}
-
\left(\mathbf{I}-\mathbf{n}_i\mathbf{n}_i^{\mathrm T}\right)(\mathbf{y}-\mathbf{x}_i). R i = y − ( I − n i n i T ) ( y − x i ) .
Comparing this with
R i = y − W i ( y − x i ) , \mathbf{R}_i=\mathbf{y}-\mathbf{W}_i(\mathbf{y}-\mathbf{x}_i), R i = y − W i ( y − x i ) ,
we obtain
W i = I − n i n i T . \boxed{\mathbf{W}_i=\mathbf{I}-\mathbf{n}_i\mathbf{n}_i^{\mathrm T}.} W i = I − n i n i T .
(b) Show that W i T W i = W i \mathbf{W}_i^{\mathrm T}\mathbf{W}_i=\mathbf{W}_i W i T W i = W i .
From part (a),
W i = I − n i n i T . \mathbf{W}_i=\mathbf{I}-\mathbf{n}_i\mathbf{n}_i^{\mathrm T}. W i = I − n i n i T .
Since n i n i T \mathbf{n}_i\mathbf{n}_i^{\mathrm T} n i n i T is symmetric, W i \mathbf{W}_i W i is also symmetric:
W i T = W i . \mathbf{W}_i^{\mathrm T}=\mathbf{W}_i. W i T = W i .
Moreover, because n i \mathbf{n}_i n i is a unit vector,
( n i n i T ) 2 = n i ( n i T n i ) n i T = n i n i T . (\mathbf{n}_i\mathbf{n}_i^{\mathrm T})^2
=
\mathbf{n}_i(\mathbf{n}_i^{\mathrm T}\mathbf{n}_i)\mathbf{n}_i^{\mathrm T}
=
\mathbf{n}_i\mathbf{n}_i^{\mathrm T}. ( n i n i T ) 2 = n i ( n i T n i ) n i T = n i n i T .
Hence
W i 2 = ( I − n i n i T ) 2 = I − 2 n i n i T + ( n i n i T ) 2 = I − n i n i T = W i . \mathbf{W}_i^{2}
=
(\mathbf{I}-\mathbf{n}_i\mathbf{n}_i^{\mathrm T})^2
=
\mathbf{I}-2\mathbf{n}_i\mathbf{n}_i^{\mathrm T}
+(\mathbf{n}_i\mathbf{n}_i^{\mathrm T})^2
=
\mathbf{I}-\mathbf{n}_i\mathbf{n}_i^{\mathrm T}
=
\mathbf{W}_i. W i 2 = ( I − n i n i T ) 2 = I − 2 n i n i T + ( n i n i T ) 2 = I − n i n i T = W i .
Therefore,
W i T W i = W i . \boxed{\mathbf{W}_i^{\mathrm T}\mathbf{W}_i=\mathbf{W}_i.} W i T W i = W i .
(c) Express w \mathbf{w} w using W i \mathbf{W}_i W i and x i \mathbf{x}_i x i ( i = 1 , 2 , 3 ) (i=1,2,3) ( i = 1 , 2 , 3 ) , assuming n 1 , n 2 , n 3 \mathbf{n}_1,\mathbf{n}_2,\mathbf{n}_3 n 1 , n 2 , n 3 are mutually orthogonal.
We want the point S ∈ Σ \mathbf{S}\in\Sigma S ∈ Σ , where
Σ = { x ∈ R 3 ∣ a T x = b } , \Sigma=\{\mathbf{x}\in\mathbb{R}^3\mid \mathbf{a}^{\mathrm T}\mathbf{x}=b\}, Σ = { x ∈ R 3 ∣ a T x = b } ,
such that the sum of squared distances from S \mathbf{S} S to the three lines l 1 , l 2 , l 3 l_1,l_2,l_3 l 1 , l 2 , l 3 is minimized.
For a point y ∈ R 3 \mathbf{y}\in\mathbb{R}^3 y ∈ R 3 , the vector from R i \mathbf{R}_i R i to y \mathbf{y} y is
y − R i = W i ( y − x i ) , \mathbf{y}-\mathbf{R}_i=\mathbf{W}_i(\mathbf{y}-\mathbf{x}_i), y − R i = W i ( y − x i ) ,
so the squared distance from y \mathbf{y} y to l i l_i l i is
∥ y − R i ∥ 2 = ∥ W i ( y − x i ) ∥ 2 . \|\mathbf{y}-\mathbf{R}_i\|^2
=
\|\mathbf{W}_i(\mathbf{y}-\mathbf{x}_i)\|^2. ∥ y − R i ∥ 2 = ∥ W i ( y − x i ) ∥ 2 .
Thus the objective function is
f ( y ) = ∑ i = 1 3 ∥ W i ( y − x i ) ∥ 2 . f(\mathbf{y})
=
\sum_{i=1}^3 \|\mathbf{W}_i(\mathbf{y}-\mathbf{x}_i)\|^2. f ( y ) = i = 1 ∑ 3 ∥ W i ( y − x i ) ∥ 2 .
Using part (b), this becomes
f ( y ) = ∑ i = 1 3 ( y − x i ) T W i ( y − x i ) . f(\mathbf{y})
=
\sum_{i=1}^3 (\mathbf{y}-\mathbf{x}_i)^{\mathrm T}\mathbf{W}_i(\mathbf{y}-\mathbf{x}_i). f ( y ) = i = 1 ∑ 3 ( y − x i ) T W i ( y − x i ) .
Expanding,
f ( y ) = y T ( ∑ i = 1 3 W i ) y − 2 y T ∑ i = 1 3 W i x i + constant . f(\mathbf{y})
=
\mathbf{y}^{\mathrm T}\left(\sum_{i=1}^3 \mathbf{W}_i\right)\mathbf{y}
-2\mathbf{y}^{\mathrm T}\sum_{i=1}^3 \mathbf{W}_i\mathbf{x}_i
+\text{constant}. f ( y ) = y T ( i = 1 ∑ 3 W i ) y − 2 y T i = 1 ∑ 3 W i x i + constant .
Now, since n 1 , n 2 , n 3 \mathbf{n}_1,\mathbf{n}_2,\mathbf{n}_3 n 1 , n 2 , n 3 are mutually orthogonal unit vectors, they form an orthonormal basis of R 3 \mathbb{R}^3 R 3 . Therefore,
n 1 n 1 T + n 2 n 2 T + n 3 n 3 T = I . \mathbf{n}_1\mathbf{n}_1^{\mathrm T}
+\mathbf{n}_2\mathbf{n}_2^{\mathrm T}
+\mathbf{n}_3\mathbf{n}_3^{\mathrm T}
=
\mathbf{I}. n 1 n 1 T + n 2 n 2 T + n 3 n 3 T = I .
Hence
∑ i = 1 3 W i = ∑ i = 1 3 ( I − n i n i T ) = 3 I − I = 2 I . \sum_{i=1}^3 \mathbf{W}_i
=
\sum_{i=1}^3 (\mathbf{I}-\mathbf{n}_i\mathbf{n}_i^{\mathrm T})
=
3\mathbf{I}-\mathbf{I}
=
2\mathbf{I}. i = 1 ∑ 3 W i = i = 1 ∑ 3 ( I − n i n i T ) = 3 I − I = 2 I .
So
f ( y ) = 2 y T y − 2 y T ∑ i = 1 3 W i x i + constant . f(\mathbf{y})
=
2\mathbf{y}^{\mathrm T}\mathbf{y}
-2\mathbf{y}^{\mathrm T}\sum_{i=1}^3 \mathbf{W}_i\mathbf{x}_i
+\text{constant}. f ( y ) = 2 y T y − 2 y T i = 1 ∑ 3 W i x i + constant .
Completing the square, we get
f ( y ) = 2 ∥ y − 1 2 ∑ i = 1 3 W i x i ∥ 2 + constant . f(\mathbf{y})
=
2\left\|
\mathbf{y}-\frac12\sum_{i=1}^3 \mathbf{W}_i\mathbf{x}_i
\right\|^2
+\text{constant}. f ( y ) = 2 y − 2 1 i = 1 ∑ 3 W i x i 2 + constant .
Therefore, the unconstrained minimizer is
w = 1 2 ∑ i = 1 3 W i x i . \mathbf{w}
=
\frac12\sum_{i=1}^3 \mathbf{W}_i\mathbf{x}_i. w = 2 1 i = 1 ∑ 3 W i x i .
Since S \mathbf{S} S is constrained to lie on the plane Σ \Sigma Σ , it is the orthogonal projection of w \mathbf{w} w onto Σ \Sigma Σ , which is why its position vector is written as
( I − a a T a T a ) w + a b a T a . \left(\mathbf{I}-\frac{\mathbf{a}\mathbf{a}^{\mathrm T}}{\mathbf{a}^{\mathrm T}\mathbf{a}}\right)\mathbf{w}
+\frac{\mathbf{a}b}{\mathbf{a}^{\mathrm T}\mathbf{a}}. ( I − a T a a a T ) w + a T a a b .
Thus,
w = 1 2 ∑ i = 1 3 W i x i . \boxed{
\mathbf{w}
=
\frac12\sum_{i=1}^3 \mathbf{W}_i\mathbf{x}_i
}. w = 2 1 i = 1 ∑ 3 W i x i .
Knowledge
矩阵秩 正定矩阵 最小二乘法 正交投影
难点思路
题目较难的部分是处理涉及到多平面的几何关系和正定矩阵的性质证明。特别是第 4 问中的最小二乘问题,需要对平面到点的距离公式有深刻理解。
解题技巧和信息
在解答此类问题时,明确矩阵的几何意义和代数性质非常关键。利用向量投影和最小二乘法的基本原理,可以有效地处理平面、直线和点之间的距离问题。
重点词汇
Rank of a matrix: 矩阵的秩
Positive definite matrix: 正定矩阵
Orthogonal projection: 正交投影
Least squares: 最小二乘法
参考资料
Gilbert Strang, Linear Algebra and Its Applications , 4th Edition, Section 6.5.
David C. Lay, Linear Algebra and Its Applications , 5th Edition, Chapter 7.