跳到主要内容

東京大学 新領域創成科学研究科 メディカル情報生命専攻 2017年8月実施 問題8

Author​

zephyr, 祭音Myyura

Description​

Let A\mathbf{A} be an n×mn \times m real matrix with positive rank rr. Such a matrix has a singular value decomposition A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T, where U\mathbf{U} and V\mathbf{V} are n×rn \times r, m×rm \times r real matrices, respectively, and satisfy UTU=Ir\mathbf{U}^T \mathbf{U} = \mathbf{I}_r, VTV=Ir\mathbf{V}^T \mathbf{V} = \mathbf{I}_r (Id\mathbf{I}_d: d×dd \times d unit matrix, MT\mathbf{M}^T: transpose of matrix M\mathbf{M}). Σ\mathbf{\Sigma} is an r×rr \times r real diagonal matrix whose diagonal elements Σkk=σk\Sigma_{kk} = \sigma_k (k=1,…,rk = 1, \ldots, r) satisfy σ1≥⋯≥σr>0\sigma_1 \geq \cdots \geq \sigma_r > 0.

(1) Describe all the positive eigenvalues and associated normalized eigenvectors of matrix ATA\mathbf{A}^T \mathbf{A}.

(2) Let TA:Rm→RnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n be a linear mapping defined by TA(x)=AxT_{\mathbf{A}} (\mathbf{x}) = \mathbf{A} \mathbf{x}. Describe the conditions on n,m,rn, m, r such that TAT_{\mathbf{A}} is surjective. Also, describe the conditions on n,m,rn, m, r such that TAT_{\mathbf{A}} is injective.

(3) The pseudoinverse of A\mathbf{A} is defined by A+=VΣ−1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T. Let B=(Im−A+A)\mathbf{B} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) and define linear mapping TB:Rm→RmT_{\mathbf{B}}: \mathbb{R}^m \to \mathbb{R}^m by TB(x)=BxT_{\mathbf{B}} (\mathbf{x}) = \mathbf{B} \mathbf{x}. Show that image Im(TB)={Bx∣x∈Rm}\mathrm{Im}(T_{\mathbf{B}}) = \{\mathbf{B} \mathbf{x} \mid \mathbf{x} \in \mathbb{R}^m\} is linearly isomorphic to kernel Ker(TA)={x∈Rm∣Ax=0n}\mathrm{Ker}(T_{\mathbf{A}}) = \{\mathbf{x} \in \mathbb{R}^m \mid \mathbf{A} \mathbf{x} = \mathbf{0}_n\} (0d\mathbf{0}_d: dd dimensional zero vector).

(4) Show that x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2 (x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x}, x2=(x−x1)\mathbf{x}_2 = (\mathbf{x} - \mathbf{x}_1)) is an orthogonal decomposition.

(5) For a given b∈Rn\mathbf{b} \in \mathbb{R}^n, let x0=A+b∈Rm\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b} \in \mathbb{R}^m. Show that x=x0\mathbf{x} = \mathbf{x}_0 minimizes (Ax−b)T(Ax−b)(\mathbf{A} \mathbf{x} - \mathbf{b})^T (\mathbf{A} \mathbf{x} - \mathbf{b}). (Hint: Ax−b=A(x−x0)+(Ax0−b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b}))


设 A\mathbf{A} 为一个 n×mn \times m 的实矩阵,且正秩为 rr。这样的矩阵有一个奇异值分解 A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T,其中 U\mathbf{U} 和 V\mathbf{V} 分别是 n×rn \times r、m×rm \times r 的实矩阵,并且满足 UTU=Ir\mathbf{U}^T \mathbf{U} = \mathbf{I}_r,VTV=Ir\mathbf{V}^T \mathbf{V} = \mathbf{I}_r(Id\mathbf{I}_d:d×dd \times d 单位矩阵,MT\mathbf{M}^T:矩阵 M\mathbf{M} 的转置)。Σ\mathbf{\Sigma} 是一个 r×rr \times r 的实对角矩阵,其对角元素 Σkk=σk\Sigma_{kk} = \sigma_k(k=1,…,rk = 1, \ldots, r)满足 σ1≥⋯≥σr>0\sigma_1 \geq \cdots \geq \sigma_r > 0。

(1) 描述矩阵 ATA\mathbf{A}^T \mathbf{A} 的所有正特征值和相关的归一化特征向量。

(2) 令 TA:Rm→RnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n 为由 TA(x)=AxT_{\mathbf{A}} (\mathbf{x}) = \mathbf{A} \mathbf{x} 定义的线性映射。描述 n,m,rn, m, r 的条件,使得 TAT_{\mathbf{A}} 是满射。同时,描述 n,m,rn, m, r 的条件,使得 TAT_{\mathbf{A}} 是单射。

(3) A\mathbf{A} 的伪逆定义为 A+=VΣ−1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T。令 B=(Im−A+A)\mathbf{B} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) 并定义线性映射 TB:Rm→RmT_{\mathbf{B}}: \mathbb{R}^m \to \mathbb{R}^m 由 TB(x)=BxT_{\mathbf{B}} (\mathbf{x}) = \mathbf{B} \mathbf{x}。证明 Im(TB)={Bx∣x∈Rm}\mathrm{Im}(T_{\mathbf{B}}) = \{\mathbf{B} \mathbf{x} \mid \mathbf{x} \in \mathbb{R}^m\} 在线性上同构于 Ker(TA)={x∈Rm∣Ax=0n}\mathrm{Ker}(T_{\mathbf{A}}) = \{\mathbf{x} \in \mathbb{R}^m \mid \mathbf{A} \mathbf{x} = \mathbf{0}_n\}(0d\mathbf{0}_d:dd 维零向量)。

(4) 证明 x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2(x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x},x2=(x−x1)\mathbf{x}_2 = (\mathbf{x} - \mathbf{x}_1))是一个正交分解。

(5) 对于给定的 b∈Rn\mathbf{b} \in \mathbb{R}^n,令 x0=A+b∈Rm\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b} \in \mathbb{R}^m。证明 x=x0\mathbf{x} = \mathbf{x}_0 最小化 (Ax−b)T(Ax−b)(\mathbf{A} \mathbf{x} - \mathbf{b})^T (\mathbf{A} \mathbf{x} - \mathbf{b})。 (提示:Ax−b=A(x−x0)+(Ax0−b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b}))

题目描述​

设 A\mathbf A 是秩为 r>0r>0 的 n×mn\times m 实矩阵,其薄奇异值分解为

A=UΣVT,\mathbf A=\mathbf U\mathbf\Sigma\mathbf V^T,

其中 U∈Rn×r\mathbf U\in\mathbb R^{n\times r}、V∈Rm×r\mathbf V\in\mathbb R^{m\times r},满足

UTU=Ir,VTV=Ir,\mathbf U^T\mathbf U=\mathbf I_r,\qquad \mathbf V^T\mathbf V=\mathbf I_r,

Σ\mathbf\Sigma 是 r×rr\times r 对角矩阵,且

Σkk=σk,σ1≥⋯≥σr>0.\Sigma_{kk}=\sigma_k,\qquad \sigma_1\ge\cdots\ge\sigma_r>0.

回答下列问题:

  1. 写出 ATA\mathbf A^T\mathbf A 的全部正特征值及对应的单位特征向量。

  2. 对线性映射

    TA:Rm→Rn,TA(x)=Ax,T_{\mathbf A}:\mathbb R^m\to\mathbb R^n,\qquad T_{\mathbf A}(\mathbf x)=\mathbf A\mathbf x,

    分别给出它为满射、为单射时 n,m,rn,m,r 应满足的条件。

  3. 定义 Moore–Penrose 伪逆

    A+=VΣ−1UT\mathbf A^+=\mathbf V\mathbf\Sigma^{-1}\mathbf U^T

    以及

    B=Im−A+A,TB(x)=Bx.\mathbf B=\mathbf I_m-\mathbf A^+\mathbf A,\qquad T_{\mathbf B}(\mathbf x)=\mathbf B\mathbf x.

    证明 Im⁡(TB)\operatorname{Im}(T_{\mathbf B}) 与 ker⁡(TA)\ker(T_{\mathbf A}) 线性同构。

  4. 对

    x1=Bx,x2=x−x1,\mathbf x_1=\mathbf B\mathbf x,\qquad \mathbf x_2=\mathbf x-\mathbf x_1,

    证明 x=x1+x2\mathbf x=\mathbf x_1+\mathbf x_2 是正交分解。

  5. 给定 b∈Rn\mathbf b\in\mathbb R^n,令 x0=A+b\mathbf x_0=\mathbf A^+\mathbf b,证明 x=x0\mathbf x=\mathbf x_0 最小化

    (Ax−b)T(Ax−b).(\mathbf A\mathbf x-\mathbf b)^T(\mathbf A\mathbf x-\mathbf b).

    可使用提示

    Ax−b=A(x−x0)+(Ax0−b).\mathbf A\mathbf x-\mathbf b =\mathbf A(\mathbf x-\mathbf x_0)+(\mathbf A\mathbf x_0-\mathbf b).

Kai​

(1)​

Given the singular value decomposition (SVD) of A\mathbf{A} as A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T, we can express ATA\mathbf{A}^T \mathbf{A} as follows:

ATA=(UΣVT)T(UΣVT)=VΣTUTUΣVT=VΣ2VT\mathbf{A}^T \mathbf{A} = (\mathbf{U} \mathbf{\Sigma} \mathbf{V}^T)^T (\mathbf{U} \mathbf{\Sigma} \mathbf{V}^T) = \mathbf{V} \mathbf{\Sigma}^T \mathbf{U}^T \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T = \mathbf{V} \mathbf{\Sigma}^2 \mathbf{V}^T

The matrix Σ2\mathbf{\Sigma}^2 is diagonal with the diagonal elements σk2\sigma_k^2 (k=1,…,rk = 1, \ldots, r). Thus, the positive eigenvalues of ATA\mathbf{A}^T \mathbf{A} are exactly the σk2\sigma_k^2, and the corresponding columns vk\mathbf v_k of V\mathbf V are normalized eigenvectors. If a singular value is repeated, every unit vector in the span of the corresponding vk\mathbf v_k is an associated normalized eigenvector.

(2)​

Surjective (onto): The mapping TA:Rm→RnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n is surjective if the range of A\mathbf{A} spans Rn\mathbb{R}^n, i.e., A\mathbf{A} has full row rank. This occurs when r=n≤mr = n \leq m.

Injective (one-to-one): The mapping TAT_{\mathbf{A}} is injective if the kernel of A\mathbf{A} contains only the zero vector, i.e., A\mathbf{A} has full column rank. This occurs when r=m≤nr = m \leq n.

(3)​

The pseudoinverse A+\mathbf{A}^+ is defined as A+=VΣ−1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T. Consider B=Im−A+A\mathbf{B} = \mathbf{I}_m - \mathbf{A}^+ \mathbf{A}.

We need to show that Im(TB)\mathrm{Im}(T_{\mathbf{B}}) is isomorphic to Ker(TA)\mathrm{Ker}(T_{\mathbf{A}}). Observe the following:

AB=A(Im−A+A)=A−AA+A=0.\mathbf{A}\mathbf{B} =\mathbf{A}(\mathbf{I}_m-\mathbf{A}^+\mathbf{A}) =\mathbf{A}-\mathbf{A}\mathbf{A}^+\mathbf{A} =0.

Thus, Im(B)⊆Ker(A)\mathrm{Im}(\mathbf{B}) \subseteq \mathrm{Ker}(\mathbf{A}).

Now, consider x∈Ker(A)\mathbf{x} \in \mathrm{Ker}(\mathbf{A}). Then Ax=0\mathbf{A} \mathbf{x} = \mathbf{0}, and

Bx=(Im−A+A)x=x\mathbf{B} \mathbf{x} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) \mathbf{x} = \mathbf{x}

Thus, x∈Im(B)\mathbf{x} \in \mathrm{Im}(\mathbf{B}). Therefore, Im(B)=Ker(A)\mathrm{Im}(\mathbf{B}) = \mathrm{Ker}(\mathbf{A}).

(4)​

Given x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2 where x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x} and x2=x−x1\mathbf{x}_2 = \mathbf{x} - \mathbf{x}_1:

x2=x−Bx=x−(Im−A+A)x=A+Ax\mathbf{x}_2 = \mathbf{x} - \mathbf{B} \mathbf{x} = \mathbf{x} - (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) \mathbf{x} = \mathbf{A}^+ \mathbf{A} \mathbf{x}

To show orthogonality:

x1Tx2=(Bx)T(A+Ax)=xTBTA+Ax\mathbf{x}_1^T \mathbf{x}_2 = (\mathbf{B} \mathbf{x})^T (\mathbf{A}^+ \mathbf{A} \mathbf{x}) = \mathbf{x}^T \mathbf{B}^T \mathbf{A}^+ \mathbf{A} \mathbf{x}

Let P=A+A=VVT\mathbf P=\mathbf A^+\mathbf A=\mathbf V\mathbf V^T. Then P\mathbf P is symmetric and idempotent, and B=Im−P\mathbf B=\mathbf I_m-\mathbf P. Hence:

x1Tx2=xT(Im−P)Px=xT(P−P2)x=0.\mathbf{x}_1^T\mathbf{x}_2 =\mathbf{x}^T(\mathbf I_m-\mathbf P)\mathbf P\mathbf{x} =\mathbf{x}^T(\mathbf P-\mathbf P^2)\mathbf{x} =0.

Thus, x1\mathbf{x}_1 and x2\mathbf{x}_2 are orthogonal.

(5)​

Let x0=A+b\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b}. We need to show that x=x0\mathbf{x} = \mathbf{x}_0 minimizes the expression.

Consider the error:

Ax−b=A(x−x0)+(Ax0−b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b})

Since x0=A+b\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b}, the residual r=Ax0−b=(AA+−In)b\mathbf r=\mathbf A\mathbf x_0-\mathbf b=(\mathbf A\mathbf A^+-\mathbf I_n)\mathbf b is orthogonal to Im⁡(A)\operatorname{Im}(\mathbf A). Thus r\mathbf r is orthogonal to A(x−x0)\mathbf A(\mathbf x-\mathbf x_0), and:

∥Ax−b∥2=∥A(x−x0)∥2+∥Ax0−b∥2≥∥Ax0−b∥2.\|\mathbf A\mathbf x-\mathbf b\|^2 =\|\mathbf A(\mathbf x-\mathbf x_0)\|^2+\|\mathbf A\mathbf x_0-\mathbf b\|^2 \geq \|\mathbf A\mathbf x_0-\mathbf b\|^2.

Therefore, x0\mathbf{x}_0 is a minimizer.

Knowledge​

奇异值分解 线性映射 广义逆矩阵 正交分解 线性代数

重点词汇​

  • singular value decomposition (SVD) 奇异值分解
  • pseudoinverse 广义逆
  • surjective 满射
  • injective 单射
  • orthogonal decomposition 正交分解

参考资料​

  1. "Linear Algebra and Its Applications" by Gilbert Strang, Chapter 7: The Singular Value Decomposition (SVD)
  2. "Matrix Computations" by Gene H. Golub and Charles F. Van Loan, Chapter 2: Matrix Analysis

Reference​