跳到主要内容

東京大学 新領域創成科学研究科 メディカル情報生命専攻 2017年8月実施 問題8

Author

zephyr, 祭音Myyura

Description

Let A\mathbf{A} be an n×mn \times m real matrix with positive rank rr. Such a matrix has a singular value decomposition A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T, where U\mathbf{U} and V\mathbf{V} are n×rn \times r, m×rm \times r real matrices, respectively, and satisfy UTU=Ir\mathbf{U}^T \mathbf{U} = \mathbf{I}_r, VTV=Ir\mathbf{V}^T \mathbf{V} = \mathbf{I}_r (Id\mathbf{I}_d: d×dd \times d unit matrix, MT\mathbf{M}^T: transpose of matrix M\mathbf{M}). Σ\mathbf{\Sigma} is an r×rr \times r real diagonal matrix whose diagonal elements Σkk=σk\Sigma_{kk} = \sigma_k (k=1,,rk = 1, \ldots, r) satisfy σ1σr>0\sigma_1 \geq \cdots \geq \sigma_r > 0.

(1) Describe all the positive eigenvalues and associated normalized eigenvectors of matrix ATA\mathbf{A}^T \mathbf{A}.

(2) Let TA:RmRnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n be a linear mapping defined by TA(x)=AxT_{\mathbf{A}} (\mathbf{x}) = \mathbf{A} \mathbf{x}. Describe the conditions on n,m,rn, m, r such that TAT_{\mathbf{A}} is surjective. Also, describe the conditions on n,m,rn, m, r such that TAT_{\mathbf{A}} is injective.

(3) The pseudoinverse of A\mathbf{A} is defined by A+=VΣ1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T. Let B=(ImA+A)\mathbf{B} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) and define linear mapping TB:RmRmT_{\mathbf{B}}: \mathbb{R}^m \to \mathbb{R}^m by TB(x)=BxT_{\mathbf{B}} (\mathbf{x}) = \mathbf{B} \mathbf{x}. Show that image Im(TB)={BxxRm}\mathrm{Im}(T_{\mathbf{B}}) = \{\mathbf{B} \mathbf{x} \mid \mathbf{x} \in \mathbb{R}^m\} is linearly isomorphic to kernel Ker(TA)={xRmAx=0n}\mathrm{Ker}(T_{\mathbf{A}}) = \{\mathbf{x} \in \mathbb{R}^m \mid \mathbf{A} \mathbf{x} = \mathbf{0}_n\} (0d\mathbf{0}_d: dd dimensional zero vector).

(4) Show that x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2 (x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x}, x2=(xx1)\mathbf{x}_2 = (\mathbf{x} - \mathbf{x}_1)) is an orthogonal decomposition.

(5) For a given bRn\mathbf{b} \in \mathbb{R}^n, let x0=A+bRm\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b} \in \mathbb{R}^m. Show that x=x0\mathbf{x} = \mathbf{x}_0 minimizes (Axb)T(Axb)(\mathbf{A} \mathbf{x} - \mathbf{b})^T (\mathbf{A} \mathbf{x} - \mathbf{b}). (Hint: Axb=A(xx0)+(Ax0b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b}))


A\mathbf{A} 为一个 n×mn \times m 的实矩阵,且正秩为 rr。这样的矩阵有一个奇异值分解 A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T,其中 U\mathbf{U}V\mathbf{V} 分别是 n×rn \times rm×rm \times r 的实矩阵,并且满足 UTU=Ir\mathbf{U}^T \mathbf{U} = \mathbf{I}_rVTV=Ir\mathbf{V}^T \mathbf{V} = \mathbf{I}_rId\mathbf{I}_dd×dd \times d 单位矩阵,MT\mathbf{M}^T:矩阵 M\mathbf{M} 的转置)。Σ\mathbf{\Sigma} 是一个 r×rr \times r 的实对角矩阵,其对角元素 Σkk=σk\Sigma_{kk} = \sigma_kk=1,,rk = 1, \ldots, r)满足 σ1σr>0\sigma_1 \geq \cdots \geq \sigma_r > 0

(1) 描述矩阵 ATA\mathbf{A}^T \mathbf{A} 的所有正特征值和相关的归一化特征向量。

(2) 令 TA:RmRnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n 为由 TA(x)=AxT_{\mathbf{A}} (\mathbf{x}) = \mathbf{A} \mathbf{x} 定义的线性映射。描述 n,m,rn, m, r 的条件,使得 TAT_{\mathbf{A}} 是满射。同时,描述 n,m,rn, m, r 的条件,使得 TAT_{\mathbf{A}} 是单射。

(3) A\mathbf{A} 的伪逆定义为 A+=VΣ1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T。令 B=(ImA+A)\mathbf{B} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) 并定义线性映射 TB:RmRmT_{\mathbf{B}}: \mathbb{R}^m \to \mathbb{R}^mTB(x)=BxT_{\mathbf{B}} (\mathbf{x}) = \mathbf{B} \mathbf{x}。证明 Im(TB)={BxxRm}\mathrm{Im}(T_{\mathbf{B}}) = \{\mathbf{B} \mathbf{x} \mid \mathbf{x} \in \mathbb{R}^m\} 在线性上同构于 Ker(TA)={xRmAx=0n}\mathrm{Ker}(T_{\mathbf{A}}) = \{\mathbf{x} \in \mathbb{R}^m \mid \mathbf{A} \mathbf{x} = \mathbf{0}_n\}0d\mathbf{0}_ddd 维零向量)。

(4) 证明 x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x}x2=(xx1)\mathbf{x}_2 = (\mathbf{x} - \mathbf{x}_1))是一个正交分解。

(5) 对于给定的 bRn\mathbf{b} \in \mathbb{R}^n,令 x0=A+bRm\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b} \in \mathbb{R}^m。证明 x=x0\mathbf{x} = \mathbf{x}_0 最小化 (Axb)T(Axb)(\mathbf{A} \mathbf{x} - \mathbf{b})^T (\mathbf{A} \mathbf{x} - \mathbf{b})。 (提示:Axb=A(xx0)+(Ax0b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b})

题目描述

A\mathbf A 是秩为 r>0r>0n×mn\times m 实矩阵,其薄奇异值分解为

A=UΣVT,\mathbf A=\mathbf U\mathbf\Sigma\mathbf V^T,

其中 URn×r\mathbf U\in\mathbb R^{n\times r}VRm×r\mathbf V\in\mathbb R^{m\times r},满足

UTU=Ir,VTV=Ir,\mathbf U^T\mathbf U=\mathbf I_r,\qquad \mathbf V^T\mathbf V=\mathbf I_r,

Σ\mathbf\Sigmar×rr\times r 对角矩阵,且

Σkk=σk,σ1σr>0.\Sigma_{kk}=\sigma_k,\qquad \sigma_1\ge\cdots\ge\sigma_r>0.

回答下列问题:

  1. 写出 ATA\mathbf A^T\mathbf A 的全部正特征值及对应的单位特征向量。

  2. 对线性映射

    TA:RmRn,TA(x)=Ax,T_{\mathbf A}:\mathbb R^m\to\mathbb R^n,\qquad T_{\mathbf A}(\mathbf x)=\mathbf A\mathbf x,

    分别给出它为满射、为单射时 n,m,rn,m,r 应满足的条件。

  3. 定义 Moore–Penrose 伪逆

    A+=VΣ1UT\mathbf A^+=\mathbf V\mathbf\Sigma^{-1}\mathbf U^T

    以及

    B=ImA+A,TB(x)=Bx.\mathbf B=\mathbf I_m-\mathbf A^+\mathbf A,\qquad T_{\mathbf B}(\mathbf x)=\mathbf B\mathbf x.

    证明 Im(TB)\operatorname{Im}(T_{\mathbf B})ker(TA)\ker(T_{\mathbf A}) 线性同构。

  4. x1=Bx,x2=xx1,\mathbf x_1=\mathbf B\mathbf x,\qquad \mathbf x_2=\mathbf x-\mathbf x_1,

    证明 x=x1+x2\mathbf x=\mathbf x_1+\mathbf x_2 是正交分解。

  5. 给定 bRn\mathbf b\in\mathbb R^n,令 x0=A+b\mathbf x_0=\mathbf A^+\mathbf b,证明 x=x0\mathbf x=\mathbf x_0 最小化

    (Axb)T(Axb).(\mathbf A\mathbf x-\mathbf b)^T(\mathbf A\mathbf x-\mathbf b).

    可使用提示

    Axb=A(xx0)+(Ax0b).\mathbf A\mathbf x-\mathbf b =\mathbf A(\mathbf x-\mathbf x_0)+(\mathbf A\mathbf x_0-\mathbf b).

Kai

(1)

Given the singular value decomposition (SVD) of A\mathbf{A} as A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T, we can express ATA\mathbf{A}^T \mathbf{A} as follows:

ATA=(UΣVT)T(UΣVT)=VΣTUTUΣVT=VΣ2VT\mathbf{A}^T \mathbf{A} = (\mathbf{U} \mathbf{\Sigma} \mathbf{V}^T)^T (\mathbf{U} \mathbf{\Sigma} \mathbf{V}^T) = \mathbf{V} \mathbf{\Sigma}^T \mathbf{U}^T \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T = \mathbf{V} \mathbf{\Sigma}^2 \mathbf{V}^T

The matrix Σ2\mathbf{\Sigma}^2 is diagonal with the diagonal elements σk2\sigma_k^2 (k=1,,rk = 1, \ldots, r). Thus, the positive eigenvalues of ATA\mathbf{A}^T \mathbf{A} are exactly the σk2\sigma_k^2, and the corresponding columns vk\mathbf v_k of V\mathbf V are normalized eigenvectors. If a singular value is repeated, every unit vector in the span of the corresponding vk\mathbf v_k is an associated normalized eigenvector.

(2)

Surjective (onto): The mapping TA:RmRnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n is surjective if the range of A\mathbf{A} spans Rn\mathbb{R}^n, i.e., A\mathbf{A} has full row rank. This occurs when r=nmr = n \leq m.

Injective (one-to-one): The mapping TAT_{\mathbf{A}} is injective if the kernel of A\mathbf{A} contains only the zero vector, i.e., A\mathbf{A} has full column rank. This occurs when r=mnr = m \leq n.

(3)

The pseudoinverse A+\mathbf{A}^+ is defined as A+=VΣ1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T. Consider B=ImA+A\mathbf{B} = \mathbf{I}_m - \mathbf{A}^+ \mathbf{A}.

We need to show that Im(TB)\mathrm{Im}(T_{\mathbf{B}}) is isomorphic to Ker(TA)\mathrm{Ker}(T_{\mathbf{A}}). Observe the following:

AB=A(ImA+A)=AAA+A=0.\mathbf{A}\mathbf{B} =\mathbf{A}(\mathbf{I}_m-\mathbf{A}^+\mathbf{A}) =\mathbf{A}-\mathbf{A}\mathbf{A}^+\mathbf{A} =0.

Thus, Im(B)Ker(A)\mathrm{Im}(\mathbf{B}) \subseteq \mathrm{Ker}(\mathbf{A}).

Now, consider xKer(A)\mathbf{x} \in \mathrm{Ker}(\mathbf{A}). Then Ax=0\mathbf{A} \mathbf{x} = \mathbf{0}, and

Bx=(ImA+A)x=x\mathbf{B} \mathbf{x} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) \mathbf{x} = \mathbf{x}

Thus, xIm(B)\mathbf{x} \in \mathrm{Im}(\mathbf{B}). Therefore, Im(B)=Ker(A)\mathrm{Im}(\mathbf{B}) = \mathrm{Ker}(\mathbf{A}).

(4)

Given x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2 where x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x} and x2=xx1\mathbf{x}_2 = \mathbf{x} - \mathbf{x}_1:

x2=xBx=x(ImA+A)x=A+Ax\mathbf{x}_2 = \mathbf{x} - \mathbf{B} \mathbf{x} = \mathbf{x} - (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) \mathbf{x} = \mathbf{A}^+ \mathbf{A} \mathbf{x}

To show orthogonality:

x1Tx2=(Bx)T(A+Ax)=xTBTA+Ax\mathbf{x}_1^T \mathbf{x}_2 = (\mathbf{B} \mathbf{x})^T (\mathbf{A}^+ \mathbf{A} \mathbf{x}) = \mathbf{x}^T \mathbf{B}^T \mathbf{A}^+ \mathbf{A} \mathbf{x}

Let P=A+A=VVT\mathbf P=\mathbf A^+\mathbf A=\mathbf V\mathbf V^T. Then P\mathbf P is symmetric and idempotent, and B=ImP\mathbf B=\mathbf I_m-\mathbf P. Hence:

x1Tx2=xT(ImP)Px=xT(PP2)x=0.\mathbf{x}_1^T\mathbf{x}_2 =\mathbf{x}^T(\mathbf I_m-\mathbf P)\mathbf P\mathbf{x} =\mathbf{x}^T(\mathbf P-\mathbf P^2)\mathbf{x} =0.

Thus, x1\mathbf{x}_1 and x2\mathbf{x}_2 are orthogonal.

(5)

Let x0=A+b\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b}. We need to show that x=x0\mathbf{x} = \mathbf{x}_0 minimizes the expression.

Consider the error:

Axb=A(xx0)+(Ax0b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b})

Since x0=A+b\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b}, the residual r=Ax0b=(AA+In)b\mathbf r=\mathbf A\mathbf x_0-\mathbf b=(\mathbf A\mathbf A^+-\mathbf I_n)\mathbf b is orthogonal to Im(A)\operatorname{Im}(\mathbf A). Thus r\mathbf r is orthogonal to A(xx0)\mathbf A(\mathbf x-\mathbf x_0), and:

Axb2=A(xx0)2+Ax0b2Ax0b2.\|\mathbf A\mathbf x-\mathbf b\|^2 =\|\mathbf A(\mathbf x-\mathbf x_0)\|^2+\|\mathbf A\mathbf x_0-\mathbf b\|^2 \geq \|\mathbf A\mathbf x_0-\mathbf b\|^2.

Therefore, x0\mathbf{x}_0 is a minimizer.

Knowledge

奇异值分解 线性映射 广义逆矩阵 正交分解 线性代数

重点词汇

  • singular value decomposition (SVD) 奇异值分解
  • pseudoinverse 广义逆
  • surjective 满射
  • injective 单射
  • orthogonal decomposition 正交分解

参考资料

  1. "Linear Algebra and Its Applications" by Gilbert Strang, Chapter 7: The Singular Value Decomposition (SVD)
  2. "Matrix Computations" by Gene H. Golub and Charles F. Van Loan, Chapter 2: Matrix Analysis

Reference