跳到主要内容

東京大学 新領域創成科学研究科 メディカル情報生命専攻 2017年8月実施 問題8

Author

zephyr

Description

Let A\mathbf{A} be an n×mn \times m real matrix with positive rank rr. Such a matrix has a singular value decomposition A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T, where U\mathbf{U} and V\mathbf{V} are n×rn \times r, m×rm \times r real matrices, respectively, and satisfy UTU=Ir\mathbf{U}^T \mathbf{U} = \mathbf{I}_r, VTV=Ir\mathbf{V}^T \mathbf{V} = \mathbf{I}_r (Id\mathbf{I}_d: d×dd \times d unit matrix, MT\mathbf{M}^T: transpose of matrix M\mathbf{M}). Σ\mathbf{\Sigma} is an r×rr \times r real diagonal matrix whose diagonal elements Σkk=σk\Sigma_{kk} = \sigma_k (k=1,,rk = 1, \ldots, r) satisfy σ1σr>0\sigma_1 \geq \cdots \geq \sigma_r > 0.

(1) Describe all the positive eigenvalues and associated normalized eigenvectors of matrix ATA\mathbf{A}^T \mathbf{A}.

(2) Let TA:RmRnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n be a linear mapping defined by TA(x)=AxT_{\mathbf{A}} (\mathbf{x}) = \mathbf{A} \mathbf{x}. Describe the conditions on n,m,rn, m, r such that TAT_{\mathbf{A}} is surjective. Also, describe the conditions on n,m,rn, m, r such that TAT_{\mathbf{A}} is injective.

(3) The pseudoinverse of A\mathbf{A} is defined by A+=VΣ1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T. Let B=(ImA+A)\mathbf{B} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) and define linear mapping TB:RmRmT_{\mathbf{B}}: \mathbb{R}^m \to \mathbb{R}^m by TB(x)=BxT_{\mathbf{B}} (\mathbf{x}) = \mathbf{B} \mathbf{x}. Show that image Im(TB)={BxxRm}\mathrm{Im}(T_{\mathbf{B}}) = \{\mathbf{B} \mathbf{x} \mid \mathbf{x} \in \mathbb{R}^m\} is linearly isomorphic to kernel Ker(TA)={xRmAx=0d}\mathrm{Ker}(T_{\mathbf{A}}) = \{\mathbf{x} \in \mathbb{R}^m \mid \mathbf{A} \mathbf{x} = \mathbf{0}_d\} (0d\mathbf{0}_d: dd dimensional zero vector).

(4) Show that x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2 (x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x}, x2=(xx1)\mathbf{x}_2 = (\mathbf{x} - \mathbf{x}_1)) is an orthogonal decomposition.

(5) For a given bRn\mathbf{b} \in \mathbb{R}^n, let x0=A+bRm\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b} \in \mathbb{R}^m. Show that x=x0\mathbf{x} = \mathbf{x}_0 minimizes (Axb)T(Axb)(\mathbf{A} \mathbf{x} - \mathbf{b})^T (\mathbf{A} \mathbf{x} - \mathbf{b}). (Hint: Axb=A(xx0)+(Ax0b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b}))


A\mathbf{A} 为一个 n×mn \times m 的实矩阵,且正秩为 rr。这样的矩阵有一个奇异值分解 A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T,其中 U\mathbf{U}V\mathbf{V} 分别是 n×rn \times rm×rm \times r 的实矩阵,并且满足 UTU=Ir\mathbf{U}^T \mathbf{U} = \mathbf{I}_rVTV=Ir\mathbf{V}^T \mathbf{V} = \mathbf{I}_rId\mathbf{I}_dd×dd \times d 单位矩阵,MT\mathbf{M}^T:矩阵 M\mathbf{M} 的转置)。Σ\mathbf{\Sigma} 是一个 r×rr \times r 的实对角矩阵,其对角元素 Σkk=σk\Sigma_{kk} = \sigma_kk=1,,rk = 1, \ldots, r)满足 σ1σr>0\sigma_1 \geq \cdots \geq \sigma_r > 0

(1) 描述矩阵 ATA\mathbf{A}^T \mathbf{A} 的所有正特征值和相关的归一化特征向量。

(2) 令 TA:RmRnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n 为由 TA(x)=AxT_{\mathbf{A}} (\mathbf{x}) = \mathbf{A} \mathbf{x} 定义的线性映射。描述 n,m,rn, m, r 的条件,使得 TAT_{\mathbf{A}} 是满射。同时,描述 n,m,rn, m, r 的条件,使得 TAT_{\mathbf{A}} 是单射。

(3) A\mathbf{A} 的伪逆定义为 A+=VΣ1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T。令 B=(ImA+A)\mathbf{B} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) 并定义线性映射 TB:RmRmT_{\mathbf{B}}: \mathbb{R}^m \to \mathbb{R}^mTB(x)=BxT_{\mathbf{B}} (\mathbf{x}) = \mathbf{B} \mathbf{x}。证明 Im(TB)={BxxRm}\mathrm{Im}(T_{\mathbf{B}}) = \{\mathbf{B} \mathbf{x} \mid \mathbf{x} \in \mathbb{R}^m\} 在线性上同构于 Ker(TA)={xRmAx=0d}\mathrm{Ker}(T_{\mathbf{A}}) = \{\mathbf{x} \in \mathbb{R}^m \mid \mathbf{A} \mathbf{x} = \mathbf{0}_d\}0d\mathbf{0}_ddd 维零向量)。

(4) 证明 x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x}x2=(xx1)\mathbf{x}_2 = (\mathbf{x} - \mathbf{x}_1))是一个正交分解。

(5) 对于给定的 bRn\mathbf{b} \in \mathbb{R}^n,令 x0=A+bRm\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b} \in \mathbb{R}^m。证明 x=x0\mathbf{x} = \mathbf{x}_0 最小化 (Axb)T(Axb)(\mathbf{A} \mathbf{x} - \mathbf{b})^T (\mathbf{A} \mathbf{x} - \mathbf{b})。 (提示:Axb=A(xx0)+(Ax0b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b})

题目描述

A\mathbf A 是秩为 r>0r>0n×mn\times m 实矩阵,其薄奇异值分解为

A=UΣVT,\mathbf A=\mathbf U\mathbf\Sigma\mathbf V^T,

其中 URn×r\mathbf U\in\mathbb R^{n\times r}VRm×r\mathbf V\in\mathbb R^{m\times r},满足

UTU=Ir,VTV=Ir,\mathbf U^T\mathbf U=\mathbf I_r,\qquad \mathbf V^T\mathbf V=\mathbf I_r,

Σ\mathbf\Sigmar×rr\times r 对角矩阵,且

Σkk=σk,σ1σr>0.\Sigma_{kk}=\sigma_k,\qquad \sigma_1\ge\cdots\ge\sigma_r>0.

回答下列问题:

  1. 写出 ATA\mathbf A^T\mathbf A 的全部正特征值及对应的单位特征向量。
  2. 对线性映射
    TA:RmRn,TA(x)=Ax,T_{\mathbf A}:\mathbb R^m\to\mathbb R^n,\qquad T_{\mathbf A}(\mathbf x)=\mathbf A\mathbf x,
    分别给出它为满射、为单射时 n,m,rn,m,r 应满足的条件。
  3. 定义 Moore–Penrose 伪逆
    A+=VΣ1UT\mathbf A^+=\mathbf V\mathbf\Sigma^{-1}\mathbf U^T
    以及
    B=ImA+A,TB(x)=Bx.\mathbf B=\mathbf I_m-\mathbf A^+\mathbf A,\qquad T_{\mathbf B}(\mathbf x)=\mathbf B\mathbf x.
    证明 Im(TB)\operatorname{Im}(T_{\mathbf B})ker(TA)\ker(T_{\mathbf A}) 线性同构。
  4. x1=Bx,x2=xx1,\mathbf x_1=\mathbf B\mathbf x,\qquad \mathbf x_2=\mathbf x-\mathbf x_1,
    证明 x=x1+x2\mathbf x=\mathbf x_1+\mathbf x_2 是正交分解。
  5. 给定 bRn\mathbf b\in\mathbb R^n,令 x0=A+b\mathbf x_0=\mathbf A^+\mathbf b,证明 x=x0\mathbf x=\mathbf x_0 最小化
    (Axb)T(Axb).(\mathbf A\mathbf x-\mathbf b)^T(\mathbf A\mathbf x-\mathbf b).
    可使用提示
    Axb=A(xx0)+(Ax0b).\mathbf A\mathbf x-\mathbf b =\mathbf A(\mathbf x-\mathbf x_0)+(\mathbf A\mathbf x_0-\mathbf b).

考点

  • 奇异值分解与谱关系:由薄 SVD 识别 ATA\mathbf A^T\mathbf A 的正特征值 σk2\sigma_k^2 及右奇异向量。
  • 线性映射的核与像:以秩 rr 和定义域、值域维数判定单射与满射,并刻画核空间。
  • Moore–Penrose 伪逆:证明 IA+A\mathbf I-\mathbf A^+\mathbf A 投影到 kerA\ker\mathbf A,并给出核空间与行空间的正交分解。
  • 最小二乘法:利用残差在 ImA\operatorname{Im}\mathbf A 上的正交性证明伪逆解使残差平方范数最小。

Kai

(1)

Given the singular value decomposition (SVD) of A\mathbf{A} as A=UΣVT\mathbf{A} = \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T, we can express ATA\mathbf{A}^T \mathbf{A} as follows:

ATA=(UΣVT)T(UΣVT)=VΣTUTUΣVT=VΣ2VT\mathbf{A}^T \mathbf{A} = (\mathbf{U} \mathbf{\Sigma} \mathbf{V}^T)^T (\mathbf{U} \mathbf{\Sigma} \mathbf{V}^T) = \mathbf{V} \mathbf{\Sigma}^T \mathbf{U}^T \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T = \mathbf{V} \mathbf{\Sigma}^2 \mathbf{V}^T

The matrix Σ2\mathbf{\Sigma}^2 is diagonal with the diagonal elements σk2\sigma_k^2 (k=1,,rk = 1, \ldots, r). Thus, the positive eigenvalues of ATA\mathbf{A}^T \mathbf{A} are exactly the σk2\sigma_k^2, and the associated normalized eigenvectors are the columns of V\mathbf{V}.

(2)

Surjective (onto): The mapping TA:RmRnT_{\mathbf{A}}: \mathbb{R}^m \to \mathbb{R}^n is surjective if the range of A\mathbf{A} spans Rn\mathbb{R}^n, i.e., A\mathbf{A} has full row rank. This occurs when r=nmr = n \leq m.

Injective (one-to-one): The mapping TAT_{\mathbf{A}} is injective if the kernel of A\mathbf{A} contains only the zero vector, i.e., A\mathbf{A} has full column rank. This occurs when r=mnr = m \leq n.

(3)

The pseudoinverse A+\mathbf{A}^+ is defined as A+=VΣ1UT\mathbf{A}^+ = \mathbf{V} \mathbf{\Sigma}^{-1} \mathbf{U}^T. Consider B=ImA+A\mathbf{B} = \mathbf{I}_m - \mathbf{A}^+ \mathbf{A}.

We need to show that Im(TB)\mathrm{Im}(T_{\mathbf{B}}) is isomorphic to Ker(TA)\mathrm{Ker}(T_{\mathbf{A}}). Observe the following:

BA=(ImA+A)A=AA+AA=AA=0\mathbf{B} \mathbf{A} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) \mathbf{A} = \mathbf{A} - \mathbf{A}^+ \mathbf{A} \mathbf{A} = \mathbf{A} - \mathbf{A} = \mathbf{0}

Thus, Im(B)Ker(A)\mathrm{Im}(\mathbf{B}) \subseteq \mathrm{Ker}(\mathbf{A}).

Now, consider xKer(A)\mathbf{x} \in \mathrm{Ker}(\mathbf{A}). Then Ax=0\mathbf{A} \mathbf{x} = \mathbf{0}, and

Bx=(ImA+A)x=x\mathbf{B} \mathbf{x} = (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) \mathbf{x} = \mathbf{x}

Thus, xIm(B)\mathbf{x} \in \mathrm{Im}(\mathbf{B}). Therefore, Im(B)=Ker(A)\mathrm{Im}(\mathbf{B}) = \mathrm{Ker}(\mathbf{A}).

(4)

Given x=x1+x2\mathbf{x} = \mathbf{x}_1 + \mathbf{x}_2 where x1=Bx\mathbf{x}_1 = \mathbf{B} \mathbf{x} and x2=xx1\mathbf{x}_2 = \mathbf{x} - \mathbf{x}_1:

x2=xBx=x(ImA+A)x=A+Ax\mathbf{x}_2 = \mathbf{x} - \mathbf{B} \mathbf{x} = \mathbf{x} - (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) \mathbf{x} = \mathbf{A}^+ \mathbf{A} \mathbf{x}

To show orthogonality:

x1Tx2=(Bx)T(A+Ax)=xTBTA+Ax\mathbf{x}_1^T \mathbf{x}_2 = (\mathbf{B} \mathbf{x})^T (\mathbf{A}^+ \mathbf{A} \mathbf{x}) = \mathbf{x}^T \mathbf{B}^T \mathbf{A}^+ \mathbf{A} \mathbf{x}

Since B\mathbf{B} is symmetric (B=ImA+A\mathbf{B} = \mathbf{I}_m - \mathbf{A}^+ \mathbf{A}):

xT(ImA+A)A+Ax=xT(A+AA+A)x=0\mathbf{x}^T (\mathbf{I}_m - \mathbf{A}^+ \mathbf{A}) \mathbf{A}^+ \mathbf{A} \mathbf{x} = \mathbf{x}^T (\mathbf{A}^+ \mathbf{A} - \mathbf{A}^+ \mathbf{A}) \mathbf{x} = \mathbf{0}

Thus, x1\mathbf{x}_1 and x2\mathbf{x}_2 are orthogonal.

(5)

Let x0=A+b\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b}. We need to show that x=x0\mathbf{x} = \mathbf{x}_0 minimizes the expression.

Consider the error:

Axb=A(xx0)+(Ax0b)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0) + (\mathbf{A} \mathbf{x}_0 - \mathbf{b})

Since x0=A+b\mathbf{x}_0 = \mathbf{A}^+ \mathbf{b}, we have Ax0=b\mathbf{A} \mathbf{x}_0 = \mathbf{b}, thus:

Axb=A(xx0)\mathbf{A} \mathbf{x} - \mathbf{b} = \mathbf{A} (\mathbf{x} - \mathbf{x}_0)

The norm to be minimized is:

(Axb)T(Axb)=(A(xx0))T(A(xx0))(\mathbf{A} \mathbf{x} - \mathbf{b})^T (\mathbf{A} \mathbf{x} - \mathbf{b}) = (\mathbf{A} (\mathbf{x} - \mathbf{x}_0))^T (\mathbf{A} (\mathbf{x} - \mathbf{x}_0))

This is minimized when x=x0\mathbf{x} = \mathbf{x}_0 since Ax0=b\mathbf{A} \mathbf{x}_0 = \mathbf{b} and A(xx0)=0\mathbf{A} (\mathbf{x} - \mathbf{x}_0) = \mathbf{0}.

Knowledge

奇异值分解 线性映射 广义逆矩阵 正交分解 线性代数

重点词汇

  • singular value decomposition (SVD) 奇异值分解
  • pseudoinverse 广义逆
  • surjective 满射
  • injective 单射
  • orthogonal decomposition 正交分解

参考资料

  1. "Linear Algebra and Its Applications" by Gilbert Strang, Chapter 7: The Singular Value Decomposition (SVD)
  2. "Matrix Computations" by Gene H. Golub and Charles F. Van Loan, Chapter 2: Matrix Analysis