Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Mahalanobis Distance

Understanding Covariance, Geometric Interpretation, and Bayesian Decision Theory

In this tutorial, we will explore the covariance matrix, its geometric interpretation, and its critical role in Bayesian Decision Theory, particularly in the context of the Mahalanobis Distance. We will derive the mathematical properties of these concepts and illustrate them with Python code and visualizations.

Resources: Adapted from Vision Dummy, SAS Blogs, Wikipedia, and Chapter 2 of Pattern Classification by Duda, Hart, and Stork Duda et al. (2000).


1. Why Do We Need Mahalanobis Distance?

In many real-world applications, we deal with multivariate data where features are often correlated and have different scales. While the Euclidean distance (∥x−y∥2\|\mathbf{x} - \mathbf{y}\|_2) is a common choice for measuring distances, it has a significant limitation: it assumes that all features are uncorrelated and have the same variance. This assumption rarely holds in practice.

Consider a dataset where one feature is measured in meters and another in kilometers. Euclidean distance would disproportionately weigh the feature with larger values. Similarly, if features are correlated, the “shape” of the data distribution is an ellipsoid rather than a sphere. Euclidean distance fails to account for this underlying structure, leading to misleading results in classification and clustering.

The Mahalanobis distance solves this by measuring the distance between a point and a distribution, normalized by the variance of each feature and the correlations between them.

Illustrative Example: Euclidean vs. Mahalanobis Distance

Let’s visualize this using a 2D Gaussian distribution with high correlation.

Source
<Figure size 1000x800 with 1 Axes>

Observation: Point 1 and Point 2 have the same Euclidean distance from the center, yet they lie in different “probability regions.” Point 2 is in a direction of lower variance (orthogonal to the main correlation), making it statistically less likely (further in Mahalanobis distance) than Point 1.


The effect of feature scaling

Case Study: Iranian Missile Similarity

In this example, we will use a dataset of Iranian Ballistic Missiles to solve a similarity problem. We want to determine which missile is more similar to the Khaybar Shikan (Missile A) based on Range and Speed.

1. The Dataset

We have the following data for three missiles:

MissileRange (km)Speed (Mach)
A. Khaybar Shikan1,4505
B. Sejjil2,00010
C. Fatāh1,40014

2. The Naive Approach (Raw Distance)

Using standard Euclidean distance, we calculate the distance between the Target (Khaybar) and the other missiles.

  • Khaybar vs. Sejjil (B):

    • ΔRange=∣2000−1450∣=550 km\Delta \text{Range} = |2000 - 1450| = 550 \text{ km}

    • ΔSpeed=∣10−5∣=5 Mach\Delta \text{Speed} = |10 - 5| = 5 \text{ Mach}

    • Draw=5502+52=302,500+25≈550.02D_{raw} = \sqrt{550^2 + 5^2} = \sqrt{302,500 + 25} \approx \mathbf{550.02}

  • Khaybar vs. Fatāh (C):

    • ΔRange=∣1400−1450∣=50 km\Delta \text{Range} = |1400 - 1450| = 50 \text{ km}

    • ΔSpeed=∣14−5∣=9 Mach\Delta \text{Speed} = |14 - 5| = 9 \text{ Mach}

    • Draw=502+92=2,500+81≈50.80D_{raw} = \sqrt{50^2 + 9^2} = \sqrt{2,500 + 81} \approx \mathbf{50.80}

❌ Problem: Fatāh looks much closer (50.8) than Sejjil (550.0). Reality Check: 50 km is negligible in missile range, but 9 Mach is a massive difference in speed class (Hypersonic vs. Medium). The raw calculation is dominated by the Range feature.

3. The Corrected Approach (Standard Scaling)

To fix this, we normalize the data using Z-scores (Z=x−μσZ = \frac{x - \mu}{\sigma}). This puts both features on the same scale.

Dataset Statistics (from Code Output):

  • Range: Mean = 1,616.67 km | Std Dev (Sample) ≈\approx 332.92 km

  • Speed: Mean = 9.67 Mach | Std Dev (Sample) ≈\approx 4.51 Mach

Note: The StandardScaler algorithm typically uses Population Std Dev (approx 272 km / 3.7 Mach) for internal transformation to ensure consistency, which results in the final distances below.

A. Khaybar vs. Sejjil (B):

  • Dscaled=1.902+1.402≈2.44D_{scaled} = \sqrt{1.90^2 + 1.40^2} \approx \mathbf{2.44}

B. Khaybar vs. Fatāh (C):

  • Dscaled=0.202+2.402≈2.45D_{scaled} = \sqrt{0.20^2 + 2.40^2} \approx \mathbf{2.45}

4. The Result

ComparisonRaw DistanceScaled Distance
Khaybar vs. Sejjil550.022.44 ✅
Khaybar vs. Fatāh50.802.45

✅ Result: Sejjil (2.44) is now correctly identified as the closest match to Khaybar, not Fatāh.

5. Why the Change?

  1. Raw Data: The Range feature (km) was numerically huge (1400-2000) compared to Speed (5-14). The algorithm ignored Speed because a change of 50 km is a tiny percentage of the total range, but a change of 9 Mach is a 180% change in Speed.

  2. Scaled Data: By dividing by the Standard Deviation, we turned Speed into the dominant factor.

  3. Conclusion: The Scaled Distance penalizes the massive speed difference in Fatāh, correctly identifying Sejjil as the more similar missile platform.


Code Logic Summary

The Python script used StandardScaler to handle this automatically. It:

  1. Calculated the Mean and Standard Deviation for each column.

  2. Transformed the values into Z-scores.

  3. Calculated Euclidean distance on the transformed (standardized) data.

This ensures features with different units (km vs Mach) contribute equally to the similarity score.

Source
=== RAW EUCLIDEAN DISTANCE (Unscaled) ===
Distance to Sejjil: 550.02
Distance to Fatāh: 50.80

=== SCALED EUCLIDEAN DISTANCE (Normalized) ===
Distance to Sejjil: 2.44
Distance to Fatāh: 2.45

--- ANALYSIS ---
Raw Distance Winner (Misleading): Fatāh
Scaled Distance Winner (Correct): Sejjil

--- DATASET STATISTICS ---
Range Std Dev: 332.92 km
Speed Std Dev: 4.51 Mach

Key Takeaway

When building ML models, Unit Mismatch is a silent killer.

Always scale your features (Standardization/Normalization) before calculating distances, training Neural Networks, or using K-Means Clustering.

2. Introduction to Covariance Matrix

The covariance matrix is a square matrix that summarizes the variances and covariances of a set of random variables. For a dataset with dd features, the covariance matrix Σ\boldsymbol{\Sigma} is a d×dd \times d matrix defined as:

Σ=[σ11σ12⋯σ1dσ21σ22⋯σ2d⋮⋮⋱⋮σd1σd2⋯σdd]\boldsymbol{\Sigma} = \begin{bmatrix} \sigma_{11} & \sigma_{12} & \cdots & \sigma_{1d} \\ \sigma_{21} & \sigma_{22} & \cdots & \sigma_{2d} \\ \vdots & \vdots & \ddots & \vdots \\ \sigma_{d1} & \sigma_{d2} & \cdots & \sigma_{dd} \end{bmatrix}

Where:

  • σii=Var⁡(xi)\sigma_{ii} = \operatorname{Var}(x_i): The variance of the ii-th feature (diagonal elements).

  • σij=Cov⁡(xi,xj)\sigma_{ij} = \operatorname{Cov}(x_i, x_j): The covariance between the ii-th and jj-th features (off-diagonal elements).

Properties:

  1. Symmetric: σij=σji\sigma_{ij} = \sigma_{ji}, so Σ=Σ⊤\boldsymbol{\Sigma} = \boldsymbol{\Sigma}^{\top}.

  2. Positive Semi-Definite: All eigenvalues are non-negative (λi≥0\lambda_i \geq 0).


3. Geometric Interpretation of Covariance Matrix

The covariance matrix defines the shape and orientation of the data distribution in the feature space.

  1. Eigenvectors (Directions): The eigenvectors of Σ\boldsymbol{\Sigma} represent the principal axes of the data distribution. They indicate the directions of maximum variance.

  2. Eigenvalues (Magnitudes): The corresponding eigenvalues represent the magnitude of variance along these eigenvector directions.

  3. Ellipsoid: In 2D, the covariance matrix defines an ellipse. In dd-dimensions, it defines a hyper-ellipsoid. The axes of this ellipsoid are aligned with the eigenvectors, and the lengths of the axes are proportional to the square roots of the eigenvalues.


Whitening Transformation

Whitening is a process that transforms data so that its covariance matrix becomes the identity matrix. This transformation removes correlations between features and scales the features to have unit variance. Whitening is particularly useful in machine learning and signal processing because it simplifies the structure of the data, making it easier to apply algorithms that assume uncorrelated features.

Figure 2.8 of Duda

Figure 2.8 Duda et al. (2000): The action of a linear transformation on the feature space will convert an arbitrary normal distribution into another normal distribution.

In the following we will see that a linear transformation W ⁣⊤W^{\!\top} applied to a zero-mean vector x−μ\mathbf{x}-\boldsymbol{\mu} produces a new vector whose covariance is obtained by “sandwiching” the old covariance between W ⁣⊤W^{\!\top} and WW. This property is fundamental to whitening, where we choose WW specifically so that the new covariance becomes the Identity matrix I\mathbf{I}.

Example: Whitening in Python

The following Python code demonstrates how to whiten a dataset:

Source
<Figure size 1200x600 with 2 Axes>

4. Linear Transformations and Covariance

To fully understand Mahalanobis distance and whitening, we must understand how covariance changes under linear transformations. The following theorem provides the mathematical foundation for transforming data to a “whitened” space.

Theorem 1 (Covariance Transformation Property)

Let x∈Rd\mathbf{x} \in \mathbb{R}^d be a random vector with mean μ=E[x]\boldsymbol{\mu} = \mathbb{E}[\mathbf{x}] and covariance matrix Σ=Cov⁡(x)\boldsymbol{\Sigma} = \operatorname{Cov}(\mathbf{x}). Let W∈Rd×dW \in \mathbb{R}^{d \times d} be a constant transformation matrix. If the transformed random vector y\mathbf{y} defines as y=W ⁣⊤(x−μ)\mathbf{y} = W^{\!\top}\bigl(\mathbf{x} - \boldsymbol{\mu}\bigr); then, the covariance matrix of the transformed vector y\mathbf{y} be as follows:

Cov⁡(y)=W ⁣⊤ΣW\boxed{\operatorname{Cov}(\mathbf{y}) = W^{\!\top}\boldsymbol{\Sigma}W}

Proof: We proceed by applying the definition of covariance and the linearity of expectation.

  1. Mean of y\mathbf{y}: First, observe that the mean of the centered vector is zero:

    E[y]=E[W ⁣⊤(x−μ)]=W ⁣⊤(E[x]−μ)=W ⁣⊤(μ−μ)=0.\begin{aligned} \mathbb{E}[\mathbf{y}] &= \mathbb{E}\bigl[W^{\!\top}(\mathbf{x} - \boldsymbol{\mu})\bigr] \\ &= W^{\!\top}\bigl(\mathbb{E}[\mathbf{x}] - \boldsymbol{\mu}\bigr) \\ &= W^{\!\top}(\boldsymbol{\mu} - \boldsymbol{\mu}) \\ &= \mathbf{0}. \end{aligned}
  2. Covariance of y\mathbf{y}:

    Cov(y)=E[(y−E[y])(y−E[y])T]\text{Cov}(\mathbf{y}) = \mathbb{E}[(\mathbf{y} - \mathbb{E}[\mathbf{y}])(\mathbf{y} - \mathbb{E}[\mathbf{y}])^T]

    Since E[y]=0\mathbb{E}[\mathbf{y}] = \mathbf{0}, the covariance definition simplifies to Cov⁡(y)=E[yy ⁣⊤]\operatorname{Cov}(\mathbf{y}) = \mathbb{E}[\mathbf{y}\mathbf{y}^{\!\top}].

    Cov⁡(y)=E[y y ⁣⊤]=E[W ⁣⊤(x−μ) (x−μ) ⁣⊤W].\begin{aligned} \operatorname{Cov}(\mathbf{y}) &= \mathbb{E}\bigl[\mathbf{y}\,\mathbf{y}^{\!\top}\bigr] \\ &= \mathbb{E}\Big[W^{\!\top}(\mathbf{x}-\boldsymbol{\mu})\,(\mathbf{x}-\boldsymbol{\mu})^{\!\top}W\Big]. \end{aligned}

    By the linearity of the expectation operator, constant matrices W ⁣⊤W^{\!\top} and WW can be pulled out of the expectation:

    Cov⁡(y)=W ⁣⊤  E[(x−μ)(x−μ) ⁣⊤]  W.\begin{aligned} \operatorname{Cov}(\mathbf{y}) &= W^{\!\top}\;\mathbb{E}\Big[(\mathbf{x}-\boldsymbol{\mu})(\mathbf{x}-\boldsymbol{\mu})^{\!\top}\Big]\;W. \end{aligned}

    Recognizing that the term inside the expectation is exactly the definition of Cov⁡(x)=Σ\operatorname{Cov}(\mathbf{x}) = \boldsymbol{\Sigma}:

    Cov⁡(y)=W ⁣⊤ΣW.□\operatorname{Cov}(\mathbf{y}) = W^{\!\top}\boldsymbol{\Sigma}W. \quad \square

Remark (Whitening Application): This property is fundamental to whitening. By choosing WW such that W ⁣⊤ΣW=IW^{\!\top}\boldsymbol{\Sigma}W = \mathbf{I}, we transform the data into a space where features are uncorrelated and have unit variance. As shown in Section 5, this WW is constructed using the eigendecomposition of Σ\boldsymbol{\Sigma}.

5. Whitening Transformation

Whitening is a preprocessing step that transforms data so that its covariance matrix becomes the Identity matrix (I\mathbf{I}). This implies the transformed features are uncorrelated and have unit variance.

Mathematical Formulation

We seek a transformation matrix WW such that if y=W⊤(x−μ)\mathbf{y} = W^{\top}(\mathbf{x} - \boldsymbol{\mu}), then Cov⁡(y)=I\operatorname{Cov}(\mathbf{y}) = \mathbf{I}.

From the proof in Section 4:

Cov⁡(y)=W ⁣⊤ΣW=I\operatorname{Cov}(\mathbf{y}) = W^{\!\top}\boldsymbol{\Sigma}W = \mathbf{I}

We can solve for WW using the Eigen Decomposition of Σ\boldsymbol{\Sigma}. Let Σ=ΦΛΦ⊤\boldsymbol{\Sigma} = \Phi \Lambda \Phi^{\top}, where:

  • Φ\Phi is the orthogonal matrix of eigenvectors.

  • Λ\Lambda is the diagonal matrix of eigenvalues.

We choose W=ΦΛ−1/2W = \Phi \Lambda^{-1/2}. Let’s verify:

Cov⁡(y)=(ΦΛ−1/2)⊤Σ(ΦΛ−1/2)=Λ−1/2Φ⊤(ΦΛΦ⊤)ΦΛ−1/2=Λ−1/2(Φ⊤Φ)Λ(Φ⊤Φ)Λ−1/2=Λ−1/2(I)Λ(I)Λ−1/2=I.\begin{aligned} \operatorname{Cov}(\mathbf{y}) &= (\Phi \Lambda^{-1/2})^{\top} \boldsymbol{\Sigma} (\Phi \Lambda^{-1/2}) \\ &= \Lambda^{-1/2} \Phi^{\top} (\Phi \Lambda \Phi^{\top}) \Phi \Lambda^{-1/2} \\ &= \Lambda^{-1/2} (\Phi^{\top} \Phi) \Lambda (\Phi^{\top} \Phi) \Lambda^{-1/2} \\ &= \Lambda^{-1/2} (\mathbf{I}) \Lambda (\mathbf{I}) \Lambda^{-1/2} \\ &= \mathbf{I}. \end{aligned}

Geometric Interpretation

  1. Rotation (Φ⊤\Phi^{\top}): Aligns the data with the principal axes (removes correlations).

  2. Scaling (Λ−1/2\Lambda^{-1/2}): Stretches or compresses the data along these axes so that all axes have length 1 (unit variance).

Result: The data cloud becomes a perfect sphere (isotropic). In this whitened space, the Euclidean distance is equivalent to the Mahalanobis distance in the original space.


6. Mahalanobis Distance

The Mahalanobis distance is a measure of the distance between a point x\mathbf{x} and a distribution with mean μ\boldsymbol{\mu} and covariance matrix Σ\mathbf{\Sigma}. It is defined as:

DM(x,μ)=(x−μ)TΣ−1(x−μ)D_M(\mathbf{x}, \boldsymbol{\mu}) = \sqrt{(\mathbf{x} - \boldsymbol{\mu})^T \mathbf{\Sigma}^{-1} (\mathbf{x} - \boldsymbol{\mu})}

This distance is central to many multivariate statistical methods, including Bayesian decision theory for classification, because it accounts for the underlying structure (variances and correlations) of the data.

The Key Insight

The covariance transformation formula from Section 4 is the fundamental mathematical foundation for both whitening transformation and Mahalanobis distance.

Recall the Mahalanobis distance:

DM(x)=(x−μ)TΣ−1(x−μ)D_M(\mathbf{x}) = \sqrt{(\mathbf{x} - \boldsymbol{\mu})^T \boldsymbol{\Sigma}^{-1} (\mathbf{x} - \boldsymbol{\mu})}

Now consider the whitened space: If we define y=W ⁣⊤(x−μ)\mathbf{y} = W^{\!\top}(\mathbf{x}-\boldsymbol{\mu}) where WW makes Cov⁡(y)=I\operatorname{Cov}(\mathbf{y}) = \mathbf{I}, then:

∥y∥2=yTy=(x−μ)TWW ⁣⊤(x−μ)\|\mathbf{y}\|_2 = \sqrt{\mathbf{y}^T \mathbf{y}} = \sqrt{(\mathbf{x} - \boldsymbol{\mu})^T W W^{\!\top} (\mathbf{x} - \boldsymbol{\mu})}

For this to equal the Mahalanobis distance, we need to show that WW ⁣⊤=Σ−1 W W^{\!\top} = \boldsymbol{\Sigma}^{-1} .

Proof:

We are given: Σ=ΦΛΦ⊤ \boldsymbol{\Sigma} = \Phi \Lambda \Phi^{\top} and W=ΦΛ−1/2 W = \Phi \Lambda^{-1/2}

We aim to show that WW ⁣⊤=Σ−1W W^{\!\top} = \boldsymbol{\Sigma}^{-1}

Step 1: Compute WW ⁣⊤ W W^{\!\top}

Substitute W=ΦΛ−1/2 W = \Phi \Lambda^{-1/2} :

WW ⁣⊤=(ΦΛ−1/2)(ΦΛ−1/2) ⁣⊤W W^{\!\top} = (\Phi \Lambda^{-1/2}) (\Phi \Lambda^{-1/2})^{\!\top}

Now, transpose the second factor:

(ΦΛ−1/2) ⁣⊤=(Λ−1/2) ⁣⊤Φ ⁣⊤(\Phi \Lambda^{-1/2})^{\!\top} = (\Lambda^{-1/2})^{\!\top} \Phi^{\!\top}

Since Λ \Lambda is a diagonal matrix with positive eigenvalues (implying invertibility), Λ−1/2 \Lambda^{-1/2} is also diagonal and symmetric, so:

(Λ−1/2) ⁣⊤=Λ−1/2(\Lambda^{-1/2})^{\!\top} = \Lambda^{-1/2}

Thus,

WW ⁣⊤=ΦΛ−1/2(Λ−1/2)Φ ⁣⊤=Φ(Λ−1/2⋅Λ−1/2)Φ ⁣⊤=ΦΛ−1Φ ⁣⊤W W^{\!\top} = \Phi \Lambda^{-1/2} (\Lambda^{-1/2}) \Phi^{\!\top} = \Phi (\Lambda^{-1/2} \cdot \Lambda^{-1/2}) \Phi^{\!\top} = \Phi \Lambda^{-1} \Phi^{\!\top}

Step 2: Compute Σ−1 \boldsymbol{\Sigma}^{-1} using the inverse of a diagonalized matrix

We are given:

Σ=ΦΛΦ ⁣⊤\boldsymbol{\Sigma} = \Phi \Lambda \Phi^{\!\top}

Since Φ \Phi is an orthogonal matrix (i.e., Φ ⁣⊤Φ=I \Phi^{\!\top} \Phi = I ), we know that:

(ΦΛΦ ⁣⊤)−1=(Φ ⁣⊤)−1Λ−1Φ−1(\Phi \Lambda \Phi^{\!\top})^{-1} = (\Phi^{\!\top})^{-1} \Lambda^{-1} \Phi^{-1}

Because Φ \Phi is orthogonal:

  • Φ−1=Φ ⁣⊤ \Phi^{-1} = \Phi^{\!\top}

  • So (Φ ⁣⊤)−1=Φ (\Phi^{\!\top})^{-1} = \Phi

Therefore,

(ΦΛΦ ⁣⊤)−1=ΦΛ−1Φ ⁣⊤(\Phi \Lambda \Phi^{\!\top})^{-1} = \Phi \Lambda^{-1} \Phi^{\!\top}

We have shown that:

WW ⁣⊤=Σ−1\boxed{W W^{\!\top} = \boldsymbol{\Sigma}^{-1}}

Key insight: The orthogonality of Φ \Phi allows us to simplify the inverse of ΦΛΦ ⁣⊤ \Phi \Lambda \Phi^{\!\top} elegantly, which is essential for establishing the result.

Key Relationship

ConceptMathematical ExpressionPurpose
Covariance TransformationCov⁡(y)=W ⁣⊤ΣW\operatorname{Cov}(\mathbf{y}) = W^{\!\top}\boldsymbol{\Sigma}WDescribes how covariance changes under linear transformation
Whitening MatrixW=ΦΛ−1/2W = \Phi \Lambda^{-1/2}Transforms data to unit covariance
Mahalanobis DistanceDM=(x−μ)TΣ−1(x−μ)D_M = \sqrt{(\mathbf{x}-\boldsymbol{\mu})^T \boldsymbol{\Sigma}^{-1}(\mathbf{x}-\boldsymbol{\mu})}Distance accounting for covariance structure
Euclidean Distance (Whitened)∣y∣2=yTy|\mathbf{y}|_2 = \sqrt{\mathbf{y}^T\mathbf{y}}Distance after covariance normalization
Key EquivalenceDM(x)=∣y∣2D_M(\mathbf{x}) = |\mathbf{y}|_2Mahalanobis = Euclidean in whitened space

7. Python Implementation: Step-by-Step Demonstration

======================================================================
STEP 1: Generate Correlated Data (Original Space)
======================================================================
Original Covariance:
[[4 3]
 [3 4]]

======================================================================
STEP 2: Eigen Decomposition (Rotation + Scaling Factors)
======================================================================

Eigenvalues (λ): [7. 1.]

Eigenvectors (Φ):
[[ 0.70710678 -0.70710678]
 [ 0.70710678  0.70710678]]

======================================================================
STEP 3: Rotation Only
======================================================================

Covariance after Rotation (should be diagonal):
[[6.54609064 0.06692453]
 [0.06692453 0.98399756]]
Off-diagonal before rotation: 3.0000
Off-diagonal after rotation: 0.0669

======================================================================
STEP 4: Scaling Only (Λ^(-1/2) normalizes variance)
======================================================================

Covariance after Scaling (should be Identity):
[[0.93515581 0.02529509]
 [0.02529509 0.98399756]]

======================================================================
STEP 5:  Whitening (W = ΦΛ^(-1/2))
======================================================================

Whitening Matrix (W = ΦΛ^(-1/2)):
[[ 0.26726124 -0.70710678]
 [ 0.26726124  0.70710678]]

Covariance after Whitening (W^T Σ W):
[[0.93515581 0.02529509]
 [0.02529509 0.98399756]]

======================================================================
STEP 6: Visualization of Transformation
======================================================================
<Figure size 1400x1200 with 4 Axes>

======================================================================
STEP 7: Mahalanobis Distance = Euclidean in Whitened Space
======================================================================

Test Point: [4 4]
Mahalanobis Distance (Original Space): 1.5119
Euclidean Distance (Whitened Space): 1.5119

✓ Mahalanobis Distance = Euclidean Distance in Whitened Space!

8. Covariance Matrix in Bayesian Decision Theory

p(x∣ωi)=1(2π)d/2∣Σi∣1/2exp⁡(−12(x−μi)TΣi−1(x−μi))p(\mathbf{x}|\omega_i) = \frac{1}{(2\pi)^{d/2}|\boldsymbol{\Sigma}_i|^{1/2}} \exp\left(-\frac{1}{2}(\mathbf{x} - \boldsymbol{\mu}_i)^T \boldsymbol{\Sigma}_i^{-1} (\mathbf{x} - \boldsymbol{\mu}_i)\right)
  • μi\boldsymbol{\mu}_i: Mean vector for class ωi\omega_i.

  • Σi\boldsymbol{\Sigma}_i: Covariance matrix for class ωi\omega_i.

The term in the exponent is exactly half the squared Mahalanobis distance. The covariance matrix determines the “tightness” and orientation of the decision boundaries.


9. Mahalanobis Distance and Its Role

The Mahalanobis distance measures the distance between a point x\mathbf{x} and a distribution (characterized by mean μ\boldsymbol{\mu} and covariance Σ\boldsymbol{\Sigma}).

Connection to Standardization (Z-Score)

To build intuition, compare this to the 1D Z-score:

z=x−μσ  ⟹  z2=(x−μ)1σ2(x−μ)z = \frac{x - \mu}{\sigma} \implies z^2 = (x-\mu)\frac{1}{\sigma^2}(x-\mu)
  • 1D: We normalize by σ2\sigma^2 (variance).

  • d-Dimensions: We normalize by Σ\boldsymbol{\Sigma} (covariance matrix).

The inverse covariance matrix Σ−1\boldsymbol{\Sigma}^{-1} effectively:

  1. Scales features by their variance (normalizing units).

  2. Rotates the space to decorrelate features.

Mahalanobis Distance Intuition

If the data distribution is a cloud of points shaped like an ellipsoid, Mahalanobis distance measures how many “standard deviations away” a point is, taking into account the stretching and rotation of that ellipsoid.

The data cloud becomes a perfect sphere (isotropic). In this whitened space, the Euclidean distance is equivalent to the Mahalanobis distance in the original space.

Is Whitening Equivalent to Per-Feature Normalization?

No, they are NOT equivalent. This is a critical distinction.

Per-Feature Normalization (Standardization)

xi′=xi−μiσix_i' = \frac{x_i - \mu_i}{\sigma_i}
  • Scales each feature independently to unit variance

  • Preserves correlations between features

Whitening

y=W ⁣⊤(x−μ),W=ΦΛ−1/2\mathbf{y} = W^{\!\top}(\mathbf{x} - \boldsymbol{\mu}), \quad W = \Phi \Lambda^{-1/2}
  • Rotates AND scales the entire space

  • Removes correlations and normalizes variance

  • Covariance becomes identity matrix I\mathbf{I}

  • Data cloud becomes a perfect sphere

Key Difference

PropertyNormalizationWhitening
Unit Variance✅✅
Zero Correlation❌✅
Identity Covariance❌✅
Euclidean = Mahalanobis❌✅

Why It Matters

For Mahalanobis distance and Bayesian decision theory, only whitening ensures that Euclidean distance in the transformed space equals Mahalanobis distance in the original space. Per-feature normalization alone is insufficient.


10. Comparison: PCA vs. Whitening

Principal Component Analysis (PCA) and Whitening are closely related linear transformations based on eigendecomposition of the covariance matrix. However, they have distinct goals and effects on the data distribution.

Mathematical Relationship

Both transformations start with the eigen decomposition of the covariance matrix Σ=ΦΛΦ ⁣⊤\boldsymbol{\Sigma} = \Phi \Lambda \Phi^{\!\top}.

  1. PCA (Rotation Only): Projects data onto the principal axes (eigenvectors). It removes correlations but preserves the variance scale.

    zPCA=Φ ⁣⊤(x−μ)\mathbf{z}_{\text{PCA}} = \Phi^{\!\top}(\mathbf{x} - \boldsymbol{\mu})

    Resulting Covariance:

    Cov⁡(zPCA)=Λ=diag⁡(λ1,λ2,…,λd)\operatorname{Cov}(\mathbf{z}_{\text{PCA}}) = \Lambda = \operatorname{diag}(\lambda_1, \lambda_2, \dots, \lambda_d)
  2. Whitening (Rotation + Scaling): Applies PCA rotation followed by a scaling step to normalize the variance of each component to 1.

    yWhite=Λ−1/2Φ ⁣⊤(x−μ)=Λ−1/2zPCA\mathbf{y}_{\text{White}} = \Lambda^{-1/2} \Phi^{\!\top}(\mathbf{x} - \boldsymbol{\mu}) = \Lambda^{-1/2} \mathbf{z}_{\text{PCA}}

    Resulting Covariance:

    Cov⁡(yWhite)=I\operatorname{Cov}(\mathbf{y}_{\text{White}}) = \mathbf{I}

Key Differences

FeaturePCAWhitening
Primary GoalDimensionality Reduction / Maximize VarianceNormalization of Variance & Correlation
TransformationRotation (Φ ⁣⊤\Phi^{\!\top})Rotation (Φ ⁣⊤\Phi^{\!\top}) + Scaling (Λ−1/2\Lambda^{-1/2})
Output Varianceλi\lambda_i (Eigenvalues)1 (Unit Variance)
Output Correlation0 (Uncorrelated)0 (Uncorrelated)
Covariance MatrixDiagonal (Λ\Lambda)Identity (I\mathbf{I})
Geometric ShapeAxis-aligned EllipsePerfect Sphere
Distance MetricEuclidean ≠\neq MahalanobisEuclidean == Mahalanobis
Information LossPossible (if k<dk < d)None (usually full rank)

Visual Demonstration: PCA vs. Whitening

Source

======================================================================
COMPARISON: Original → PCA → Whitening
======================================================================
<Figure size 1800x600 with 3 Axes>

======================================================================
Distance Metrics for a Test Point
======================================================================

Test Point: [2. 1.]
Mahalanobis Distance (Original): 1.0690
Euclidean Distance (PCA Space):  2.2361
Euclidean Distance (White Space): 1.0690

✅ Conclusion: Only Whitening preserves Mahalanobis Distance as Euclidean Distance.

Why This Matters for Mahalanobis Distance

  1. PCA aligns the data with the principal axes but preserves the different variances along those axes. If you measure Euclidean distance in PCA space, a point 2 units away along a high-variance axis is treated the same as a point 2 units away along a low-variance axis. This distorts the statistical probability.

  2. Whitening normalizes the variances. In this space, 1 unit equals 1 standard deviation in every direction. Therefore, Euclidean distance in whitened space correctly represents the statistical “surprise” (Mahalanobis distance) of a point.

Conclusion

The Mahalanobis distance is the natural generalization of the Z-score to multivariate dimensions. By incorporating the covariance matrix, it provides a statistically valid measure of distance that accounts for feature scales and correlations. This makes it indispensable for:

  • Anomaly Detection: Identifying points that are statistically unlikely.

  • Classification: Bayesian decision boundaries rely on this distance metric.

  • Preprocessing: Whitening transforms data to simplify downstream algorithms (e.g., PCA, Neural Networks).


Key Takeaways

  • Euclidean Distance assumes spherical data; Mahalanobis Distance assumes elliptical data.

  • The Covariance Matrix Σ\boldsymbol{\Sigma} captures the shape, scale, and orientation of the data distribution.

  • Whitening transforms Σ\boldsymbol{\Sigma} to I\mathbf{I}, effectively “spherifying” the data cloud.

  • Linear Transformation Property: If y=W⊤(x−μ)\mathbf{y} = W^{\top}(\mathbf{x}-\boldsymbol{\mu}), then Cov⁡(y)=W⊤ΣW\operatorname{Cov}(\mathbf{y}) = W^{\top}\boldsymbol{\Sigma}W.

  • In the Whitened Space, Euclidean distance equals Mahalanobis distance.


References

Duda et al. (2000)

Duda, R. O., Hart, P. E., & Stork, D. G. (2000). Pattern Classification (2nd ed.). Wiley-Interscience.

References
  1. Duda, R. O., Hart, P. E., & Stork, D. G. (2000). Pattern Classification (2nd Edition) (2nd ed.). Wiley-Interscience. https://file.fouladi.ir/courses/pr/books/%5BDuda%5D_PatternClassification.pdf