Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Feature Maps: Bridging to Kernel Methods

Feature Maps: Bridging to Kernel Methods

1. Introduction to Feature Maps

A feature map ϕ\phi transforms input data into a higher-dimensional space:

ϕ:Rp→Rdwhere d≫p\phi: \mathbb{R}^p \rightarrow \mathbb{R}^d \quad \text{where } d \gg p

Key Motivation: Enable linear models to solve non-linear problems by:

  • Explicit mapping for classification/clustering

  • Basis expansion for regression

For advanced feature creation techniques, see Feature-Engine’s MathFeatures.

2. Classification: Two-Ring Problem

2.1 Original 2D Space (Linear Failure)

Source
<Figure size 1500x600 with 2 Axes>

The map ϕ(x,y)=[x,y,x2+y2]\phi(x,y) = [x, y, x^2+y^2] makes classes linearly separable by converting radial distance to a linear feature.

Output
Loading...
Source
<Figure size 1000x700 with 1 Axes>
Source
Loading...

3. Clustering: Two-Ring Problem

Source
<Figure size 1200x600 with 2 Axes>

Regression Example: Using Feature Maps for Non-Linear Data

Goal: To demonstrate how mapping features to a higher-dimensional space can allow a linear model to fit non-linear data.

Motivation: Standard linear regression models assume a linear relationship between the features (independent variables) and the target (dependent variable). What happens when the underlying relationship is non-linear? A simple linear model will perform poorly.

Consider a scenario where the data follows a curve, for example, a quadratic relationship. A straight line (from linear regression) won’t capture this curve effectively.

Solution Idea: We can transform the original features into a higher-dimensional space where the relationship becomes linear. This transformation is called a feature map, denoted by Φ.

Let’s illustrate this with an example.

1. Generating Non-Linear Data

We’ll create synthetic data where y depends quadratically on x, plus some random noise to make it more realistic.

Source
<Figure size 800x600 with 1 Axes>

As you can see, the data clearly follows a curve, not a straight line.

2. Attempting Simple Linear Regression (Original Space)

Let’s see how a standard linear regression model performs on this data without any feature transformation.

Source
<Figure size 800x600 with 1 Axes>
Linear Regression Score (R^2): 0.4260

The linear model tries its best to fit a straight line through the curved data, but it’s clearly a poor fit. The R² score will likely be low, indicating that the model doesn’t explain much of the variance in the data.

3. Applying a Feature Map (Polynomial Features)

Now, let’s apply a feature map. Since we know the underlying relationship is quadratic (y ≈ ax² + bx + c), a suitable feature map Φ would transform our single feature x into two features: x and x². So, Φ(x) = [x, x²].

We can achieve this using Scikit-Learn’s PolynomialFeatures.

Original X (first 5 samples):
 [[-0.75275929]
 [ 2.70428584]
 [ 1.39196365]
 [ 0.59195091]
 [-2.06388816]]

Transformed X_poly (first 5 samples) [x, x^2]:
 [[-0.75275929  0.56664654]
 [ 2.70428584  7.3131619 ]
 [ 1.39196365  1.93756281]
 [ 0.59195091  0.35040587]
 [-2.06388816  4.25963433]]

Notice that our data is now represented in a 2-dimensional feature space [x, x²].

4. Linear Regression in the Higher-Dimensional Feature Space

Now, we train a linear regression model, but using the transformed features (X_poly). The model will learn weights for both x and x², effectively fitting a model of the form: y = w₁*x + w₂*x² + b where w₁, w₂ are the weights (coefficients) and b is the intercept (bias). This is a linear model with respect to the new features x and x².

Original linear model:

y^=wTx=∑i=1pwixi\hat{y} = \mathbf{w}^T\mathbf{x} = \sum_{i=1}^{p} w_i x_i

A transformation ϕ\phi that maps features to higher dimensions:

ϕ:Rp→Rd(d≫p)\phi: \mathbb{R}^p \rightarrow \mathbb{R}^d \quad (d \gg p)

New model becomes:

y^=wTϕ(x)\hat{y} = \mathbf{w}^T\phi(\mathbf{x})

For p=1, degree=2:

[1,x1]→ϕ[1,x1,x12][1, x_1] \xrightarrow{\phi} [1, x_1, x_1^2]
Source
<Figure size 800x600 with 1 Axes>
Polynomial Regression Score (R^2): 0.8525
Coefficients (w1, w2): [[0.93366893 0.56456263]]
Intercept (b): [1.78134581]

5. Conclusion

By mapping the original single feature x to a higher-dimensional space [x, x²], we enabled a standard linear regression model to perfectly capture the non-linear (quadratic) relationship present in the data. The resulting fit is much better, as indicated visually and by the significantly higher R² score.

This example illustrates the power of feature maps: transforming data into a space where linear models can become effective, even if the original relationship was non-linear.

Transition to Kernel Trick: While this explicit feature mapping works well for low-dimensional data and simple transformations, calculating and storing these higher-dimensional features can become computationally expensive or even infeasible if the original data or the target feature space is very high-dimensional (e.g., using polynomial features of a very high degree or other complex maps). The Kernel Trick provides a mathematical shortcut to achieve the same result as working in the high-dimensional feature space without explicitly computing the coordinates of the data in that space. It allows algorithms that only depend on dot products between data points (like Support Vector Machines) to operate implicitly in the high-dimensional feature space, making complex mappings computationally tractable. This regression example helps build the intuition for why such implicit mappings are desirable.

Notebook Cell
Notebook Cell