Supervised LearningNon-Linear RegressionRegression

Decision Tree Regressor (CART)

Primary task · Regression

Decision Tree Regressor (CART) applies the Decision Trees (CART) learning mechanism to continuous targets, producing numeric predictions instead of class labels.

← Directory
Visual intuition

From data to learned behaviour

A decision tree repeatedly asks simple feature questions. Each question divides the data into more homogeneous regions until leaves contain sufficiently similar targets. The final model is therefore both a prediction rule and a visible hierarchy of decisions.

Infographic
1Training data2Best split3Branches4Leaves5PredictionTraining transforms evidence into a reusable model state
Conceptual simulation

Watch the learning mechanism form

The structure below is synchronized with the same training state used by the prediction simulation.

Mechanism view
Training control centre

Control both simulations together

Reset regenerates the synthetic data and model state. Train animates to completion. Pause freezes the animation. Train Step advances one learning stage.

Step 0 / 8
Model simulation

Inspect the learned prediction / representation

Synthetic data are generated locally in your browser.

Model description

Understand Decision Tree Regressor (CART) after watching it learn

This section connects the animation to the actual statistical or computational idea behind the model.

Deep description

Decision Tree Regressor (CART) Decision Tree Regressor (CART) applies the Decision Trees (CART) learning mechanism to continuous targets, producing numeric predictions instead of class labels.

What is learned. During training, the algorithm builds or adjusts feature thresholds, branches and terminal leaf values. The core learning mechanism is: Recursively partitions feature space into axis-aligned rectangular regions using purity split criteria like Gini Impurity or Information Gain.

How training becomes inference. Place all samples at the root → score candidate splits → choose the strongest split → recurse into child nodes → stop according to complexity rules → predict from terminal leaves. Once training stops, the fitted state is reused on unseen inputs rather than being reconstructed from scratch. The resulting output is: A continuous numeric prediction; some probabilistic variants can also provide uncertainty or intervals.

Why practitioners use it. Completely transparent white-box interpretability, handles numerical and categorical features naturally, no feature scaling required. Typical fits include House-price estimation, insurance severity, demand estimation, interpretable non-linear numeric prediction.

What to verify before trusting it. High variance; prone to severe overfitting on noisy data without pruning or depth constraints. The visual simulation is intentionally simplified, so real use should still validate preprocessing, data independence, hyperparameters, uncertainty and task-appropriate metrics.

Internal statefeature thresholds, branches and terminal leaf values
Typical outputA continuous numeric prediction; some probabilistic variants can also provide uncertainty or intervals.
Good fitHouse-price estimation, insurance severity, demand estimation, interpretable non-linear numeric prediction.
Main cautionHigh variance; prone to severe overfitting on noisy data without pruning or depth constraints.
1Training data→
2Learning objective→
3Internal model state→
4Prediction / representation→
5Evaluation
Intuition

What the model is trying to learn

A decision tree repeatedly asks simple feature questions. Each question divides the data into more homogeneous regions until leaves contain sufficiently similar targets. The final model is therefore both a prediction rule and a visible hierarchy of decisions.

Mathematical lens

Core logic

At every candidate node, the algorithm evaluates possible feature thresholds and selects the split that maximises impurity reduction (classification) or reduces squared/absolute target error (regression). Complexity is controlled by depth, leaf size and pruning constraints.

Training sequence

How learning progresses

Place all samples at the root → score candidate splits → choose the strongest split → recurse into child nodes → stop according to complexity rules → predict from terminal leaves.

Original mechanism

Taxonomy description

Recursively partitions feature space into axis-aligned rectangular regions using purity split criteria like Gini Impurity or Information Gain.

Evaluation guide

How to evaluate this model responsibly

ValidationK-Fold; Group K-Fold for repeated entities; time-aware splits for temporal targets.
MetricsMAE and RMSE together, plus R²; inspect residuals rather than trusting one aggregate score.
HPORandom/Bayesian optimisation for continuous hyperparameters; use nested CV when model selection is intensive.
Post-processingInverse target transforms, clipping only with domain justification, and prediction intervals where uncertainty matters.
Hyperparameters

Key parameters

max_depthTypical: None

Maximum depth of the tree.

min_samples_splitTypical: 2

Minimum samples required to split a node.

min_samples_leafTypical: 1

Minimum samples in a terminal leaf.

criterionTypical: gini

Split-quality function.

Use & trade-offs

Where it fits

Typical applications

House-price estimation, insurance severity, demand estimation, interpretable non-linear numeric prediction.

Strengths

Completely transparent white-box interpretability, handles numerical and categorical features naturally, no feature scaling required.

Limitations

High variance; prone to severe overfitting on noisy data without pruning or depth constraints.

Code example

Minimal Python implementation

# STEP 1 · Build a nonlinear two-feature regression problem.
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, r2_score

rng = np.random.default_rng(42)
X = rng.uniform(-3, 3, size=(240, 2))
y = (1.2 + 0.75*X[:,0]**2 - 0.45*X[:,1]
     + 0.65*np.sin(X[:,0]*X[:,1]) + rng.normal(0, 0.35, 240))
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.25, random_state=42)
print("STEP 1 · Train/test:", X_train.shape, X_test.shape)

# STEP 2 · Fit Decision Tree Regressor (CART) to the curved target.
from sklearn.tree import DecisionTreeRegressor
model = DecisionTreeRegressor(max_depth=5, random_state=42)
model.fit(X_train, y_train)
print("STEP 2 · Model fitted")

# STEP 3 · Evaluate held-out nonlinear predictions.
pred = model.predict(X_test)
rmse = mean_squared_error(y_test, pred) ** 0.5
r2 = r2_score(y_test, pred)
print("STEP 3 · RMSE:", round(rmse, 3))
print("R²:", round(r2, 3))
print("First predictions:", np.round(pred[:4], 2).tolist())
Expected / representative output
STEP 1 · Prepare the miniature example
STEP 2 · Fit / train the model
STEP 3 · Inspect predictions / metrics
Three numeric predictions.