Supervised LearningMulti-Class ClassificationClassification

Decision Tree Classifier (CART)

Primary task · Classification

Decision Tree Classifier (CART) applies the Decision Trees (CART) learning mechanism to categorical targets. Decision Trees (CART) is a supervised learning method in the multi-class classification family. This page summarizes its mechanism, practical uses, important trade-offs, and a browser-based concept explorer.

← Directory
Visual intuition

From data to learned behaviour

A decision tree repeatedly asks simple feature questions. Each question divides the data into more homogeneous regions until leaves contain sufficiently similar targets. The final model is therefore both a prediction rule and a visible hierarchy of decisions.

Infographic
1Training data2Best split3Branches4Leaves5PredictionTraining transforms evidence into a reusable model state
Conceptual simulation

Watch the learning mechanism form

The structure below is synchronized with the same training state used by the prediction simulation.

Mechanism view
Training control centre

Control both simulations together

Reset regenerates the synthetic data and model state. Train animates to completion. Pause freezes the animation. Train Step advances one learning stage.

Step 0 / 8
Model simulation

Inspect the learned prediction / representation

Synthetic data are generated locally in your browser.

Model description

Understand Decision Tree Classifier (CART) after watching it learn

This section connects the animation to the actual statistical or computational idea behind the model.

Deep description

Decision Tree Classifier (CART) Decision Tree Classifier (CART) applies the Decision Trees (CART) learning mechanism to categorical targets. Decision Trees (CART) is a supervised learning method in the multi-class classification family. This page summarizes its mechanism, practical uses, important trade-offs, and a browser-based concept explorer.

What is learned. During training, the algorithm builds or adjusts feature thresholds, branches and terminal leaf values. The core learning mechanism is: Recursively partitions feature space into axis-aligned rectangular regions using purity split criteria like Gini Impurity or Information Gain.

How training becomes inference. Place all samples at the root → score candidate splits → choose the strongest split → recurse into child nodes → stop according to complexity rules → predict from terminal leaves. Once training stops, the fitted state is reused on unseen inputs rather than being reconstructed from scratch. The resulting output is: Class probabilities or class labels, depending on the decision threshold and API used.

Why practitioners use it. Completely transparent white-box interpretability, handles numerical and categorical features naturally, no feature scaling required. Typical fits include Customer churn diagnosis, operational decision rules, medical triage workflows, feature importance inspection.

What to verify before trusting it. High variance; prone to severe overfitting on noisy data without pruning or depth constraints. The visual simulation is intentionally simplified, so real use should still validate preprocessing, data independence, hyperparameters, uncertainty and task-appropriate metrics.

Internal statefeature thresholds, branches and terminal leaf values
Typical outputClass probabilities or class labels, depending on the decision threshold and API used.
Good fitCustomer churn diagnosis, operational decision rules, medical triage workflows, feature importance inspection.
Main cautionHigh variance; prone to severe overfitting on noisy data without pruning or depth constraints.
1Training data→
2Learning objective→
3Internal model state→
4Prediction / representation→
5Evaluation
Intuition

What the model is trying to learn

A decision tree repeatedly asks simple feature questions. Each question divides the data into more homogeneous regions until leaves contain sufficiently similar targets. The final model is therefore both a prediction rule and a visible hierarchy of decisions.

Mathematical lens

Core logic

At every candidate node, the algorithm evaluates possible feature thresholds and selects the split that maximises impurity reduction (classification) or reduces squared/absolute target error (regression). Complexity is controlled by depth, leaf size and pruning constraints.

Training sequence

How learning progresses

Place all samples at the root → score candidate splits → choose the strongest split → recurse into child nodes → stop according to complexity rules → predict from terminal leaves.

Original mechanism

Taxonomy description

Recursively partitions feature space into axis-aligned rectangular regions using purity split criteria like Gini Impurity or Information Gain.

Evaluation guide

How to evaluate this model responsibly

ValidationStratified K-Fold; Group/StratifiedGroup K-Fold when samples share subjects or entities.
MetricsF1, ROC-AUC, PR-AUC, log loss and a confusion matrix; use balanced accuracy for imbalanced classes.
HPORandom search or Bayesian optimisation after a reasonable baseline; nested CV when tuning and unbiased performance estimation must be separated.
Post-processingTune decision thresholds and calibrate probabilities when downstream decisions use risk scores.
Hyperparameters

Key parameters

max_depthTypical: None

Maximum depth of the tree.

min_samples_splitTypical: 2

Minimum samples required to split a node.

min_samples_leafTypical: 1

Minimum samples in a terminal leaf.

criterionTypical: gini

Split-quality function.

Use & trade-offs

Where it fits

Typical applications

Customer churn diagnosis, operational decision rules, medical triage workflows, feature importance inspection.

Strengths

Completely transparent white-box interpretability, handles numerical and categorical features naturally, no feature scaling required.

Limitations

High variance; prone to severe overfitting on noisy data without pruning or depth constraints.

Code example

Minimal Python implementation

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score, confusion_matrix

# STEP 1 · Build a three-class classification problem.
X, y = make_classification(n_samples=210, n_features=6, n_informative=5,
                           n_redundant=0, n_classes=3, n_clusters_per_class=1,
                           class_sep=1.15, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, stratify=y, random_state=42)
print("STEP 1 · Train/test:", X_train.shape, X_test.shape)

# STEP 2 · Grow a multiclass CART tree.
model = DecisionTreeClassifier(max_depth=5, random_state=42).fit(X_train, y_train)
print("STEP 2 · Leaves:", model.get_n_leaves())

# STEP 3 · Inspect three-class predictions.
pred = model.predict(X_test)
print("STEP 3 · Accuracy:", round(accuracy_score(y_test, pred), 3))
print("Confusion matrix:")
print(confusion_matrix(y_test, pred))
Expected / representative output
STEP 1 · Prepare the miniature example
STEP 2 · Fit / train the model
STEP 3 · Inspect predictions / metrics
A list of five predicted class labels.