Data Analytics · Interactive concept

Correlation Analysis

See how direction, strength, outliers and nonlinearity change correlation—and why a coefficient must always be read with the scatter plot.

Start here

What does a correlation coefficient actually describe?

Correlation summarises the direction and strength of association between variables. Pearson correlation measures linear association on the original numeric scale; Spearman correlation measures monotonic association after converting values to ranks. A coefficient is not a causal estimate and should be interpreted together with the scatter plot and data-generating context.

Current correlation——
Range−1 to +1sign = direction; magnitude = strength
Pearson correlation

Linear association

Correlation summarises the direction and strength of association between variables. Pearson correlation measures linear association on the original numeric scale; Spearman correlation measures monotonic association after converting values to ranks. A coefficient is not a causal estimate and should be interpreted together with the scatter plot and data-generating context. The coefficient is only a compressed summary of paired observations. Its meaning depends on pairing, missing-data handling, shape, influential points and subgroups, so the scatter plot and data-generating context are part of the analysis rather than optional decoration.

r = covariance(X, Y) / (sd(X) × sd(Y))
  • Pair observations correctly and handle missing pairs deliberately.
  • Plot the variables to check shape, clusters and outliers.
  • Compute Pearson r for linear association or Spearman rho for monotonic ranked association when appropriate.
  • Interpret sign as direction and magnitude as strength relative to context.
  • Check whether subgroups or influential points change the coefficient.
Spearman correlation

Monotonic association

Spearman correlation applies Pearson correlation to ranks. It is useful when the relationship is monotonic but not well described by equal numeric spacing.

ρ = Pearson correlation of rank(X) and rank(Y)
Always inspect the scatter plotThe same coefficient can hide clusters, curvature or an influential outlier.
Correlation ≠ causationConfounding, reverse causality and selection effects can produce association without a causal relationship.
Feature selection

Correlation can screen features—but only with safeguards.

Pair observations correctly and handle missing pairs deliberately. Plot the variables to check shape, clusters and outliers. Compute Pearson r for linear association or Spearman rho for monotonic ranked association when appropriate.

Open Correlation Feature Selection Playground →
What correlation can miss

Weak r does not mean “useless feature”.

Use correlation to summarise a defined bivariate association: Pearson for linear association on a numeric scale, or Spearman for monotonic association in ranked values. Reconsider a single coefficient when the scatter plot shows curvature, separated subgroups, influential outliers, restricted range or a causal question that correlation cannot answer.

Detailed concept notes

Mechanism, distinctions and checks

UnivariateOne variable: distribution, counts, centre and spread.
BivariateTwo variables: association or group differences.
MultivariateSeveral variables: interactions, confounding, conditional patterns.
Confirmatory analysisTests or models a pre-specified claim; should be distinguished from open-ended exploration.

Common mistakes

  • Generating dozens of plots without a question or interpretation.
  • Using correlation alone to judge a relationship without the scatter plot.
  • Letting information from the future/test set influence model-building decisions.

Verification

  • For a four- or five-pair example, centre the values and inspect the covariance/sign pattern (or rank the pairs for Spearman), then compare the hand calculation with the displayed coefficient.
  • Check that X and Y contain the same paired observations after missing-value handling, and inspect the scatter plot for curvature, clusters and influential points.
  • Move one influential point or add a subgroup and predict whether Pearson and Spearman should strengthen, weaken or diverge before recomputing them.