UMAP UMAP is an unsupervised learning method in the non-linear manifold learning family. This page summarizes its mechanism, practical uses, important trade-offs, and a browser-based concept explorer.
What is learned. During training, the algorithm builds or adjusts the parameters and internal representation used by UMAP. The core learning mechanism is: Constructs a fuzzy simplicial set representation of high-dimensional data and optimizes low-dimensional embedding preserving both local and global topology.
How training becomes inference. Prepare data → initialise the model state → evaluate the current objective → update parameters or structure → validate progress → use the final state for inference. Once training stops, the fitted state is reused on unseen inputs rather than being reconstructed from scratch. The resulting output is: A lower-dimensional embedding or transformed representation designed to preserve selected structure.
Why practitioners use it. Significantly faster than t-SNE, preserves global structure much better, scales to large datasets. Typical fits include Single-cell RNA sequencing visualization, embedding space exploration, high-dimensional clustering prep.
What to verify before trusting it. Non-deterministic; sensitive to hyperparameters (n_neighbors, min_dist); distances between clusters cannot be strictly interpreted. The visual simulation is intentionally simplified, so real use should still validate preprocessing, data independence, hyperparameters, uncertainty and task-appropriate metrics.
Internal statethe parameters and internal representation used by UMAP
Typical outputA lower-dimensional embedding or transformed representation designed to preserve selected structure.
Good fitSingle-cell RNA sequencing visualization, embedding space exploration, high-dimensional clustering prep.
Main cautionNon-deterministic; sensitive to hyperparameters (n_neighbors, min_dist); distances between clusters cannot be strictly interpreted.