LoRA (Low-Rank Adaptation) LoRA (Low-Rank Adaptation) is a ensemble learning & modern enablers method in the peft family. This page summarizes its mechanism, practical uses, important trade-offs, and a browser-based concept explorer.
What is learned. During training, the algorithm builds or adjusts low-rank adapter matrices that modify a frozen base model. The core learning mechanism is: Freezes pre-trained foundation model weights and injects trainable low-rank decomposition matrices (A and B) into Transformer attention projection layers.
How training becomes inference. Prepare data → initialise the model state → evaluate the current objective → update parameters or structure → validate progress → use the final state for inference. Once training stops, the fitted state is reused on unseen inputs rather than being reconstructed from scratch. The resulting output is: A more efficient training/adaptation configuration rather than a conventional predictive target.
Why practitioners use it. Reduces trainable parameters by >99% and GPU memory by 60-70%; zero added inference latency when merged back into base weights. Typical fits include Efficient fine-tuning of LLMs (e.g., LLaMA, Mistral, Qwen) and Diffusion models on consumer GPUs.
What to verify before trusting it. Slightly lower expressive capacity than full parameter fine-tuning on fundamentally novel domains. The visual simulation is intentionally simplified, so real use should still validate preprocessing, data independence, hyperparameters, uncertainty and task-appropriate metrics.
Internal statelow-rank adapter matrices that modify a frozen base model
Typical outputA more efficient training/adaptation configuration rather than a conventional predictive target.
Good fitEfficient fine-tuning of LLMs (e.g., LLaMA, Mistral, Qwen) and Diffusion models on consumer GPUs.
Main cautionSlightly lower expressive capacity than full parameter fine-tuning on fundamentally novel domains.