Follow the transformation
Split the data before fitting learned preprocessing.
Apply transformations separately by column type.
Fit preprocessing only on the training fold.
Target Transformations is a preprocessing operation that changes feature representation before modelling.
Target Transformations is a preprocessing operation that changes feature representation before modelling. Preprocessing is part of the learned pipeline because statistics such as means, categories or quantiles must be estimated on training data and then applied unchanged to validation/test data.
Target Transformations matters because preprocessing defines the representation a model actually sees. Learned transformations must be fitted inside the training boundary so their parameters do not leak information from validation or test cases.
Treat this as a sequence of observable decisions rather than one opaque command. Stage 1: Split the data before fitting learned preprocessing. Stage 2: Apply transformations separately by column type. Stage 3: Fit preprocessing only on the training fold. Final checkpoint: Bundle preprocessing and estimator in one pipeline so cross-validation repeats the correct sequence automatically.
Split the data before fitting learned preprocessing.
Apply transformations separately by column type.
Fit preprocessing only on the training fold.
Split the data before fitting learned preprocessing. At this stage of Target Transformations, keep the incoming data or object separate from the learned parameter, transformed object, or statistic so the change can be reproduced and independently checked.
# Step 1 — Import the module so its functions/classes are available to the rest of this example.
import math
# Step 2 — Compute the right-hand expression and store its result in `y` for the next step.
y=1000
# Step 3 — Compute the right-hand expression and store its result in `logy` for the next step.
logy=math.log1p(y)
# Step 4 — Compute the right-hand expression and store its result in `back` for the next step.
back=math.expm1(logy)
# Step 5 — Display the current value explicitly so the result/state can be inspected during execution.
print(round(logy,3), round(back,1))log1p compresses a positive skewed target; inverse transformation returns to the original unit for interpretation.
For Target Transformations, trace representative source rows/columns into the result and reconcile row counts, dtypes, keys or missing values that the operation could change.
Standardisation(x - training mean) / training standard deviation.Robust scalingUses median and quantile range; less influenced by extremes.One-hot encodingCreates indicator columns for categories without imposing numeric order.PipelineChains preprocessing and estimator so fitting/evaluation stay leakage-safe.Use Target Transformations when the model/analysis requires a deliberate representation of raw features and the transformation can be fit without leaking future or held-out information.
Avoid transformations that are unnecessary for the chosen model, cannot be reproduced at inference time, or learn from data that should remain held out.
Build a tiny, inspectable example of Target Transformations. First split the data before fitting learned preprocessing. Then apply transformations separately by column type. Write the expected result before running it, and explain one condition that would make the result misleading or invalid.
Before trusting a result from Target Transformations, which check provides the strongest evidence that you understand and applied it correctly?
Step 1Split the data before fitting learned preprocessing.Step 2Apply transformations separately by column type.Step 3Fit preprocessing only on the training fold.Step 4Transform validation/test with the stored training parameters.