In a LinkedIn post I asked a question that comes up early in machine learning: when should we normalize, and when should we standardize? The short answer fits on a card. The useful answer needs more room.
Revisiting the subject for this article, I wanted to refine some simplifications in that post. Standardization does not make a distribution normal. Logistic regression and SVM do not generally require every feature to follow a Gaussian distribution. Using a mean and standard deviation does not make the standard scaler resistant to outliers, either.
The question remains worthwhile. I would change the starting point: before choosing a transformation, understand what the algorithm will do with differences between observations.
Two customers, two units
Imagine finding similar customers using age and annual income. A difference of ten years and one of ten thousand euros enter Euclidean distance as ten and ten thousand. Income dominates numerically before we have decided how much either feature ought to matter.
Express income in thousands of euros and the second difference becomes ten. The customer did not change; the geometry presented to the model did. Scaling addresses this problem when variables with different units share a comparison.
Algorithms respond differently. Nearest neighbors and k-means depend directly on distances. With an RBF SVM, scale affects the kernel’s distances. In regularized linear models it also affects coefficient penalties. Decision trees are generally much less sensitive to monotonic transformations of an individual feature because they split at thresholds, though implementation and the rest of the pipeline still matter.
Not scaling is an option too. It should be a reasoned choice rather than a forgotten step.
When you need a range
Min-max scaling transforms a value using (x − minimum) / (maximum − minimum), usually putting training observations between zero and one. Zero corresponds to the observed minimum, one to the maximum.
I will call this normalization, as in my post. The word can also mean normalizing a vector to unit norm, used for some representations. That is a different operation. Saying “min-max” avoids ambiguity when it matters.
We do not need a non-normal distribution to use min-max. We need that range to make sense for the intended task. Nor does it guarantee that future data stays between zero and one: an observation exceeding the training maximum can transform to a value above one. Clipping it is another decision and can conceal a meaningful change.
When you need centering and comparable spread
Standardization uses (x − mean) / standard deviation. Nonconstant training features end up with zero mean and unit variance under the calculation’s convention. Skewness, tails and multiple groups have not disappeared.
Its interpretation differs from min-max: a value of two is two training standard deviations above the training mean. This can be useful with non-Gaussian data. The choice depends on the model and on what the scale means, not on passing a normality test as a universal prerequisite.
An extraordinary customer arrives
Add one enormous income to several ordinary ones. Min-max uses this new maximum and compresses the ordinary group near zero. StandardScaler changes too: the extreme observation moves the mean and increases the standard deviation. Neither automatically solves the problem.
RobustScaler uses the median and, by default, the interquartile range. These statistics are generally less affected by extreme observations than the mean and standard deviation. It does not remove outliers or guarantee a good model. It can leave extreme points far from the central group, which may or may not suit our purpose.
The scikit-learn preprocessing documentation distinguishes these alternatives. When the concrete effect matters, I would visualize the transformation before relying on the label “robust.”
We also need to decide what this extraordinary customer represents. A unit error? A real but unusual person? A population worth studying separately? Scaling can improve a geometry, but it cannot answer those quality and context questions.
When we fit matters as much as the formula
Means, standard deviations, minima and maxima are learned from data. Calculating them on the complete dataset before splitting lets evaluation data influence part of the learning process.
The practical solution is to put transformation inside a pipeline. During cross-validation it is fitted within each training fold. Validation and test receive only the already learned transformation.
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
model = make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=5))
# X_train and y_train contain only the training partition.
model.fit(X_train, y_train)
predictions = model.predict(X_test)
Reproducing this needs a labeled problem and an appropriate split. The fragment shows the order, not a completed evaluation. Temporal observations or repeated people require more care than a random partition.
A fitted scaler stores parameters that belong to the system. Deployment must preserve and apply them consistently. Recalculating a mean on every incoming batch changes the representation received by an already trained model. A deliberate design can do this, but it should not happen accidentally.
The answer I would give now
If asked to choose, I would start with the algorithm and units. Then examine ranges, extreme values and distribution. If numerical boundaries are needed, examine min-max and handling of out-of-range inputs. For centered features with comparable spread, try standardization. For influential extremes, investigate their origin and consider robust scaling.
Compare options under the same protocol and leave the final test alone. Do not change the recipe after each result until a pleasing one appears. Inspect constant variables too: division by zero spread lacks the usual interpretation even when a library handles it without an exception.
One question helps me finish: if tomorrow we replace euros with thousands of euros, should the result change? If not, our treatment of scale must reflect that. If yes, because the magnitude carries relevant meaning, we should explain why. That conversation is usually more productive than deciding which of the two names sounds more correct.

Found this useful? If you would like to support this space, you can buy me a coffee.
Buy me a coffee Optional support through PayPal. You choose the amount.

