Decide what “new” means
Predicting another measurement from a familiar batch is a different task from predicting a new material family. Both can matter, but they require different evidence. A score without a description of the holdout leaves that distinction unresolved.
Before selecting a model, write down the future case you want to understand. That statement should guide how observations are kept apart.
Look beyond row identity
Different row identifiers do not guarantee independent observations. Samples can share preparation conditions, a source publication, or a time period. Related information crossing the train–test boundary can make performance look stronger than the intended use warrants.
Audit the data, identify meaningful grouping columns, and inspect a split’s actual membership. Choose from the installed strategies according to the scientific question rather than accepting a default without review.
Preserve the comparison
Compare supported baselines with the same inputs and evaluation conditions. Keep model selection separate from the final test, and record when test results were exposed. Once a result influences a choice, it is part of the experimental history.
Report the metric with its scope: the held-out population, sample count, units, and limitations. This makes the number useful to someone deciding whether it answers their question.