In which of the following situations is it preferable to impute missing feature values with their median value over the mean value?
Imputing missing values with the median is often preferred over the mean in scenarios where the data contains a lot of extreme outliers. The median is a more robust measure of central tendency in such cases, as it is not as heavily influenced by outliers as the mean. Using the median ensures that the imputed values are more representative of the typical data point, thus preserving the integrity of the dataset's distribution. The other options are not specifically relevant to the question of handling outliers in numerical data. Reference:
Data Imputation Techniques (Dealing with Outliers).
Charisse
22 days agoTrinidad
27 days agoAvery
1 month agoFranklyn
1 month agoNicolette
1 month agoEdgar
2 months agoEleonora
2 months agoLorriane
2 months agoCordelia
2 months agoZona
2 months agoBlondell
2 months agoChuck
3 months ago