Deal of The Day! Hurry Up, Grab the Special Discount - Save 25% - Ends In 00:00:00 Coupon code: SAVE25
Welcome to Pass4Success

- Free Preparation Discussions

Databricks Machine Learning Associate Exam - Topic 1 Question 48 Discussion

A data scientist has replaced missing values in their feature set with each respective feature variable's median value. A colleague suggests that the data scientist is throwing away valuable information by doing this.Which of the following approaches can they take to include as much information as possible in the feature set?
D) Create a binary feature variable for each feature that contained missing values indicating whether each row's value has been imputed
A) Impute the missing values using each respective feature variable's mean value instead of the median value
B) Refrain from imputing the missing values in favor of letting the machine learning algorithm determine how to handle them
C) Remove all feature variables that originally contained missing values from the feature set
E) Create a constant feature variable for each feature that contained missing values indicating the percentage of rows from the feature that was originally missing

Databricks Machine Learning Associate Exam - Topic 1 Question 48 Discussion

Actual exam question for Databricks's Databricks Machine Learning Associate exam
Question #: 48
Topic #: 1
[All Databricks Machine Learning Associate Questions]

A data scientist has replaced missing values in their feature set with each respective feature variable's median value. A colleague suggests that the data scientist is throwing away valuable information by doing this.

Which of the following approaches can they take to include as much information as possible in the feature set?

Show Suggested Answer Hide Answer
Suggested Answer: D

By creating a binary feature variable for each feature with missing values to indicate whether a value has been imputed, the data scientist can preserve information about the original state of the data. This approach maintains the integrity of the dataset by marking which values are original and which are synthetic (imputed). Here are the steps to implement this approach:

Identify Missing Values: Determine which features contain missing values.

Impute Missing Values: Continue with median imputation or choose another method (mean, mode, regression, etc.) to fill missing values.

Create Indicator Variables: For each feature that had missing values, add a new binary feature. This feature should be '1' if the original value was missing and imputed, and '0' otherwise.

Data Integration: Integrate these new binary features into the existing dataset. This maintains a record of where data imputation occurred, allowing models to potentially weight these observations differently.

Model Adjustment: Adjust machine learning models to account for these new features, which might involve considering interactions between these binary indicators and other features.

Reference

'Feature Engineering for Machine Learning' by Alice Zheng and Amanda Casari (O'Reilly Media, 2018), especially the sections on handling missing data.

Scikit-learn documentation on imputing missing values: https://scikit-learn.org/stable/modules/impute.html


Contribute your Thoughts:

0/2000 characters
Jeanice
3 days ago
Letting the algorithm handle it might lead to unexpected results.
upvoted 0 times
...
Bettina
8 days ago
Surprised that people still debate median vs mean for imputation!
upvoted 0 times
...
Stephaine
13 days ago
Removing features just because they have missing values seems extreme.
upvoted 0 times
...
Mireya
19 days ago
I think option D is a smart way to keep track of imputed values!
upvoted 0 times
...
Selene
24 days ago
Using the mean can skew results if there are outliers.
upvoted 0 times
...
Odette
29 days ago
Removing features with missing values seems too extreme; I feel like we might lose important data that could be useful for the model.
upvoted 0 times
...
Brice
1 month ago
I practiced a question similar to this, and I think creating a binary feature for imputed values could help retain some information about the missingness.
upvoted 0 times
...
Aimee
1 month ago
I think letting the algorithm handle missing values could be a good idea, but I’m not confident about how that would impact the model's performance.
upvoted 0 times
...
Katie
1 month ago
I remember discussing how using the mean instead of the median can be sensitive to outliers, but I'm not sure if that's a better option here.
upvoted 0 times
...

Save Cancel