Deal of The Day! Hurry Up, Grab the Special Discount - Save 25% - Ends In 00:00:00 Coupon code: SAVE25
Welcome to Pass4Success

- Free Preparation Discussions

Amazon MLS-C01 Exam - Topic 1 Question 134 Discussion

[Modeling]A data scientist must build a custom recommendation model in Amazon SageMaker for an online retail company. Due to the nature of the company's products, customers buy only 4-5 products every 5-10 years. So, the company relies on a steady stream of new customers. When a new customer signs up, the company collects data on the customer's preferences. Below is a sample of the data available to the data scientist.How should the data scientist split the dataset into a training and test set for this use case?
D) Randomly select 10% of the users. Split off all interaction data from these users for the test set.
A) Shuffle all interaction data. Split off the last 10% of the interaction data for the test set.
B) Identify the most recent 10% of interactions for each user. Split off these interactions for the test set.
C) Identify the 10% of users with the least interaction data. Split off all interaction data from these users for the test set.

Amazon MLS-C01 Exam - Topic 1 Question 134 Discussion

Actual exam question for Amazon's MLS-C01 exam
Question #: 134
Topic #: 1
[All MLS-C01 Questions]

[Modeling]

A data scientist must build a custom recommendation model in Amazon SageMaker for an online retail company. Due to the nature of the company's products, customers buy only 4-5 products every 5-10 years. So, the company relies on a steady stream of new customers. When a new customer signs up, the company collects data on the customer's preferences. Below is a sample of the data available to the data scientist.

How should the data scientist split the dataset into a training and test set for this use case?

Show Suggested Answer Hide Answer
Suggested Answer: D

Contribute your Thoughts:

0/2000 characters
Nieves
1 day ago
Not sure if any of these options are ideal for such low-frequency purchases.
upvoted 0 times
...
Kandis
6 days ago
Totally agree with B! Recent interactions are key.
upvoted 0 times
...
Jordan
12 days ago
Wait, why would you use the least active users? That seems off.
upvoted 0 times
...
Titus
17 days ago
I think A is better. Random shuffling could work!
upvoted 0 times
...
Glen
22 days ago
Option B makes the most sense for new customers.
upvoted 0 times
...
Nydia
27 days ago
I recall that random sampling can sometimes be misleading. Option D seems risky because it might not represent the overall user behavior accurately.
upvoted 0 times
...
Tamekia
1 month ago
I practiced a similar question where we had to consider user behavior over time. I think option C could be useful, but I'm worried about losing too much data from those less active users.
upvoted 0 times
...
Brandon
1 month ago
I'm not entirely sure, but I feel like option A could lead to data leakage since it shuffles everything. We might miss out on how preferences change over time.
upvoted 0 times
...
Yvette
1 month ago
I remember we discussed the importance of temporal data in our last class. I think option B makes the most sense since it focuses on the most recent interactions for each user.
upvoted 0 times
...

Save Cancel