Which of the following issues should a data scientist be most concerned about when generating a synthetic data set?
If synthetic data don't accurately mirror the real-world distributions and relationships, any models trained on them will perform poorly in deployment. Representativeness is the critical concern when generating synthetic data.
Currently there are no comments in this discussion, be the first to comment!