A credit card company wants to build a credit scoring model to help predict whether a new credit card applicant will default on a credit card payment. The company has collected data from a large number of sources with thousands of raw attributes. Early experiments to train a classification model revealed that many attributes are highly correlated, the large number of features slows down the training speed significantly, and that there are some overfitting issues.
The Data Scientist on this project would like to speed up the model training time without losing a lot of information from the original dataset.
Which feature engineering technique should the Data Scientist use to meet the objectives?
ahquiceno
Highly Voted 3 years, 7 months agoDr_Kiko
3 years, 5 months agoVinceCar
2 years, 5 months ago[Removed]
Highly Voted 3 years, 6 months agoSophieSu
3 years, 6 months agorodrigus
2 years, 1 month agoxicocaio
Most Recent 7 months agoGiodefa96
8 months, 4 weeks agogeoan13
1 year, 5 months agoMickey321
1 year, 8 months agoMickey321
1 year, 8 months agokaike_reis
1 year, 8 months agovbal
1 year, 11 months agoJK1977
1 year, 11 months agoGOSD
1 year, 11 months agooso0348
2 years agoPaolo991
2 years, 1 month agoSneep
2 years, 3 months agoAninina
2 years, 3 months agoovokpus
2 years, 10 months agoovokpus
2 years, 10 months ago