Features Engineering using Deep Learning Models
Abstract :
The curse of dimensionality refers to all or any the problems that arise when working with data within the higher dimensions, which did not exist within the lower dimensions. As the number of features increase, the number of samples conjointly increases proportionately. As a result, the number of features will increase the model becomes more complicated. If the amount of features increases there are a lot of possibilities of over fitting. A machine learning model that is trained on an outsize range of features gets increasingly dependent on the data. Avoiding over fitting is also a serious motivation for playing dimensionality reduction in features engineering. We now live in an era of data deluge where large volumes of data are accumulating altogether aspects of our lives. Data streams coming from diverse domains contribute to the emerging paradigm of massive data. It’s going to be an excellent opportunity for the large data scientist amongst the vast amount and array of data. By discovering associations, analyzing patterns and predicting trends within the data, big data has the potential to vary our society and improve the standard of our life. Big data typically refers to the subsequent three types supported data sources from physical, cyber, and social worlds. While big data brings great opportunities, unpredictable challenges are on the way at an equivalent time. It can't be stored, analyzed and processed by traditional data management technologies and requires adaptation of some new workflows, platforms and architectures. The sector of machine learning which is beneficial to accomplish tasks of prediction, classification, and association about large amounts of knowledge is getting more and more attention from researchers within the current time. However, because the big data era is coming some characteristics of massive data will bring great challenges to the normal machine learning methods. Volume is that the important aspect of massive data. As data volumes and varieties grow, processing and consuming the insights generated becomes challenging. The curse of dimensionality refers to varied challenges that arise when analyzing data in high-dimensional spaces that don't occur in low-dimensional settings. High dimensionality including large sample sizes creates issues like heavy computational cost and algorithmic instability.