Thursday, 19 May 2022

Feature Selection in Machine Learning


Today we will see what is Feature Selection and Various Methods used in this process.

Description

  • Feature selection is a process where from all the features only the optimal subset of features are selected for the learning algorithm(model as well we call sometimes).
  • Different methods are used to select best and required features for the model.

  • Feature Selection is needed for :

    • Easy to interrupt or understand

    • Optimising training times

    • Reduces overfitting

    • Removing highly correlated variables

  • Feature Selection Methods are below :

    • Filter : using features

      • Variance

      • Correlation

      • Univariate Selection

        • Mutual Information

        • Chi Square

        • F-Test

        • ROC_AUC or RMSE

    • Wrapper : using predictive ML

      • Step backward selection

      • Step forward selection

      • Exhaustive search selection

    • Embedded

      • Linear

      • LASSO

      • Tree Importance


Filter

  • Filter methods are generally used as a preprocessing step

  • The selection of features is independent of any machine learning algorithms. 

  • Instead, features are selected on the basis of their scores in various statistical tests for their correlation with the outcome variable. 

  • Visualization



Filter - Variance 

  • Remove columns which as constant values, we can check wth standard deviation with 0

  • Remove columns which have standard deviation of 0.01, like almost constant. 

    • But need to be careful while doing this. 

    • Also known as quasi constant features.

    • Eg: we have feature called A have values 1 with frequency 9999 and 0 with frequency1

  • Remove columns which are same or remove duplicate columns


Filter – Correlation

  • Remove columns which are highly correlated within independent variables

  • Process : 

    • Finding highly correlated column groups

    • Find best column, which contributes more in target (Y) prediction


Filter – Univariate Selection

Mutual Information

  • Measures how much each independent values contributes in the correct prediction on Y

  • Works with only the numeric data

  • Used with the continous values in X

  • Feature Selection Methods :

    • Regression : sklearn.feature_selection.mutual_info_regression
    • Classification : sklearn.feature_selection.mutual_info_classif
  • Best Feature Selection Methods :

    • sklearn.feature_selection.SelectKBest
    • sklearn.feature_selection.SelectPercentile
  • Higher the mutual info score, more significant the features

  • Caution: Rarely Used in Real Life


Chi – Square

  • Measures how much each independent values contributes in the correct prediction on Y

  • Uses Fisher Score

  • Should be used to evaluate categorical variables in the classification task

  • Feature Selection Methods :

    • Classification : sklearn.feature_selection.chi2

  • Best Feature Selection Methods :

    • sklearn.feature_selection.SelectKBest
    • sklearn.feature_selection.SelectPercentile
  • Here, lower the p-values more significant the features

  • Note: Fisher score or any univariate selection model, with very big dataset shows small p value, therefore looking the low p-value we cannot conclude these are highly predicted features. But in reality this is because of large sample size

  • Caution: Rarely Used in Real Life


F Test

  • Measures how much each independent values contributes in the correct prediction on Y

  • Works by selecting the best features based on univariate statistical test (ANOVA)

  • Used with the continous values in X

  • F-Test Method

    • Assumptions :

      • Normal distribution

      • Linear relationship between the features and the target

    • So while using this we need to think on the assumptions as well

  • Feature Selection Methods :

    • Regression : sklearn.feature_selection.f_regression

    • Classification : sklearn.feature_selection.f_classif

  • Best Feature Selection Methods :

    • sklearn.feature_selection.SelectKBest

    • sklearn.feature_selection.SelectPercentile

  • Here, lower the p-values more significant the features

  • Caution: Rarely Used in Real Life, but can be used for investigation


ROC_AUC or RMSE

  • Uses ML to define most significant features

  • But this is Time and Memory Consuming

  • Features are selected using ROC_AUC or MSE

  • And the highest ranked features are selected

  • Get the numerical features, each X feature is used to fit for prediction Y value

  • ROC_AUC :

    • Features with the values > 0.5 is selected, as considered more significant

    • Less then that would be the features would be random

  • RMSE

    • Smaller the RMSE are more significant



Wrapper Methods

  • Is also known as Greedy Search

  • As this scans all the possible combination

  • Often not feasible due to number of features are high in the dataset

  • UPSIDE : Provides best features subset

  • DOWNSIDE : Extremely computable expensive, ML Model Specific feature selected

  • Combined with CV, will overcome the overfitting also

  • Visualization


Step Forward

  • Begins with one feature at a time and keeps adding one by one

  • Algorithm

    • Select each features at a time and get the best one using accuracy

    • Later add the second feature with all the feature space and together calculate the accuracy

    • From the best combination accuracy, select the second feature

    • Continue this until all the k number of features are selected or no further improvement is observed 

  • Method : mlxtend


Step Backward

  • Opposite of Step Forward

  • Evaluates all the features first, and removes one feature and evaluate algorithm performance

  • Algorithm

    • Select all features at a time and get the accuracy

    • Later remove one feature, from all the feature space

    • Calculate the accuracy, and remove the least significant feature at each iteration which improves the performance of the model.

    • Continue this until all the k number of features are selected or no improvement is observed 

  • Method : mlxtend


Exhaustive

  • Try all the possible combinations, like tries all the single, group (combination) features

  • PROS: Computationally expensive, at times unfeasible

  • CONS: Provides best subset

  • Algorithm

    • all possible combinations of 1 feature

    • all possible combinations of 2 feature

    • all possible combinations of n feature

    • and selects the one that results the best in performance

  • Method : mlxtend



Embedded

  • Embedded methods combine the qualities’ of filter and wrapper methods

  • It’s implemented by algorithms that have their own built-in feature selection methods

  • Visualization


Regularization

  • In regression we always try to find Low Cost Funtion Score, for the best fitted line, with less errors in Training Set

  • Regression Coefficient :

    • y = β0 +  β1X1 + β2X2 + .. + βnXn

  • The coefficients are predictors the Bs are directly proportional how the features contributes to the final value y

  • Regularization consists of adding penalty to the different parameters

  • Hence the model is less likely to fit the noise of the training data and improving the generalization

  • In general the below models are used when the Simple Regression model predicts overfitting results, so these techniques are used to reduce model complexity and prevent over fitting

  • Types of Linear Model

    • L1 Regularization : Lasso

    • L2 Regularization : Ridge

    • L1|L2 Regularization : Elasticnet

  • Assumptions :

    • Linear relationship between X and Y

    • Xs are independent

    • Xs are not correlated to other independent feature

    • Xs are normally distributed

    • Xs should be scaled as the coefficients are directly influenced by the scale of X

    • So to select important features its very much important all the X should be in the Same Scale


Lasso Reguralization

  • Adding penalty to the coefficient

    • penalty is lambda

  • This is done to avoid overfitting

  • Formula : 

  • Formula to find Coefficient:

    • 1/2m * ∑(y – ypred)ˆ2 = ႓ ∑ϴ ˆ1

    • m : number of observations

    • y : observed

    • ypred : predicted output

    • ႓ : lamda regularization parameter, penalty

    • ϴ : coefficient, its power one to make it always positive

  • Note: when λ → 0 , the cost function becomes similar to the linear regression cost function

  • So lower the penalty it will resemble the linear regression model

  • The higher the penalty the bigger the generalization

  • This ML model, tries to minimize the difference between y and ypred, works to reduce all the three including the regularization component

  • Here the ϴ is at times moved to 0 so some features could be removed, as the formula doesnt takes the Square of Cofficient only the magnitude is taken into consideration

  • Hence this is used for Feature Selection, not for model optimization

  • Model Optimization : Reduce over fitting

  • Method: 

    • sklearn.linear_model.Lasso

  • Increase in the penalty will remove the features, so we need to keep this in mind


Ridge Reguralization

  • Works same as Lasso, in Ridge the ϴ may be changed to close to 0 but never the ϴ is never turned to 0, as the coefficient is always squared

  • Hence Not suitable for feature selection, but used for model optimization

  • Model Optimization is done to reduce over fitting

  • Formula :


Tree Based

Decision Tree

  • Most popular way to use in the feature selections

  • Highly accurate in training set

  • Provides good generalization

  • Helps in bringing Interpretobility

  • Algorithm

    • First Node: Highest Impurity

    • Impurity decreases further each level

    • Until the impurity is 0

  • More the feature decreases the impurity, the more the the feature importance


Random Forest

  • Contains several decision trees

  • Decrease of impurity for each feature is averaged accross trees

  • Algorithm

    • Build random forests

    • Calculate feature importance

    • Remove least important feature

    • Repeat till condition is met

  • Less of real life use, although doesn't give more benifit in removing feature

  • Limitation

    • If correlated are not removed both the variable will show equal low importance

    • If a feature removed is correlated to another feature in the dataset, by removing the correlated features the tree importance of the feature will be revealed

    • Highly baised with the categorical variables



Reference and Credits

https://www.analyticsvidhya.com/blog/2016/12/introduction-to-feature-selection-methods-with-an-example-or-how-to-select-the-right-variables/

https://www.datacamp.com/community/tutorials/feature-selection-python

https://towardsdatascience.com/feature-importance-and-forward-feature-selection-752638849962

https://towardsdatascience.com/ridge-and-lasso-regression-a-complete-guide-with-python-scikit-learn-e20e34bcbf0b

https://towardsdatascience.com/of-suppandi-regularization-and-lasso-the-feature-selector-2a09acfdbc1b

https://miro.medium.com/max/1400/1*MO2JGos1gZnyfQueFrzzsw.png



2 comments:

  1. Very technical yet easy to understand. Underrated topic in ML but very important. Covered it nicely.

    ReplyDelete

Scarcity Brings Efficiency: Python RAM Optimization

  In today’s world, with the abundance of RAM available, we rarely think about optimizing our code. But sooner or later, we hit the limits a...