Description
- Feature selection is a process where from all the features only the optimal subset of features are selected for the learning algorithm(model as well we call sometimes).
- Different methods are used to select best and required features for the model.
Feature Selection is needed for :
Easy to interrupt or understand
Optimising training times
Reduces overfitting
Removing highly correlated variables
Feature Selection Methods are below :
Filter : using features
Variance
Correlation
Univariate Selection
Mutual Information
Chi Square
F-Test
ROC_AUC or RMSE
Wrapper : using predictive ML
Step backward selection
Step forward selection
Exhaustive search selection
Embedded
Linear
LASSO
Tree Importance
Filter
Filter methods are generally used as a preprocessing step
The selection of features is independent of any machine learning algorithms.
Instead, features are selected on the basis of their scores in various statistical tests for their correlation with the outcome variable.
Visualization
Filter - Variance
Remove columns which as constant values, we can check wth standard deviation with 0
Remove columns which have standard deviation of 0.01, like almost constant.
But need to be careful while doing this.
Also known as quasi constant features.
Eg: we have feature called A have values 1 with frequency 9999 and 0 with frequency1
Remove columns which are same or remove duplicate columns
Filter – Correlation
Remove columns which are highly correlated within independent variables
Process :
Finding highly correlated column groups
Find best column, which contributes more in target (Y) prediction
Filter – Univariate Selection
Mutual Information
Measures how much each independent values contributes in the correct prediction on Y
Works with only the numeric data
Used with the continous values in X
Feature Selection Methods :
- Regression : sklearn.feature_selection.mutual_info_regression
- Classification : sklearn.feature_selection.mutual_info_classif
Best Feature Selection Methods :
- sklearn.feature_selection.SelectKBest
- sklearn.feature_selection.SelectPercentile
Higher the mutual info score, more significant the features
Caution: Rarely Used in Real Life
Chi – Square
Measures how much each independent values contributes in the correct prediction on Y
Uses Fisher Score
Should be used to evaluate categorical variables in the classification task
Feature Selection Methods :
Classification : sklearn.feature_selection.chi2
Best Feature Selection Methods :
- sklearn.feature_selection.SelectKBest
- sklearn.feature_selection.SelectPercentile
Here, lower the p-values more significant the features
Note: Fisher score or any univariate selection model, with very big dataset shows small p value, therefore looking the low p-value we cannot conclude these are highly predicted features. But in reality this is because of large sample size
Caution: Rarely Used in Real Life
F Test
Measures how much each independent values contributes in the correct prediction on Y
Works by selecting the best features based on univariate statistical test (ANOVA)
Used with the continous values in X
F-Test Method
Assumptions :
Normal distribution
Linear relationship between the features and the target
So while using this we need to think on the assumptions as well
Feature Selection Methods :
Regression : sklearn.feature_selection.f_regression
Classification : sklearn.feature_selection.f_classif
Best Feature Selection Methods :
sklearn.feature_selection.SelectKBest
sklearn.feature_selection.SelectPercentile
Here, lower the p-values more significant the features
Caution: Rarely Used in Real Life, but can be used for investigation
ROC_AUC or RMSE
Uses ML to define most significant features
But this is Time and Memory Consuming
Features are selected using ROC_AUC or MSE
And the highest ranked features are selected
Get the numerical features, each X feature is used to fit for prediction Y value
ROC_AUC :
Features with the values > 0.5 is selected, as considered more significant
Less then that would be the features would be random
RMSE
Smaller the RMSE are more significant
Wrapper Methods
Is also known as Greedy Search
As this scans all the possible combination
Often not feasible due to number of features are high in the dataset
UPSIDE : Provides best features subset
DOWNSIDE : Extremely computable expensive, ML Model Specific feature selected
Combined with CV, will overcome the overfitting also
Visualization
Step Forward
Begins with one feature at a time and keeps adding one by one
Algorithm
Select each features at a time and get the best one using accuracy
Later add the second feature with all the feature space and together calculate the accuracy
From the best combination accuracy, select the second feature
Continue this until all the k number of features are selected or no further improvement is observed
Method : mlxtend
Step Backward
Opposite of Step Forward
Evaluates all the features first, and removes one feature and evaluate algorithm performance
Algorithm
Select all features at a time and get the accuracy
Later remove one feature, from all the feature space
Calculate the accuracy, and remove the least significant feature at each iteration which improves the performance of the model.
Continue this until all the k number of features are selected or no improvement is observed
Method : mlxtend
Exhaustive
Try all the possible combinations, like tries all the single, group (combination) features
PROS: Computationally expensive, at times unfeasible
CONS: Provides best subset
Algorithm
all possible combinations of 1 feature
all possible combinations of 2 feature
all possible combinations of n feature
and selects the one that results the best in performance
Method : mlxtend
Embedded
Embedded methods combine the qualities’ of filter and wrapper methods
It’s implemented by algorithms that have their own built-in feature selection methods
Visualization
Regularization
In regression we always try to find Low Cost Funtion Score, for the best fitted line, with less errors in Training Set
Regression Coefficient :
y = β0 + β1X1 + β2X2 + .. + βnXn
The coefficients are predictors the Bs are directly proportional how the features contributes to the final value y
Regularization consists of adding penalty to the different parameters
Hence the model is less likely to fit the noise of the training data and improving the generalization
In general the below models are used when the Simple Regression model predicts overfitting results, so these techniques are used to reduce model complexity and prevent over fitting
Types of Linear Model
L1 Regularization : Lasso
L2 Regularization : Ridge
L1|L2 Regularization : Elasticnet
Assumptions :
Linear relationship between X and Y
Xs are independent
Xs are not correlated to other independent feature
Xs are normally distributed
Xs should be scaled as the coefficients are directly influenced by the scale of X
So to select important features its very much important all the X should be in the Same Scale
Lasso Reguralization
Adding penalty to the coefficient
penalty is lambda
This is done to avoid overfitting
Formula :


Formula to find Coefficient:
1/2m * ∑(y – ypred)ˆ2 = ႓ ∑ϴ ˆ1
m : number of observations
y : observed
ypred : predicted output
႓ : lamda regularization parameter, penalty
ϴ : coefficient, its power one to make it always positive
Note: when λ → 0 , the cost function becomes similar to the linear regression cost function
So lower the penalty it will resemble the linear regression model
The higher the penalty the bigger the generalization
This ML model, tries to minimize the difference between y and ypred, works to reduce all the three including the regularization component
Here the ϴ is at times moved to 0 so some features could be removed, as the formula doesnt takes the Square of Cofficient only the magnitude is taken into consideration
Hence this is used for Feature Selection, not for model optimization
Model Optimization : Reduce over fitting
Method:
sklearn.linear_model.Lasso
Increase in the penalty will remove the features, so we need to keep this in mind
Ridge Reguralization
Works same as Lasso, in Ridge the ϴ may be changed to close to 0 but never the ϴ is never turned to 0, as the coefficient is always squared
Hence Not suitable for feature selection, but used for model optimization
Model Optimization is done to reduce over fitting
Formula :

Tree Based
Decision Tree
Most popular way to use in the feature selections
Highly accurate in training set
Provides good generalization
Helps in bringing Interpretobility
Algorithm
First Node: Highest Impurity
Impurity decreases further each level
Until the impurity is 0
More the feature decreases the impurity, the more the the feature importance
Random Forest
Contains several decision trees
Decrease of impurity for each feature is averaged accross trees
Algorithm
Build random forests
Calculate feature importance
Remove least important feature
Repeat till condition is met
Less of real life use, although doesn't give more benifit in removing feature
Limitation
If correlated are not removed both the variable will show equal low importance
If a feature removed is correlated to another feature in the dataset, by removing the correlated features the tree importance of the feature will be revealed
Highly baised with the categorical variables
Reference and Credits
https://www.datacamp.com/community/tutorials/feature-selection-python
https://towardsdatascience.com/feature-importance-and-forward-feature-selection-752638849962
https://miro.medium.com/max/1400/1*MO2JGos1gZnyfQueFrzzsw.png






















