Introduction
We use history for our future prediction but in reality it could fail as continuous change occurs every and the historical data is not general enough for the accurate prediction. In the past, most top-level executive positions were occupied by men. However, it is important to ensure that qualified candidates of diverse backgrounds are also given equal opportunities. Another example is that the majority of loan applications, around 90%, are often rejected, leaving only 10% approved. This skewed ratio can lead to incorrect decisions, resulting in good loan applications being unfairly rejected. When we use these unbalanced data for model learn (get trained), could result in bias output supporting the high class in prediction. In the real world we have many such scenarios where we can observe imbalance in data.
Another one reason of biased results is the way we collect data. Sampling is a method of data collection, a small subset of the population. One of the biggest problems with sampling is that if it is done in an imbalanced way, resulting biased data.
To counter such imbalanced datasets, we use a techniques called up-sampling and down-sampling. Let's see more in detail below.
Upsampling
Upsampling is a technique that is used to increase the number of instances in a dataset. It is commonly used when there is an imbalanced distribution of data, with one class being significantly underrepresented. This process involves generating synthetic data points using different techniques or generative models that are similar to the existing data and adding them to the dataset. However, it is important to ensure that by adding the new dataset, the model should not perform overfitting. After this process, the count of the classes is almost balanced.
Methods
Upsampling can be used in scenarios where the dataset has an imbalanced distribution or is small. To address class imbalance in machine learning, one effective technique is upsampling. This involves oversampling the minority class by duplicating its examples in the training dataset before fitting the model. While this balances the class distribution, it doesn’t offer any new information to the model. Upsampling may be more effective in addressing the problem and improving model performance, especially when identifying rare events or anomalies. It is worth noting that the methods explained below are not exhaustive, but they provide some practical insights.
Resample
- Resample, does upsample a dataset by simply copying records from minority classes.
- Logic
- The default strategy involves executing the step of the bootstrapping procedure.
- The bootstrap method involves creating subset from dataset multiple times.
- From the size choosen it basically creates a random selects with or without replacement of the dataset.
- Sampling without replacement, in which a subset of the observations are selected randomly, and once an observation is selected it cannot be selected again.
- Sampling with replacement, in which a subset of observations are selected randomly, and an observation may be selected more than once.
- The resample function is specifically used to resample a dataset n_samples times, with the default option being to sample with replacement.
- Sample Code
from collections import Counter
from sklearn.datasets import make_classification
from sklearn.utils import resample
from matplotlib import pyplot
from numpy import where
import pandas as pd
import numpy as np
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
# perform upsampling
df = pd.DataFrame(X, columns=['Column_A', 'Column_B', 'Column_C', 'Column_D'])
df['Y'] = y
over_samples = df[df["Y"] == 1]
under_samples = df[df["Y"] == 0]
print("Original Under Sample Shape", under_samples.shape)
df_res = resample(under_samples,
replace=True,
n_samples=len(over_samples),
random_state=42)
print("Resampled Under Sample Shape", df_res.shape)
X_Res = np.append(df_res[['Column_A', 'Column_B', 'Column_C', 'Column_D']].to_numpy(), over_samples[['Column_A', 'Column_B', 'Column_C', 'Column_D']].to_numpy(), axis=0)
y_res = np.append(df_res['Y'].to_numpy(), over_samples['Y'].to_numpy())
print('Resampled dataset shape %s' % Counter(y_res))
RandomOverSampler
- Randomly selecting the rows from minority class and performing duplication.
- Logic
- Random oversampling involves duplicating examples from the minority class and adding them to the training dataset.
- This is done by making exact copies of the minority class examples, which are selected randomly with replacement from the training dataset.
- Source Code
from collections import Counter
from sklearn.datasets import make_classification
from imblearn.over_sampling import RandomOverSampler
from matplotlib import pyplot
from numpy import where
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
ros = RandomOverSampler(random_state=24)
X_Res, y_res = ros.fit_resample(X, y)
print('Resampled dataset shape %s' % Counter(y_res))
Resampled dataset shape Counter({1: 894, 0: 894})# scatter plot of examples by class label
for label, _ in Counter(y_res).items():
row_ix = where(y_res == label)[0]
pyplot.scatter(X_Res[row_ix, 0], X_Res[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
SMOTE (Synthetic Minority Oversampling Technique)
- The Synthetic Minority Oversampling Technique (SMOTE) is a widely used method for creating new examples. It involves selecting examples that are similar in the feature space, drawing a line between them, and generating a new sample at a point along that line.
- Logic
- Works based on KNearest Neighbours Algorithm, by generating the synthetic data points which falls under the existing of minority class.
- To be more specific, SMOTE randomly selects an instance from the minority class and identifies its k nearest minority class neighbours (usually k=5). Then, it randomly selects a neighbour from the k nearest neighbours and creates a synthetic example at a randomly chosen point between the two examples in the feature space.
- To generate the synthetic instances, SMOTE uses a convex combination of the two selected instances, a and b, which are connected by a line segment in the feature space.
- Input record should not contain null values.
- Visualization
![]() |
| Credits: https://rikunert.com/wp-content/uploads/2017/11/the-basic-principle-of-the-synthetic-minority-oversample-technique-smote-algorithm-5452514.png |
- Source Code
from collections import Counter
from sklearn.datasets import make_classification
from imblearn.over_sampling import SMOTE
from matplotlib import pyplot
from numpy import where
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
# perform upsampling
sm = SMOTE(random_state=42)
X_res, y_res = sm.fit_resample(X, y)
print('Resampled dataset shape %s' % Counter(y_res))
Resampled dataset shape Counter({1: 894, 0: 894})# scatter plot of examples by class label post upsampling
for label, _ in Counter(y_res).items():
row_ix = where(y_res == label)[0]
pyplot.scatter(X_res[row_ix, 0], X_res[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
ADASYN (Adaptive Synthetic)
- Oversampling can be done using the Adaptive Synthetic (ADASYN) algorithm. This algorithm works similarly to SMOTE(can be said as extension), but it generates a varying number of samples based on the local distribution of the class that needs to be oversampled.
- Logic
- ADASYN is a technique that generates synthetic samples for minority classes using the feature space of the original dataset. It calculates the density distribution of each minority class sample and generates synthetic samples according to the density distribution. The main idea behind ADASYN is to use a weighted distribution for different minority class examples based on their level of difficulty in learning. It generates more synthetic data for minority class examples that are harder to learn compared to those that are easier to learn.
- Randomly selects an instance from the minority class and identifies its k nearest minority class neighbours (usually k=5).
- Find the ratio ri = majority_class in neighbourhood / K
- The rᵢ value indicates the dominance of the majority class in each specific neighbourhood. Higher rᵢ neighbourhoods contain more majority class examples and are more difficult to learn.
- Because rᵢ is higher for neighbourhoods dominated by majority class examples, more synthetic minority class examples will be generated for those neighbourhoods. This is what gives the ADASYN algorithm its adaptive nature, generating more data for harder-to-learn neighbourhoods.
- The main difference between ADASYN and SMOTE is that ADASYN generates synthetic samples adaptively based on the density distribution of minority class samples, while SMOTE generates synthetic samples by interpolating between minority class samples. This adaptive approach helps to focus more on difficult-to-learn samples, which can potentially lead to better classification performance.
- However, ADASYN has a limitation of increased computational complexity due to the generation of synthetic samples. This may affect the training time of machine learning models.
- Visualization
![]() |
| Credits: https://miro.medium.com/v2/resize:fit:720/format:webp/1*qgPPSwNO8XWqCFTTe1rGwg.jpeg |
- Sample Code
from collections import Counter
from sklearn.datasets import make_classification
from imblearn.over_sampling import ADASYN
from matplotlib import pyplot
from numpy import where
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
# perform upsampling
ada = ADASYN(random_state=42)
X_res, y_res = ada.fit_resample(X, y)
print('Resampled dataset shape %s' % Counter(y_res))
Resampled dataset shape Counter({0: 897, 1: 894})Pros
- Helps in increasing underrepresented class dataset
- Making model less bias
- Cost effective as does not requires additional dataset collection
- Overfitting when the synthetic data points are very much similar to existing data points
- Quality of synthetic data could result in impacting model accuracy negatively
Downsampling
Downsampling is a method used to decrease the size of a dataset by removing some of the instances. When the data is imbalance remove the data points of the majority class. This is often done to reduce computational complexity and training time, and to eliminate redundancy or irrelevant instances from the dataset. While reduction need to make sure the important information is not lost and the sample represent the entire dataset.
Methods
Scenarios to use upsampling is when the dataset is large, downsampling may be a good option to reduce the computational complexity and training time of the model. If the goal is to improve model efficiency or reduce the risk of overfitting, downsampling may be a better option. Below are few methods explained, though these are not exhaustive list.
Resample
- Similar to upsampling the same method can be used for downsampling as well. The simplest way of downsampling majority classes is by randomly removing records from that category.
- Logic is same as upsampling, instead of increase the randomly records are removed.
- Sample Code
from collections import Counter
from sklearn.datasets import make_classification
from sklearn.utils import resample
from matplotlib import pyplot
from numpy import where
import pandas as pd
import numpy as np
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
# perform downsampling
df = pd.DataFrame(X, columns=['Column_A', 'Column_B', 'Column_C', 'Column_D'])
df['Y'] = y
over_samples = df[df["Y"] == 1]
under_samples = df[df["Y"] == 0]
print("Original Over Sample Shape", under_samples.shape)
df_res = resample(over_samples,
replace=True,
n_samples=len(under_samples),
random_state=42)
print("Resampled Over Sample Shape", df_res.shape)
X_Res = np.append(df_res[['Column_A', 'Column_B', 'Column_C', 'Column_D']].to_numpy(), under_samples[['Column_A', 'Column_B', 'Column_C', 'Column_D']].to_numpy(), axis=0)
y_res = np.append(df_res['Y'].to_numpy(), under_samples['Y'].to_numpy())
print('Resampled dataset shape %s' % Counter(y_res))
# scatter plot of examples by class label
for label, _ in Counter(y_res).items():
row_ix = where(y_res == label)[0]
pyplot.scatter(X_Res[row_ix, 0], X_Res[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
RandomUnderSampler
- Randomly selecting the rows from majority class and performing replacement.
- Logic is same as upsampling, except the strategies impact the majority class instead of the minority class.
- Source Code
from collections import Counter
from sklearn.datasets import make_classification
from imblearn.under_sampling import RandomUnderSamplerfrom matplotlib import pyplot
from numpy import where
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
# perform downsampling
rus = RandomUnderSampler(random_state=42)
X_Res, y_res = rus.fit_resample(X, y)
print('Resampled dataset shape %s' % Counter(y_res))
Resampled dataset shape Counter({0: 106, 1: 106})Cluster Centroid
- Centroid algorithm uses KMeans method to replace the samples under the majority class.
- This method handles the imbalanced datasets and involves under-sampling the majority class by replacing a group of majority samples with the cluster centroid of a KMeans algorithm. This newly generated set is created using the centroids of the K-means method instead of the original samples. By doing this, the majority class(es) are transformed while the minority class remains unchanged.
- Logic
- Choose random data from the majority class.
- Calculate the Euclidean distance between the random data and its k nearest neighbours.
- Calculate centroids for clusters and use them to represent the majority class, reducing samples.
- Repeat the procedure until the desired proportion of minority class is met.
- Sample Code
from collections import Counter
from sklearn.datasets import make_classification
from sklearn.utils import resample
from imblearn.under_sampling import ClusterCentroids
from sklearn.cluster import MiniBatchKMeans
from matplotlib import pyplot
from numpy import where
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
# perform downsampling
cc = ClusterCentroids(
estimator=MiniBatchKMeans(n_init=1, random_state=0), random_state=42
)
X_Res, y_res = cc.fit_resample(X, y)
print('Resampled dataset shape %s' % Counter(y_res))
Resampled dataset shape Counter({0: 106, 1: 106})# scatter plot of examples by class label
for label, _ in Counter(y_res).items():
row_ix = where(y_res == label)[0]
pyplot.scatter(X_Res[row_ix, 0], X_Res[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
Near Miss
- Near Miss, where the majority class instances that are near to the minority class instances are removed.
- There are three versions of this algorithm, each version has a different criterion for selecting which majority class instances to keep and which ones to discard :
- NearMiss-1: The data is balanced by selecting the samples from the majority class for which the average distance of the k nearest samples of the minority class is the smallest.
- NearMiss-2: The data is balanced by selecting the samples from the majority class for which the average distance to the farthest samples of the minority class is the smallest.
- NearMiss-3: This is a 2-step algorithm: first, for each minority sample, their m nearest-neighbours will be kept; then, the majority samples selected are the on for which the average distance to the k nearest neighbours is the largest.
- Logic
- This algorithm has two important parameters - version and k neighbours
Calculates the distance between all the points in the larger class with the points in the smaller class.
- Based on the passed parameter perform elimination.
- NearMiss-1: This version keeps instances from the majority class that are closest to the instances in the minority class.
- NearMiss-2: This version removes instances from the majority class that are farthest from the minority class.
- NearMiss-3: This version is somewhat opposite to NearMiss-1; it keeps instances from the majority class that are farthest from the instances in the minority class.
- Sample Code
from collections import Counter
from sklearn.datasets import make_classification
from sklearn.utils import resample
from imblearn.under_sampling import NearMiss
from matplotlib import pyplot
from numpy import where
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
# perform downsampling
cc = NearMiss()
X_Res, y_res = cc.fit_resample(X, y)
print('Resampled dataset shape %s' % Counter(y_res))
Resampled dataset shape Counter({0: 106, 1: 106})# scatter plot of examples by class label
for label, _ in Counter(y_res).items():
row_ix = where(y_res == label)[0]
pyplot.scatter(X_Res[row_ix, 0], X_Res[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
Tomek (T-Links)
- Here a pair of data points from different classes (nearest-neighbours) are dropped, the objective is to drop the sample that corresponds to the majority and thereby minimalizing the count of the dominating label.
- This also increases the border space between the two labels and thus improving the performance accuracy.
- Logic
- Identifying Tomek Links:
- A Tomek link is a pair of instances (x, y) where x is from the majority class and y is from the minority class, leaving no other smaller instance between.
- Removing Instances in Tomek Links:
- Once Tomek links are identified, the instances from the majority class in these links are removed.
- This helps in creating a cleaner and more separable boundary between the classes.
- Resulting Effect:
- The removal of instances involved in Tomek links can enhance the performance of a machine learning model by reducing noise in the dataset and improving the separation between classes.
- Representation
![]() |
| Credits: https://www.analyticsvidhya.com/blog/2020/11/handling-imbalanced-data-machine-learning-computer-vision-and-nlp/ |
- Source Code
from collections import Counter
from sklearn.datasets import make_classification
from sklearn.utils import resample
from imblearn.under_sampling import TomekLinks
from matplotlib import pyplot
from numpy import where
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# perform downsampling
cc = TomekLinks()
X_Res, y_res = cc.fit_resample(X, y)
print('Resampled dataset shape %s' % Counter(y_res))
Resampled dataset shape Counter({1: 890, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y_res).items():
row_ix = where(y_res == label)[0]
pyplot.scatter(X_Res[row_ix, 0], X_Res[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
AIIKNN (All K-Nearest-Neighbours)
- It works by removing some of the instances from the majority class that are considered to be "misclassified" or near the decision boundary.
- The goal is to create a more balanced dataset by eliminating instances from the majority class that may be causing confusion for the classifier.
- Logic
- AllKNN is a modification of RepeatedEditedNearestNeighbours (RENN) that expands the size of the neighbourhood being considered each time it is run.
- The RENN algorithm runs the EditedNearestNeighbour (ENN) algorithm a specified number of times, removing more outlier points with each rerun.
- ENN removes misclassified instances near the decision boundary from the majority class.
- Sample Code
from collections import Counter
from sklearn.datasets import make_classification
from sklearn.utils import resample
from imblearn.under_sampling import AllKNN
from matplotlib import pyplot
from numpy import where
X, y = make_classification(n_classes=2, class_sep=2,
weights=[0.1, 0.9], n_features=4, n_clusters_per_class=1, n_samples=1000, random_state=10)
# summarize class distribution
print('Original dataset shape %s' % Counter(y))
Original dataset shape Counter({1: 894, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y).items():
row_ix = where(y == label)[0]
pyplot.scatter(X[row_ix, 0], X[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
# perform downsampling
allknn = AllKNN()
X_Res, y_res = allknn.fit_resample(X, y)
print('Resampled dataset shape %s' % Counter(y_res))
Resampled dataset shape Counter({1: 872, 0: 106})# scatter plot of examples by class label
for label, _ in Counter(y_res).items():
row_ix = where(y_res == label)[0]
pyplot.scatter(X_Res[row_ix, 0], X_Res[row_ix, 1], label=str(label))
pyplot.legend()
pyplot.show()
- Helps in reducing the bias where the class is over populated
- Reduce noise and improve accuracy
- Saves training time and reduce computational power
- Can result in information loss
- Post downsampling remaining dataset can lead to overfitting or biased results
Conclusion
Upsampling is increase on the minority class and downsampling is vice-versa decrease on the majority class.
Using these equalization process prevents the model from inclining towards the majority class. Points to remember, the final result of up-sampling and down-sampling needs to preserve the distribution of the data and boundary line between the target classes remains same. While performing up-sampling and down-sampling, these techniques are only applied to the training datasets and no changes are made to validation and testing data.
Here we have discussed only few machine learning approaches, but similar problem exists in computer vision and natural language programming (NLP). Just as teaser, lets me share for computer vision - image augmentation and nlp - class weights are some common methods, and there are more advanced methods to generate synthetic data for better predictions.
Credits and References
- https://media.istockphoto.com/id/526326349/vector/sphere-seesaw-imbalance-horizontal-contrasts-comparison.jpg?s=612x612&w=0&k=20&c=nvfqdfaD4yzS8z9CGVWkooxQRHn0PAVcrE2goUfxekI=
- https://imbalanced-learn.org/stable/references/over_sampling.html
- https://imbalanced-learn.org/stable/references/under_sampling.html
- https://medium.com/@rithpansanga/choosing-the-right-size-a-look-at-the-differences-between-upsampling-and-downsampling-methods-daae83915c19
- https://www.nomidl.com/machine-learning/what-is-upsampling-and-downsampling/
- https://www.analyticsvidhya.com/blog/2020/11/handling-imbalanced-data-machine-learning-computer-vision-and-nlp/
- https://elitedatascience.com/imbalanced-classes
- https://arxiv.org/pdf/1106.1813.pdf
- https://www.kaggle.com/code/residentmario/advanced-under-sampling-and-data-cleaning
- https://machinelearningmastery.com/undersampling-algorithms-for-imbalanced-classification/
- https://medium.com/@ruinian/an-introduction-to-adasyn-with-code-1383a5ece7aa
- https://analyticsindiamag.com/using-near-miss-algorithm-for-imbalanced-datasets/

.png)

.png)
.png)


.png)

.png)
.png)

.png)
.png)
.png)
.png)

.png)
