Basics of Probability
- #GoodToKnowStuff: Christiaan Huygens published a book on the subject in 1657. In the 19th century, what is considered the classical definition of probability was completed by Pierre Laplace. - History from wiki
- Probability
- Probability is a measure that defines the likelihood or possibility of an event occurring, tough some event cannot be predicted with total certainty.
- Probability rules
- Probability = 0 means outcome will never happen
- Probability = 1 means outcome will always happen
- Always probability values lie between Probability >=0 and Probability <=1
- Random Variables
- A random variable is a variable with an unknown value or a function that gives values to each of the results of an experiment. Basically when the outcome is unknown, it is referred to random variable.
- Two Types of Random Variables
- Discrete - Having specific values such as toss is coin can always give head or tail.
- Continuous - Having any possible value within a range, eg avg salary, temperature.
- Random variables are used to quantify results of random events.
- Event
- An event is single outcome or combination of outcomes with their probability.
- Eg:
- Tossing 1 coin, the result would be either head or tail. P(Head) would be 50%. Here after tossing the coin the result either its head or tail is called as event. This is an example of single outcome, for combination of outcomes lets take an example of tossing two coins together. The outcome here would be the permutation of head and tail of the two coin [HH, HT, TT, TH].
- Event Spaces
- An event space contains all possible events and combination of events for a given experiment or happening.
- The event space is sometimes called the sample space.
- Eg:
- Tossing a coin, the event space is head or tail.
- When rolling a normal six-sided die and recording the uppermost face, the event space is {1,2,3,4,5,6}
- Notations of Probability
- P(A) : Probability of event A occuring
- P(A U B) : P(A or B) Probability of either A or B event occurring
- P(A ∩ B) : P(A and B) Probability of A and B event occurring
- P(A | B) : P( A given B) Probability of A given B event occurs
- Independent Events
- Outcome of two or more events that does not affect each other are called as independent events.
- Eg: Considering tossing two coins, the probability of getting heads does not depend on each other.
- To find if the two events are independent, the following two probability relationships must hold:
- P(A and B) = P(A) * P(B) = P(AB)
- Here the chances of both A and B will happening at a same time which is the P(A and B), is the product of their probabilities P(A) * P(B). This is also called as joint probability.
- Example: What are the chances of receiving double sixes while rolling two dice? The result of the first die roll is event B, while the result of the second die roll is event A. Because each dice has a 1 in 6 chance of landing on a six, and the outcomes are independent, the chances are simply 1/6 multiplied by 1/6.
- P(A | B) = P(A) [and similarly P(B|A) = P(B)]
- Probability of A given B event has occurrent will not change, it will be same as probability of A - P(A).
- Example: What are the odds of getting two sixes if the first die has already finished rolling and is a 6 (event B is a 6). The answer is 1/6 since the probability of the second dice landing on 6 (event A) is 1/6 and we already know B is a six.
- P(A or B) = P(A) + P(B) - P(AB)
- Chances of either getting A or B, as both are independent, there are three possibilities:
- A occurs and B does not occurs
- B occurs and A does not occurs
- Both A and B occurs
- P(A or B) = P(A) + P(B) - P(AB), which means probability of A or B either occurring depends on the probability of A occurring plus probability of B occurring minus probability of AB occurring together to avoid counting the outcomes twice.
- Example: lets understand this with an example of deck of 52 cards,
- P(Spade) = 13/52
- P(Queen) = 4/52
- Here if we simply add both together the probability will be wrong, as Queen is counted twice. To avoid that:
- P(Spade and Queen) = 1/52 (chances are only once)
- Hence the formula is P(A or B) = P(A) + P(B) - P(AB), which is P(Spade or Queen) = P(Spade) + P(Queen) - P(Spade and Queen)
- Mutually Exclusive Events
- Two events which cannot happen together are mutually exclusive events, also known as probability of disjoint.
- Eg:
- When flipping a coin, the outcomes of head and tail are mutually exclusive. Because the chance of acquiring both the head and tail at the same time is nil.
- The events "2" and "5" on a six-sided die are mutually exclusive. When we tossed one dice, we couldn't get both events 2 and 5 at the same time.
- To find when two events are mutually exclusive, the following probability relationships must hold:
- P(AB) = 0
- Stating that probability of both the events occurring together will be zero.
- P(A or B) = P(A) + P(B)
- In general the formula is P(A or B) = P(A) + P(B) - P(AB), as there is a possibility of event A and B occurring together, but in mutually exclusive scenario this is not possible.
- Mutually exclusive example:
- P(Queen or King)
- P(Queen) = 4/52
- P(King) = 4/52
- Here the P(Queen and King) occurring together the chances are 0.
- Hence the formula for P(A or B) = P(A) + P(B) is good enough.
- Not mutually exclusive example:
- P(Spade or Queen) : Here there is a possibility of both occurring together hence need to exclude this while calculating to avoid double counting.
- Note: Independent event cannot be mutually exclusive. For independent events, the probability of two events is product of both, but in mutually exclusive the product of two events are always 0.
- Collective/Cumulative Exhaustive Events
- Events that cumulatively describes all the outcomes.
- Conditionally Independent Events
- If two events[P(A), P(B)] are individually independent or not, but when these two events are independent given the occurrence of third event.
- Example:
- Take a random sample of school children and for each child obtain data on:
- Foot Size (F)
- Literacy Score (L)
- The two will be (positively) correlated, in that the bigger the foot size the higher the literacy score.
- The random variables F and L are not independent.
- Obviously a bigger foot size is not the direct cause for a higher literacy score. What correlates the two is the child's age (A), which is the confounder.
- By conditioning on age, we no longer consider the relationship between foot size and literacy for the whole sample, but per each age group separately, and makes foot size and literacy score independent.
- So this was just an example of two random variables F and L that were:
- dependent when not conditioned on A
- independent when conditioned on A
- Hence F is conditionally independent of L given A: P(L|F, A) = P (L| A) And so: P(L|F) = P(L)
- Reference: https://math.stackexchange.com/questions/23093/could-someone-explain-conditional-independence#:~:text=If%20two%20people%20live%20in,an%20effect%20on%20the%20other.
- Discrete Probability Function
- A discrete probability distribution counts occurrences that have countable or finite outcomes.
- Example: Consider you have 1 - green ball, 2 - yellow balls, 3 - blue balls and 4 - red balls, total 10 balls. When randomly we pick a ball blindfolded, and the finite possible outcomes are green, yellow, blue and red. Probability function for this is x/10, such as probability of green ball occurring is 1/10 = 10%, yellow ball 2/10 = 20%, blue ball 3/10 = 30% and red ball 4/10 = 40% and the total of these together is 100%.
- Continuous Probability Function
- A continuous probability distribution counts occurrences that have any value or infinite outcomes.
- Probabilities are measured over intervals, not single points.
- Example: Blood pressure levels
- Unconditional Probability
- An unconditional probability is the likelihood that one out of several possible results will occur.
- Unconditional is also sometime called as marginal.
- Here possibility that an event will occur regardless of whether other events have occurred or other conditions exist. Ideally does not depend of the prior event.
- Example: P(Queen) is the probability of an event Queen returning
- Conditional Probability
- The probability of an event occurring given the outcome of another event.
- This can also be used to incorporate additional information (which is given event) to update unconditional probabilities. Here the given additional information could gain better prediction. As conditional probability operates as if B is the event space and A is an event inside this new restricted space.
- P(A|B) defines as probability of event A occurring given B event has occurred.
- Example: P(Queen|Spades) is the probability of card returning Queen given the cards are Spades.
- General Formula: P(A|B) = P(AB) / P(B)
- In simple terms could also be said as P(Conditional) = P(Joint)/P(Unconditional)
- Example:
- P(Queen | Spades) = P(Queen and Spades) / P(Spades)
- P(Queen | Spades) = (1/52) / (13/52)
- P(Queen | Spades) = (1/13) #Reality check, this is true as Queen occurs only once in Spades
- Formula For Independent : P(A|B) = P(A)
- As there is no dependence of A giving B occurred.
- Joint Probability
- The probability of two events occurring together.
- P(AB) is referred as both A and B event occurring at the same time.
- The joint probability of two independent events is the product of the probability of each.
- Formula: P(AB) = P(A) * P(B)
- Example:
- P(Queen and Spades)
- = P(Queen) * P(Spades)
- = (4/52) * (13/52)
- = (1/13) * (1/4)
- = (1/52) #Reality check, this is true as there is only chance of this to occur in whole set of decks Queen and Spade occurring together is only once.
- Conditional probability can be used to calculate joint probability, deriving from the above conditional formula.
- Formula: P(AB) = P(A|B) * P(B)
- Example:
- P(Queen and Spades)
- = P(Queen | Spades) * P(Spades)
- = (1/13) * (13/52)
- = (1/13) * (1/4)
- = (1/52)
- Joint probability depends on the two events acting independently from one another. To determine whether they are truly independent, it's important to establish whether one's outcome affects the other. If they do, they are dependent, which means they lead to conditional probability. If they don't, you end up with joint probability.
- Law of Total Probability
- In cases where the probability of occurrence of one event depends on the occurrence of other events, we use the law of total probability theorem.
Credits: https://www.embibe.com/exams/theorems-on-probability/
- Total probability theorem indicates the probability that an event will occur given each of the sample space’s partitions
- Total Probability Theorem Formula:
- P(A)=P(A∩B)+P(A∩B^)
- Bayes Theorem
- Bayes' Theorem was named after 18th-century mathematician Thomas Bayes.
- Bayes formula is used determine conditional probability, based on the previous outcome of the similar circumstances.
- In other words Bayes formula uses basic ideas from probability to develop a precise expression for how information about one event can be used to learn from another event.
- Formula:
- P(A|B) = P(AB) / P(B)
- Here,
- P(AB) = P(B|A) * P(A)
- P(AB) is the probability of event A and B occurring together
- P(A ∣ B) is the conditional probability of event A occurring, given that B is true.
- P(B ∣ A) is the conditional probability of event B occurring, given that A is true.
- P(A) and P(B) are the probabilities of A and B occurring independently of one another.
- The Bayes theorem is based on finding P(A | B) when P(B | A) is given.
- While conditional probability deals with the probability of an event given another event, Bayes' Theorem is a more specific application that involves updating probabilities based on new evidence. Bayes' Theorem is a powerful tool in Bayesian statistics and is widely used in fields such as machine learning, medical diagnosis, and other areas where probabilistic reasoning is essential.
Recap with an Example
1000 mobiles are given to a company for testing, with following features:
- 600 smart foldable phones
- Of the 600 smart foldable phones 150 are AI enabled
- 400 smart flip phones
- Of the 400 smart flip phones 200 are AI enabled
Unconditional Probabilities
- P(Smart Foldable Phones) = 600/1000 = 60%
- P(Smart Flip Phones) = 400/1000 = 40%
Conditional Probabilities
- P(AI Enabled | Smart Foldable Phones) = 150/600 = 25%
- P(AI Enabled | Smart Flip Phones) = 200/400 = 50%
Joint Probabilities [Formula: P(AB) = P(A|B) * P(B)]
- P(AI Enabled and Smart Foldable Phones)
- (150/600) * (600/1000) = 25%*60% = 15%
- 15% of 1000 phones = 150 of smart foldable phones are AI enabled
- P(AI Enabled and Smart Flip Phones)
- (200/400) * (400/1000) = 50%*40% = 20%
- 20% of 1000 phones = 200 of smart flip phones are AI enabled
Total Probability
- P(AI Enabled Phones) =
- P(AI Enabled and Smart Foldable Phones) + P(AI Enabled and Smart Flip Phones)
- 15% + 20% = 35%
- 35% of 1000 phones = 350 phones are AI enabled
Bayes Rule [Formula: P(A|B) = P(AB) / P(B)]
- P(Smart Foldable Phones | AI Enabled) / P(AI Enabled)
- 15% / 35% = 42.9%
- 350 AI enabled phones, and of them 150 are smart foldable phones : 150 / 350 = 42.9%
Bayes Theorem Example
Suppose there are three types of managers: the underperformers beat the market only 25% of the time, the in-line performers beat the market 50% of the time, and the outperformers beat the market 75% of the time. Our prior belief is that a manager has a 60% probability of being an in-line performer, a 20% chance of being an underperformer, and a 20% chance of being an outperformer. If the manager beats the market two years in a row what should our updated beliefs be?
Our prior beliefs can be summarized as: P[MO] = 20%, P[MI] = 60% and P[MU] = 20%.
A Binomial Tree:
Associated Probability Matrix:
Above matrix is defined by deriving, for example MO beats market is 15% which is 75% of 20%.
The unconditional probability of observing the manager beat the market two years in a row, given our prior beliefs is P(2B) = 27.50%.
P(2B) = P(2B|MO) * P(MO) + P(2B|MI) * P(MI) + P(2B|MU) * P(MU)
Using Bayes’ theorem, for example, we can calculate our posterior belief that the manager is an outperformer as 40.91%:
P(MO|2B) = (P(2B|MO) * P(MO)) / (P2B) = 0.56 × 0.20/0.275 = 40.91%
Similarly, we can show that the posterior probability that the manager is an in-line performer is 54.55% and that the posterior probability that the manager is an underperformer is 4.55%.
As we would expect, given that the manager beat the market two years in a row, the posterior probability that the manager is an outperformer has increased, from 20% to 40.91%, and the posterior probability that the manager is an underperformer has decreased, from 20% to 4.55%.
Although the probabilities changed, the sum of the probabilities remains equal to 100%.
Credits: Bionic Study Notes - This reference very much helps in understanding bayes theorem concept.
Description
- Supervised Classification Model
- Naive Bayes, calculates the probability of every X factor, then selects the highest probability
- Naive Bayes, is based on the Bayes Theorem
- Bayes Theorem :
- P(A|B) = P(B|A)P(A) / P(B)
- P(A|B) : Probability of A happening, given the condition of B has occurred
- B is the evidence
- A is the hypothesis
- P(B) : Probability of B happening
- P(A) : Probability of A happening
- P(B|A) : Probability of B happening, given the condition of A has occurred
- Naive Bayes algorithms are mostly used in
- sentiment analysis
- spam filtering
- recommendation systems
- Types of Naive Bayes Classifier :
- Multinomial
- Predictors are categorical, more than two
- Usually used in document classification problem,
- eg : whether the document belongs to sports, health, technology, etc
- Bernoulli
- Predictors are boolean
- eg: spam or ham email
- Gaussian
- Used when the features are normally distributed
Formula
- Formula is derived from Bayes Theorem
- P(Hypothesis | Evidence) = ( P(Evidence | Hypothesis) * P(Hypothesis) ) / P(Evidence)
- Hypothesis = prediction variable
- These are independent features
- Evidence 1, Evidence 2, Evidence 3, ... Evidence n
- P(x| Hypothesis)
- Likelihood
- Probability of the evidence, given the belief is true.
- P(Hypothesis)
- Class Prior Probability
- The probability before the evidence is considered.
- P(Evidence)
- Predictor Prior Probability
- Probability of the evidence, under any circumstance.
- P(Hypothesis | Evidence)
- Posterior Probability
- Before the actual event occurred predicting given the x event has occurred, with the assumption event c could be related to the event x.
- So if we have multiple parameters, the formula would be
Max probability is taken, after multiple prediction values
Assumption
- Each independent variable should be not correlated
- eg: if its high sunny day, humidity would also be high
- Equal or Weighted importance given to each independent variable,
- eg: to predict if it will rain or not, if we would give equal importance to sunny, humidity
- eg: to predict if it will rain or not, if we would give more importance to cloudy than humidity
- The word 'Naive' is used because of the above two assumptions
Algorithm Explained with an Example
Problem Statement:
From the given attributes, predict if the car was stolen or not, with the assumption the given attributes are independent with no correlation and equal importance.
Dataset:
Model Training:
First step is to calculate the Likelihood for each attribute against the target, by molding the frequency table to likelihood table.
Color:
Independent probability of each color:
P(Red) = 5/10 = 1/2
P(Yellow) = 5/10 = 1/2
Frequency and Likelihood Table of the attribute: P(Color | Stolen)
Independent probability of each type:
P(Sports) = 6/10 = 3/5
P(SUV) = 4/10 = 2/5
Frequency and Likelihood Table of the attribute: P(Type | Stolen)
Independent probability of each origin:
P(Domestic) = 5/10 = 1/2
P(Imported) = 5/10 = 1/2
Frequency and Likelihood Table of the attribute: P(Origin | Stolen)
Class Prior Probability:
In our example we have two possible outcomes for Stolen i.e. Yes or No.
P(Yes) = 5/10 = 1/2
P(No) = 5/10 = 1/2
Model Testing:
Lets test for the given input:
Using Formula:
Given:
- P(Yes | Red) : 3/5
- P(Yes | SUV) : 1/5
- P(Yes | Domestic) : 2/5
- P(Yes) : 1/2
- P(Red) : 1/2
- P(SUV) : 2/5
- P(Domestic) : 1/2
Apply:
P(Yes | Red, SUV, Domestic)
= P(Yes | Red) * P(Yes | SUV) * P(Yes | Domestic) * P(Yes) / P(Red) * P(SUV) * P(Domestic)
= ((3/5) * (1/5) * (2/5) * (1/2)) / ((1/2) * (2/5) * (1/2))
= .24 = 24%
Calculate the posterior probability P(No | X): P(No | Red, SUV, Domestic)
Given:
- P(No | Red) : 2/5
- P(No | SUV) : 3/5
- P(No | Domestic) : 3/5
- P(No) : 1/2
- P(Red) : 1/2
- P(SUV) : 2/5
- P(Domestic) : 1/2
Apply:
P(No | Red, SUV, Domestic)
= P(No | Red) * P(No | SUV) * P(No | Domestic) * P(No) / P(Red) * P(SUV) * P(Domestic)
= ((2/5) * (3/5) * (3/5) * (1/2)) / ((1/2) * (2/5) * (1/2))
= .72 = 72%
Now lets predict, y = Max of [ P(Yes | X), P(No | X)]:
y = Max( Yes = .24 and No = .72)
y = No
Hence our example gets classified as ’NO’ the car is not stolen.
Credits: https://www.kdnuggets.com/2020/06/naive-bayes-algorithm-everything.html
Zero Frequency Problem
One of the disadvantage of the above explained problem is when there are 0 values in the frequency. For example lets say we have one more column added = 'Kilometers Run'. Here for the brand new car it would be 0. And this will get a zero when all the probabilities are multiplied.
An approach to overcome this 'zero-frequency problem' in a Bayesian environment is to add one to the count for every attribute value-class combination when an attribute value doesn’t occur with every class value. Now the 'Kilometers Run' for the brand new car will be one, including the cars which already have values in it (for example if the car has already ran 200kms will now become 201kms) and this will solve the problem of 'zero frequency problem'.
Evaluation
- ROC
- ROC stands for Receiver Operating Characteristics Curve
- Used to find performance matrix of classification model, under various Threshold parameters
- This curve is plot using two parameters:
- True Positive Rate
- TPR = TP / (TP + FN)
- Intuitively, Percentage of identifying actual true values
- False Positive Rate
- FPR = FP / (FP + TN)
- Intuitively, Percentage of identifying actual false values
- AUC
- AUC stands for Area Under ROC Curve
- Also written as AUROC (Area Under the Receiver Operating Characteristics)
- AUC ranges in value from 0 to 1. A model whose predictions are 100% wrong has an AUC of 0.0; one whose predictions are 100% correct has an AUC of 1.0.
- If the value is 1, it means our model is overfitted and if the value is 0 it means it underfitted
- But Higher the AUC value, better the performance
- Accuracy :
- Percentage of correct predictions
- But in real life scenarios, we may be more keen in looking for precision and recall
- Because eg: we have a scenario where in when need to find if the there is fraud transaction, as out of all the transactions we would have only 1% of fraud transactions, we may have good results in accuracy, but our intention would be to find less False Negatives.
- Accuracy of model: TP+TN / (TP+TN+FP+FN)
- Misclassification
- Percentage of incorrect predictions
- Formala :
- 1 – Accuracy
- or
- FP+FN / (TP+TN+FP+FN)
- Classification Matrix
- Precision
- The precision is the ratio tp / (tp + fp) where tp is the number of true positives and fp the number of false positives.
- The precision is intuitively the ability, indentify the Type I error, find score based on how False Positive value 0-1, can be converted to percentages.
- Recall
- The recall is the ratio tp / (tp + fn) where tp is the number of true positives and fn the number of false negatives.
- The recall is intuitively the ability, identify the Type II error, find score based on how False Negative value 0-1, can be converted to percentages.
- Specificity
- The specificity is the ratio tn / (tn + fn) where tn is the number of true negatives and fn the number of false negatives.
- The recall is intuitively the ability, negative correct prediction, could also consider this as Type II error
- F1 score
- The harmonic mean of the precision and recall.
- F1 score reaches its best value at 1 and worst score at 0.
- Confusion Matrix
- An easy and popular method of diagnosing model performance.
- TN : True Negative
- TP : True Positive
- FN : False Negative - Type II
- FP : False Positive - Type I
- Always Type I and Type II models are inversely proportions, intuitively if Type I increases Type II error decreases
Example
import pandas as pdimport numpy as npimport matplotlib.pyplot as pltimport seaborn as sns
%matplotlib inline
import warningswarnings.filterwarnings('ignore')
from scipy import statsfrom scipy.stats import zscorefrom sklearn import metricsfrom sklearn.model_selection import train_test_split
from sklearn.naive_bayes import GaussianNB
# Install a pip package in the current Jupyter kernelimport sys!{sys.executable} -m pip install mlxtend
from mlxtend.plotting import plot_confusion_matrix
df = pd.read_csv('Bank_Personal_Loan_Modelling.csv', delimiter=',')
## obervations :## no NAN values## values are only either int or float## total 5000 recordsdf.info()'''<class 'pandas.core.frame.DataFrame'>RangeIndex: 5000 entries, 0 to 4999Data columns (total 14 columns):ID 5000 non-null int64Age 5000 non-null int64Experience 5000 non-null int64Income 5000 non-null int64ZIP Code 5000 non-null int64Family 5000 non-null int64CCAvg 5000 non-null float64Education 5000 non-null int64Mortgage 5000 non-null int64Personal Loan 5000 non-null int64Securities Account 5000 non-null int64CD Account 5000 non-null int64Online 5000 non-null int64CreditCard 5000 non-null int64dtypes: float64(1), int64(13)memory usage: 547.0 KB'''
def calculate_score(model,X_train,y_train,X_test,y_test):
predicted_labels = model.predict(X_train) accuracy_score = metrics.accuracy_score(y_train, predicted_labels) print("Accuracy for Train Data: ") print(accuracy_score)
predicted_labels = model.predict(X_test) print("Actual Labels") print(y_test.as_matrix(columns=None)) print("Predicted Labels") print(predicted_labels) accuracy_score = metrics.accuracy_score(y_test, predicted_labels) print("Accuracy for Test Data: ") print(accuracy_score) print("Confusion Matrix") cm = metrics.confusion_matrix(y_test,predicted_labels) print(cm)
fig, ax = plot_confusion_matrix(conf_mat=cm, show_absolute=True, show_normed=True, colorbar=True) plt.show() print("Classification Matrix") print(metrics.classification_report(y_test,predicted_labels)) result={} result['Accuracy']=accuracy_score result['MSE']=1-accuracy_score return result
# Naive bayes model
naive_model = GaussianNB()naive_model.fit(x_train, y_train)
calculate_score(naive_model, x_train, y_train, x_test, y_test)'''Accuracy for Train Data: 0.8839004861309694Actual Labels[0 0 0 ... 0 0 0]Predicted Labels[1 0 0 ... 0 0 0]Accuracy for Test Data: 0.8806666666666667Confusion Matrix[[1231 131] [ 48 90]]
Classification Matrix precision recall f1-score support
0 0.96 0.90 0.93 1362 1 0.41 0.65 0.50 138
micro avg 0.88 0.88 0.88 1500 macro avg 0.68 0.78 0.72 1500weighted avg 0.91 0.88 0.89 1500
{'Accuracy': 0.8806666666666667, 'MSE': 0.11933333333333329}'''
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
%matplotlib inline
import warnings
warnings.filterwarnings('ignore')
from scipy import stats
from scipy.stats import zscore
from sklearn import metrics
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import GaussianNB
# Install a pip package in the current Jupyter kernel
import sys
!{sys.executable} -m pip install mlxtend
from mlxtend.plotting import plot_confusion_matrix
df = pd.read_csv('Bank_Personal_Loan_Modelling.csv', delimiter=',')
## obervations :
## no NAN values
## values are only either int or float
## total 5000 records
df.info()
'''
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 5000 entries, 0 to 4999
Data columns (total 14 columns):
ID 5000 non-null int64
Age 5000 non-null int64
Experience 5000 non-null int64
Income 5000 non-null int64
ZIP Code 5000 non-null int64
Family 5000 non-null int64
CCAvg 5000 non-null float64
Education 5000 non-null int64
Mortgage 5000 non-null int64
Personal Loan 5000 non-null int64
Securities Account 5000 non-null int64
CD Account 5000 non-null int64
Online 5000 non-null int64
CreditCard 5000 non-null int64
dtypes: float64(1), int64(13)
memory usage: 547.0 KB
'''
def calculate_score(model,X_train,y_train,X_test,y_test):
predicted_labels = model.predict(X_train)
accuracy_score = metrics.accuracy_score(y_train, predicted_labels)
print("Accuracy for Train Data: ")
print(accuracy_score)
predicted_labels = model.predict(X_test)
print("Actual Labels")
print(y_test.as_matrix(columns=None))
print("Predicted Labels")
print(predicted_labels)
accuracy_score = metrics.accuracy_score(y_test, predicted_labels)
print("Accuracy for Test Data: ")
print(accuracy_score)
print("Confusion Matrix")
cm = metrics.confusion_matrix(y_test,predicted_labels)
print(cm)
fig, ax = plot_confusion_matrix(conf_mat=cm,
show_absolute=True,
show_normed=True,
colorbar=True)
plt.show()
print("Classification Matrix")
print(metrics.classification_report(y_test,predicted_labels))
result={}
result['Accuracy']=accuracy_score
result['MSE']=1-accuracy_score
return result
naive_model = GaussianNB()
naive_model.fit(x_train, y_train)
calculate_score(naive_model, x_train, y_train, x_test, y_test)
'''
Accuracy for Train Data:
0.8839004861309694
Actual Labels
[0 0 0 ... 0 0 0]
Predicted Labels
[1 0 0 ... 0 0 0]
Accuracy for Test Data:
0.8806666666666667
Confusion Matrix
[[1231 131]
[ 48 90]]
Classification Matrix
precision recall f1-score support
0 0.96 0.90 0.93 1362
1 0.41 0.65 0.50 138
micro avg 0.88 0.88 0.88 1500
macro avg 0.68 0.78 0.72 1500
weighted avg 0.91 0.88 0.89 1500
{'Accuracy': 0.8806666666666667, 'MSE': 0.11933333333333329}
'''
Advantages
- They are fast and easy to implement.
- Suited when the dimensionality of the inputs is high.
Drawbacks
- Requirement of predictors to be independent.
- If categorical variable has a category (in test data set), which was not observed in training data set, then model will assign a 0 (zero) probability and will be unable to make a prediction. This is often known as “Zero Frequency”. To solve this, we can use the smoothing technique. One of the simplest smoothing techniques is called Laplace estimation.
Credits and References
- https://towardsdatascience.com/naive-bayes-classifier-81d512f50a7c
- https://www.investopedia.com/
- https://medium.com/machine-learning-101/chapter-1-supervised-learning-and-naive-bayes-classification-part-1-theory-8b9e361897d5
- http://www.statsoft.com/textbook/naive-bayes-classifier
- https://machinelearningmastery.com/naive-bayes-for-machine-learning/
- https://www.analyticsvidhya.com/blog/2017/09/naive-bayes-explained/
- https://insightimi.wordpress.com/2020/04/04/naive-bayes-classifier-from-scratch-with-hands-on-examples-in-r/
- https://scikit-learn.org/stable/modules/naive_bayes.html (each naive bayes type model with example)
- https://www.saedsayad.com/images/Bayes_rule.png
- SchweserNotes Quantitative Analysis 2023
- https://www.machinelearningplus.com/predictive-modeling/how-naive-bayes-algorithm-works-with-example-and-full-code/
- https://www.machinelearningplus.com/predictive-modeling/how-naive-bayes-algorithm-works-with-example-and-full-code/
- https://www.kdnuggets.com/2020/06/naive-bayes-algorithm-everything.html














