Thursday, 25 May 2023

Naive Bayes

Basics of Probability

  • #GoodToKnowStuff: Christiaan Huygens published a book on the subject in 1657. In the 19th century, what is considered the classical definition of probability was completed by Pierre Laplace. - History from wiki
  • Probability
    • Probability is a measure that defines the likelihood or possibility of an event occurring, tough some event cannot be predicted with total certainty.
    • Probability rules
      • Probability = 0 means outcome will never happen
      • Probability = 1 means outcome will always happen
      • Always probability values lie between Probability >=0 and Probability <=1
  • Random Variables
    • A random variable is a variable with an unknown value or a function that gives values to each of the results of an experiment. Basically when the outcome is unknown, it is referred to random variable.
    • Two Types of Random Variables
      • Discrete - Having specific values such as toss is coin can always give head or tail.
      • Continuous - Having any possible value within a range, eg avg salary, temperature.
    • Random variables are used to quantify results of random events.
  • Event
    • An event is single outcome or combination of outcomes with their probability.
    • Eg: 
      • Tossing 1 coin, the result would be either head or tail. P(Head) would be 50%. Here after tossing the coin the result either its head or tail is called as event. This is an example of single outcome, for combination of outcomes lets take an example of tossing two coins together. The outcome here would be the permutation of head and tail of the two coin [HH, HT, TT, TH].
  • Event Spaces
    • An event space contains all possible events and combination of events for a given experiment or happening. 
    • The event space is sometimes called the sample space.
    • Eg:
      • Tossing a coin, the event space is head or tail.
      • When rolling a normal six-sided die and recording the uppermost face, the event space is {1,2,3,4,5,6}
  • Notations of Probability
    • P(A) : Probability of event A occuring
    • P(A U B) : P(A or B) Probability of either A or B event occurring
    • P(A ∩ B) : P(A and B) Probability of A and B event occurring
    • P(A | B) : P( A given B) Probability of A given B event occurs
  • Independent Events
    • Outcome of two or more events that does not affect each other are called as independent events.
    • Eg: Considering tossing two coins, the probability of getting heads does not depend on each other.
    • To find if the two events are independent, the following two probability relationships must hold:
      • P(A and B) = P(A) * P(B) = P(AB) 
        • Here the chances of both A and B will happening at a same time which is the P(A and B), is the product of their probabilities P(A) * P(B). This is also called as joint probability.
        • Example: What are the chances of receiving double sixes while rolling two dice? The result of the first die roll is event B, while the result of the second die roll is event A. Because each dice has a 1 in 6 chance of landing on a six, and the outcomes are independent, the chances are simply 1/6 multiplied by 1/6.
      • P(A | B) = P(A) [and similarly P(B|A) = P(B)]
        • Probability of A given B event has occurrent will not change, it will be same as probability of A - P(A).
        • Example: What are the odds of getting two sixes if the first die has already finished rolling and is a 6 (event B is a 6). The answer is 1/6 since the probability of the second dice landing on 6 (event A) is 1/6 and we already know B is a six.
      • P(A or B) = P(A) + P(B) - P(AB)
        • Chances of either getting A or B, as both are independent, there are three possibilities:
          • A occurs and B does not occurs
          • B occurs and A does not occurs
          • Both A and B occurs
        • P(A or B) = P(A) + P(B) - P(AB), which means probability of A or B either occurring depends on the probability of A occurring plus probability of B occurring minus probability of AB occurring together to avoid counting the outcomes twice.
        • Example: lets understand this with an example of deck of 52 cards, 
          • P(Spade) = 13/52
          • P(Queen) = 4/52
          • Here if we simply add both together the probability will be wrong, as Queen is counted twice. To avoid that:
          • P(Spade and Queen) = 1/52 (chances are only once)
          • Hence the formula is P(A or B) = P(A) + P(B) - P(AB), which is P(Spade or Queen) = P(Spade) + P(Queen) - P(Spade and Queen)
  • Mutually Exclusive Events
    • Two events which cannot happen together are mutually exclusive events, also known as probability of disjoint.
    • Eg:
      • When flipping a coin, the outcomes of head and tail are mutually exclusive. Because the chance of acquiring both the head and tail at the same time is nil.
      • The events "2" and "5" on a six-sided die are mutually exclusive. When we tossed one dice, we couldn't get both events 2 and 5 at the same time.
    • To find when two events are mutually exclusive, the following probability relationships must hold:
      • P(AB) = 0
        • Stating that probability of both the events occurring together will be zero.
      • P(A or B) = P(A) + P(B)
        • In general the formula is P(A or B) = P(A) + P(B) - P(AB), as there is a possibility of event A and B occurring together, but in mutually exclusive scenario this is not possible.
        • Mutually exclusive example:
          • P(Queen or King)
            • P(Queen) = 4/52
            • P(King) = 4/52
          • Here the P(Queen and King) occurring together the chances are 0.
          • Hence the formula for P(A or B) = P(A) + P(B) is good enough.
        • Not mutually exclusive example:
          • P(Spade or Queen) : Here there is a possibility of both occurring together hence need to exclude this while calculating to avoid double counting.
    • Note: Independent event cannot be mutually exclusive. For independent events, the probability of two events is product of both, but in mutually exclusive the product of two events are always 0.
  • Collective/Cumulative Exhaustive Events
    • Events that cumulatively describes all the outcomes.
  • Conditionally Independent Events
    • If two events[P(A), P(B)] are individually independent or not, but when these two events are independent given the occurrence of third event.
    • Example:
      • Take a random sample of school children and for each child obtain data on:
        • Foot Size (F)
        • Literacy Score (L)
      • The two will be (positively) correlated, in that the bigger the foot size the higher the literacy score.
      • The random variables F and L are not independent.
      • Obviously a bigger foot size is not the direct cause for a higher literacy score. What correlates the two is the child's age (A), which is the confounder.
      • By conditioning on age, we no longer consider the relationship between foot size and literacy for the whole sample, but per each age group separately, and makes foot size and literacy score independent.
      • So this was just an example of two random variables F and L that were:
        • dependent when not conditioned on A
        • independent when conditioned on A
      • Hence F is conditionally independent of L given A: P(L|F, A) = P (L| A) And so: P(L|F) = P(L)
    • Reference: https://math.stackexchange.com/questions/23093/could-someone-explain-conditional-independence#:~:text=If%20two%20people%20live%20in,an%20effect%20on%20the%20other.
  • Discrete Probability Function
    • A discrete probability distribution counts occurrences that have countable or finite outcomes.
    • Example: Consider you have 1 - green ball, 2 - yellow balls, 3 - blue balls and 4 - red balls, total 10 balls. When randomly we pick a ball blindfolded, and the finite possible outcomes are green, yellow, blue and red. Probability function for this is x/10, such as probability of green ball occurring is 1/10 = 10%, yellow ball 2/10 = 20%, blue ball 3/10 = 30% and red ball 4/10 = 40% and the total of these together is 100%.
  • Continuous Probability Function
    • A continuous probability distribution counts occurrences that have any value or infinite outcomes.
    • Probabilities are measured over intervals, not single points.
    • Example: Blood pressure levels
  • Unconditional Probability
    • An unconditional probability is the likelihood that one out of several possible results will occur. 
    • Unconditional is also sometime called as marginal.
    • Here possibility that an event will occur regardless of whether other events have occurred or other conditions exist. Ideally does not depend of the prior event.
    • Example: P(Queen) is the probability of an event Queen returning
  • Conditional Probability
    • The probability of an event occurring given the outcome of another event.
    • This can also be used to incorporate additional information (which is given event) to update unconditional probabilities. Here the given additional information could gain better prediction. As conditional probability operates as if B is the event space and A is an event inside this new restricted space.
    • P(A|B) defines as probability of event A occurring given B event has occurred.
    • Example: P(Queen|Spades) is the probability of card returning Queen given the cards are Spades. 
    • General Formula: P(A|B) = P(AB) / P(B)
    • In simple terms could also be said as P(Conditional) = P(Joint)/P(Unconditional)
    • Example: 
      • P(Queen | Spades) = P(Queen and Spades) / P(Spades)
      • P(Queen | Spades) = (1/52) / (13/52)
      • P(Queen | Spades) = (1/13) #Reality check, this is true as Queen occurs only once in Spades
    • Formula For Independent : P(A|B) = P(A)
    • As there is no dependence of A giving B occurred.
  • Joint Probability
    • The probability of two events occurring together.
    • P(AB) is referred as both A and B event occurring at the same time.
    • The joint probability of two independent events is the product of the probability of each. 
      • Formula: P(AB) = P(A) * P(B)
      • Example:
        • P(Queen and Spades)
          • = P(Queen) * P(Spades)
          • = (4/52) * (13/52)
          • = (1/13) * (1/4)
          • = (1/52) #Reality check, this is true as there is only chance of this to occur in whole set of decks Queen and Spade occurring together is only once.
    • Conditional probability can be used to calculate joint probability, deriving from the above conditional formula.
      • Formula: P(AB) = P(A|B) * P(B)
      • Example: 
        • P(Queen and Spades) 
          • = P(Queen | Spades) * P(Spades)
          • = (1/13) * (13/52)
          • = (1/13) * (1/4)
          • = (1/52)
    • Joint probability depends on the two events acting independently from one another. To determine whether they are truly independent, it's important to establish whether one's outcome affects the other. If they do, they are dependent, which means they lead to conditional probability. If they don't, you end up with joint probability.
  • Law of Total Probability
    • In cases where the probability of occurrence of one event depends on the occurrence of other events, we use the law of total probability theorem.
Credits: https://www.embibe.com/exams/theorems-on-probability/
    • Total probability theorem indicates the probability that an event will occur given each of the sample space’s partitions
    • Total Probability Theorem Formula:
      • P(A)=P(A∩B)+P(A∩B^)
  • Bayes Theorem
    • Bayes' Theorem was named after 18th-century mathematician Thomas Bayes.
    • Bayes formula is used determine conditional probability, based on the previous outcome of the similar circumstances.
    • In other words Bayes formula uses basic ideas from probability to develop a precise expression for how information about one event can be used to learn from another event.
    • Formula:
      • P(A|B) = P(AB) / P(B)
        • Here,
        • P(AB) =  P(B|A) * P(A)
      • P(AB) is the probability of event A and B occurring together
      • P(A ∣ B) is the conditional probability of event A occurring, given that B is true.
      • P(B ∣ A) is the conditional probability of event B occurring, given that A is true.
      • P(A) and P(B) are the probabilities of A and B occurring independently of one another.
    • The Bayes theorem is based on finding P(A | B) when P(B | A) is given. 
    • While conditional probability deals with the probability of an event given another event, Bayes' Theorem is a more specific application that involves updating probabilities based on new evidence. Bayes' Theorem is a powerful tool in Bayesian statistics and is widely used in fields such as machine learning, medical diagnosis, and other areas where probabilistic reasoning is essential.

Recap with an Example

1000 mobiles are given to a company for testing, with following features:
  • 600 smart foldable phones
  • Of the 600 smart foldable phones 150 are AI enabled
  • 400 smart flip phones
  • Of the 400 smart flip phones 200 are AI enabled
Unconditional Probabilities
  • P(Smart Foldable Phones) = 600/1000 = 60%
  • P(Smart Flip Phones) = 400/1000 = 40%
Conditional Probabilities
  • P(AI Enabled | Smart Foldable Phones) = 150/600 = 25%
  • P(AI Enabled | Smart Flip Phones) = 200/400 = 50%
Joint Probabilities [Formula: P(AB) = P(A|B) * P(B)]
  • P(AI Enabled and Smart Foldable Phones) 
    • (150/600) * (600/1000) = 25%*60% = 15%
    • 15% of 1000 phones = 150 of smart foldable phones are AI enabled
  • P(AI Enabled and Smart Flip Phones)
    • (200/400) * (400/1000) = 50%*40% = 20%
    • 20% of 1000 phones = 200 of smart flip phones are AI enabled
Total Probability
  • P(AI Enabled Phones) = 
    • P(AI Enabled and Smart Foldable Phones) + P(AI Enabled and Smart Flip Phones)
    • 15% + 20% = 35%
    • 35% of 1000 phones = 350 phones are AI enabled
Bayes Rule [Formula: P(A|B) = P(AB) / P(B)]
  • P(Smart Foldable Phones | AI Enabled) / P(AI Enabled)
  • 15% / 35% = 42.9%
  • 350 AI enabled phones, and of them 150 are smart foldable phones : 150 / 350 = 42.9% 

Bayes Theorem Example

Suppose there are three types of managers: the underperformers beat the market only 25% of the time, the in-line performers beat the market 50% of the time, and the outperformers beat the market 75% of the time. Our prior belief is that a manager has a 60% probability of being an in-line performer, a 20% chance of being an underperformer, and a 20% chance of being an outperformer. If the manager beats the market two years in a row what should our updated beliefs be?

Our prior beliefs can be summarized as: P[MO] = 20%, P[MI] = 60% and P[MU] = 20%.

A Binomial Tree:

Associated Probability Matrix:

Above matrix is defined by deriving, for example MO beats market is 15% which is 75% of 20%.

The unconditional probability of observing the manager beat the market two years in a row, given our prior beliefs is P(2B) = 27.50%.
P(2B) = P(2B|MO) * P(MO) + P(2B|MI) * P(MI) + P(2B|MU) * P(MU)
Using Bayes’ theorem, for example, we can calculate our posterior belief that the manager is an outperformer as 40.91%:

P(MO|2B) = (P(2B|MO) * P(MO)) / (P2B)  = 0.56 × 0.20/0.275 = 40.91%

Similarly, we can show that the posterior probability that the manager is an in-line performer is 54.55% and that the posterior probability that the manager is an underperformer is 4.55%.

As we would expect, given that the manager beat the market two years in a row, the posterior probability that the manager is an outperformer has increased, from 20% to 40.91%, and the posterior probability that the manager is an underperformer has decreased, from 20% to 4.55%.

Although the probabilities changed, the sum of the probabilities remains equal to 100%.

Credits: Bionic Study Notes - This reference very much helps in understanding bayes theorem concept.

Description

  • Supervised Classification Model
  • Naive Bayes, calculates the probability of every X factor, then selects the highest probability
  • Naive Bayes, is based on the Bayes Theorem
  • Bayes Theorem :
    • P(A|B) = P(B|A)P(A) / P(B)
      • P(A|B) : Probability of A happening, given the condition of B has occurred
        • B is the evidence
        • A is the hypothesis
      • P(B) : Probability of B happening
      • P(A) : Probability of A happening
      • P(B|A) : Probability of B happening, given the condition of A has occurred
  • Naive Bayes algorithms are mostly used in 
    • sentiment analysis
    • spam filtering
    • recommendation systems 
  • Types of Naive Bayes Classifier :
    • Multinomial
      • Predictors are categorical, more than two
      • Usually used in document classification problem, 
        • eg : whether the document belongs to sports, health, technology, etc
    • Bernoulli
      • Predictors are boolean
        • eg: spam or ham email
    • Gaussian
      • Used when the features are normally distributed


Formula

  • Formula is derived from Bayes Theorem
    • P(Hypothesis | Evidence) = ( P(Evidence | Hypothesis) * P(Hypothesis) ) / P(Evidence)
      • Hypothesis = prediction variable
      • These are independent features
        • Evidence 1, Evidence 2, Evidence 3, ... Evidence n
    • P(x| Hypothesis)
      • Likelihood
      • Probability of the evidence, given the belief is true.
    • P(Hypothesis)
      • Class Prior Probability
      • The probability before the evidence is considered.
    • P(Evidence)
      • Predictor Prior Probability
      • Probability of the evidence, under any circumstance.
    • P(Hypothesis | Evidence)
      • Posterior Probability
      • Before the actual event occurred predicting given the x event has occurred, with the assumption event c could be related to the event x.
  • So if we have multiple parameters, the formula would be

  • Max probability is taken, after multiple prediction values


Assumption

  • Each independent variable should be not correlated
    • eg: if its high sunny day, humidity would also be high
  • Equal or Weighted importance given to each independent variable,
    • eg: to predict if it will rain or not, if we would give equal importance to sunny, humidity
    • eg: to predict if it will rain or not, if we would give more importance to cloudy than humidity
  • The word 'Naive' is used because of the above two assumptions


Algorithm Explained with an Example

Problem Statement:
From the given attributes, predict if the car was stolen or not, with the assumption the given attributes are independent with no correlation and equal importance.

Dataset:
Model Training:
First step is to calculate the Likelihood for each attribute against the target, by molding the frequency table to likelihood table.

Color:
Independent probability of each color: 
P(Red) = 5/10 = 1/2
P(Yellow) = 5/10 = 1/2
Frequency and Likelihood Table of the attribute: P(Color | Stolen)
Type:
Independent probability of each type: 
P(Sports) = 6/10 = 3/5
P(SUV) = 4/10 = 2/5
Frequency and Likelihood Table of the attribute: P(Type | Stolen)
Origin:
Independent probability of each origin: 
P(Domestic) = 5/10 = 1/2
P(Imported) = 5/10 = 1/2
Frequency and Likelihood Table of the attribute: P(Origin | Stolen)
Class Prior Probability:
In our example we have two possible outcomes for Stolen i.e. Yes or No.
P(Yes) = 5/10 = 1/2
P(No) = 5/10 = 1/2

Model Testing:
Lets test for the given input:

Using Formula:

Calculate the posterior probability P(Yes | X): P(Yes | Red, SUV, Domestic)
Given:
  • P(Yes | Red) : 3/5
  • P(Yes | SUV) : 1/5
  • P(Yes | Domestic) : 2/5
  • P(Yes) : 1/2
  • P(Red) : 1/2
  • P(SUV) : 2/5
  • P(Domestic) : 1/2
Apply:
P(Yes | Red, SUV, Domestic) 
= P(Yes | Red) * P(Yes | SUV) * P(Yes | Domestic) * P(Yes) / P(Red) * P(SUV) * P(Domestic)
= ((3/5) * (1/5) * (2/5) * (1/2)) / ((1/2) * (2/5) * (1/2))
= .24 = 24%

Calculate the posterior probability P(No | X): P(No | Red, SUV, Domestic)
Given:
  • P(No | Red) : 2/5
  • P(No | SUV) : 3/5
  • P(No | Domestic) : 3/5
  • P(No) : 1/2
  • P(Red) : 1/2
  • P(SUV) : 2/5
  • P(Domestic) : 1/2
Apply:
P(No | Red, SUV, Domestic) 
= P(No | Red) * P(No | SUV) * P(No | Domestic) * P(No) / P(Red) * P(SUV) * P(Domestic)
= ((2/5) * (3/5) * (3/5) * (1/2)) / ((1/2) * (2/5) * (1/2))
= .72 = 72%

Now lets predict, y = Max of [ P(Yes | X), P(No | X)]:
y = Max( Yes = .24 and No = .72)
y = No

Hence our example gets classified as ’NO’ the car is not stolen.

Credits: https://www.kdnuggets.com/2020/06/naive-bayes-algorithm-everything.html

Zero Frequency Problem

One of the disadvantage of the above explained problem is when there are 0 values in the frequency. For example lets say we have one more column added = 'Kilometers Run'. Here for the brand new car it would be 0. And this will get a zero when all the probabilities are multiplied. 

An approach to overcome this 'zero-frequency problem' in a Bayesian environment is to add one to the count for every attribute value-class combination when an attribute value doesn’t occur with every class value. Now the 'Kilometers Run' for the brand new car will be one, including the cars which already have values in it (for example if the car has already ran 200kms will now become 201kms) and this will solve the problem of 'zero frequency problem'.

Evaluation

  • ROC
    • ROC stands for Receiver Operating Characteristics Curve
    • Used to find performance matrix of classification model, under various Threshold parameters
    • This curve is plot using two parameters:
      • True Positive Rate
        • TPR = TP / (TP + FN)
        • Intuitively, Percentage of identifying actual true values
      • False Positive Rate
        • FPR = FP / (FP + TN)
        • Intuitively, Percentage of identifying actual false values
  • AUC
    • AUC stands for Area Under ROC Curve
    • Also written as AUROC (Area Under the Receiver Operating Characteristics)
    • AUC ranges in value from 0 to 1. A model whose predictions are 100% wrong has an AUC of 0.0; one whose predictions are 100% correct has an AUC of 1.0.
    • If the value is 1, it means our model is overfitted and if the value is 0 it means it underfitted
    • But Higher the AUC value, better the performance
  • Accuracy :
    • Percentage of correct predictions
    • But in real life scenarios, we may be more keen in looking for precision and recall
    • Because eg: we have a scenario where in when need to find if the there is fraud transaction, as out of all the transactions we would have only 1% of fraud transactions, we may have good results in accuracy, but our intention would be to find less False Negatives. 
    • Accuracy of model: TP+TN / (TP+TN+FP+FN)
    • Misclassification
      • Percentage of incorrect predictions
      • Formala : 
        • 1 – Accuracy
        • or 
        • FP+FN / (TP+TN+FP+FN)
  • Classification Matrix
    • Precision
      • The precision is the ratio tp / (tp + fp) where tp is the number of true positives and fp the number of false positives. 
      • The precision is intuitively the ability, indentify the Type I error, find score based on how False Positive value 0-1, can be converted to percentages.
    • Recall
      • The recall is the ratio tp / (tp + fn) where tp is the number of true positives and fn the number of false negatives. 
      • The recall is intuitively the ability, identify the Type II error, find score based on how False Negative value 0-1, can be converted to percentages.
    • Specificity
      • The specificity is the ratio tn / (tn + fn) where tn is the number of true negatives and fn the number of false negatives. 
      • The recall is intuitively the ability, negative correct prediction, could also consider this as Type II error
    • F1 score
      • The harmonic mean of the precision and recall.
      • F1 score reaches its best value at 1 and worst score at 0.
  • Confusion Matrix
    • An easy and popular method of diagnosing model performance.
      • TN : True Negative
      • TP : True Positive
      • FN : False Negative - Type II 
      • FP : False Positive - Type I
    • Always Type I and Type II models are inversely proportions, intuitively if Type I increases Type II error decreases


Example

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns

%matplotlib inline

import warnings
warnings.filterwarnings('ignore')

from scipy import stats
from scipy.stats import zscore
from sklearn import metrics
from sklearn.model_selection import train_test_split

from sklearn.naive_bayes import GaussianNB

# Install a pip package in the current Jupyter kernel
import sys
!{sys.executable} -m pip install mlxtend

from mlxtend.plotting import plot_confusion_matrix

df = pd.read_csv('Bank_Personal_Loan_Modelling.csv', delimiter=',')

## obervations :
## no NAN values
## values are only either int or float
## total 5000 records
df.info()
'''
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 5000 entries, 0 to 4999
Data columns (total 14 columns):
ID 5000 non-null int64
Age 5000 non-null int64
Experience 5000 non-null int64
Income 5000 non-null int64
ZIP Code 5000 non-null int64
Family 5000 non-null int64
CCAvg 5000 non-null float64
Education 5000 non-null int64
Mortgage 5000 non-null int64
Personal Loan 5000 non-null int64
Securities Account 5000 non-null int64
CD Account 5000 non-null int64
Online 5000 non-null int64
CreditCard 5000 non-null int64
dtypes: float64(1), int64(13)
memory usage: 547.0 KB
'''


def calculate_score(model,X_train,y_train,X_test,y_test):

predicted_labels = model.predict(X_train)
accuracy_score = metrics.accuracy_score(y_train, predicted_labels)
print("Accuracy for Train Data: ")
print(accuracy_score)

predicted_labels = model.predict(X_test)
print("Actual Labels")
print(y_test.as_matrix(columns=None))
print("Predicted Labels")
print(predicted_labels)
accuracy_score = metrics.accuracy_score(y_test, predicted_labels)
print("Accuracy for Test Data: ")
print(accuracy_score)
print("Confusion Matrix")
cm = metrics.confusion_matrix(y_test,predicted_labels)
print(cm)

fig, ax = plot_confusion_matrix(conf_mat=cm,
show_absolute=True,
show_normed=True,
colorbar=True)
plt.show()
print("Classification Matrix")
print(metrics.classification_report(y_test,predicted_labels))
result={}
result['Accuracy']=accuracy_score
result['MSE']=1-accuracy_score
return result

# Naive bayes model
naive_model = GaussianNB()
naive_model.fit(x_train, y_train)

calculate_score(naive_model, x_train, y_train, x_test, y_test)
'''
Accuracy for Train Data:
0.8839004861309694
Actual Labels
[0 0 0 ... 0 0 0]
Predicted Labels
[1 0 0 ... 0 0 0]
Accuracy for Test Data:
0.8806666666666667
Confusion Matrix
[[1231 131]
[ 48 90]]

Classification Matrix
precision recall f1-score support

0 0.96 0.90 0.93 1362
1 0.41 0.65 0.50 138

micro avg 0.88 0.88 0.88 1500
macro avg 0.68 0.78 0.72 1500
weighted avg 0.91 0.88 0.89 1500

{'Accuracy': 0.8806666666666667, 'MSE': 0.11933333333333329}
'''


Advantages

  • They are fast and easy to implement.
  • Suited when the dimensionality of the inputs is high.


Drawbacks

  • Requirement of predictors to be independent. 
  • If categorical variable has a category (in test data set), which was not observed in training data set, then model will assign a 0 (zero) probability and will be unable to make a prediction. This is often known as “Zero Frequency”. To solve this, we can use the smoothing technique. One of the simplest smoothing techniques is called Laplace estimation. 


Credits and References

  • https://towardsdatascience.com/naive-bayes-classifier-81d512f50a7c
  • https://www.investopedia.com/
  • https://medium.com/machine-learning-101/chapter-1-supervised-learning-and-naive-bayes-classification-part-1-theory-8b9e361897d5
  • http://www.statsoft.com/textbook/naive-bayes-classifier
  • https://machinelearningmastery.com/naive-bayes-for-machine-learning/
  • https://www.analyticsvidhya.com/blog/2017/09/naive-bayes-explained/
  • https://insightimi.wordpress.com/2020/04/04/naive-bayes-classifier-from-scratch-with-hands-on-examples-in-r/
  • https://scikit-learn.org/stable/modules/naive_bayes.html (each naive bayes type model with example)
  • https://www.saedsayad.com/images/Bayes_rule.png
  • SchweserNotes Quantitative Analysis 2023
  • https://www.machinelearningplus.com/predictive-modeling/how-naive-bayes-algorithm-works-with-example-and-full-code/
  • https://www.machinelearningplus.com/predictive-modeling/how-naive-bayes-algorithm-works-with-example-and-full-code/
  • https://www.kdnuggets.com/2020/06/naive-bayes-algorithm-everything.html


Thursday, 11 May 2023

Enterprise Mobile Security


Introduction

Mobile has become our an additional finger which knows all most of sensitive information. Leading this makes it very much important to secure and especially for enterprise. The mobile device has different composition regular computer, embedded in nature makes it bit more challenging in securing.

Communicating Methods

Communication between from one mobile to another mobile is through electro magnetic waves (zero and one) but the mobile itself does not have the capability to transfer through long distance. Hence mobile devices can communicate using different methods such as Cellular Network - (5G which is the fifth generation of cellular wireless standards providing capabilities beyond 4G), Bluetooth, Wi-Fi, Near field communications and so on. There are various security concerns with this, i.e., Traffic monitoring, Location tracking, Wide access to mobile devices. Let's look into each of these methods.

Cellular Network

  • Cellular network or also known as mobile network, are the main mode of the communications, which connects to and from end nodes through the service provider network.  
  • Each phone communicates with the service provider by radio waves through a local antenna at a cellular base station or cell site. The cellular consists of below components:

Credits: https://www.electroschematics.com/wp-content/uploads/2010/03/Mobile-Communication.png

  • Cellular Layout
    • Mobile network is divided into different geographical areas known as cells. This geographical are is divided hexagonal cells - an antenna coverages a cell with certain frequencies. 
    • Each cell has a transceiver - mobile tower that make a wireless connection to the mobile device and the base station is a land station in the land mobile service below the tower. Both serves the similar purpose to produce network signals to the consumers.
    • Base Station
      • Base stations provide the cell with the network coverage and connects with the tower, which can be used for transmission of voice, data, and other types of content.
      • The base stations are meant to improve the signal frequency and communication between interconnected devices such as computers or smartphones.
    • Tower
      • Tower is where the antennas and electric communications equipments are placed to create a cell or adjacent cells.
      • The cell towers distributes the signals are generated by the base station. 
    • All base stations and towers in a city are connected via a high-speed link or fibre optics to a mobile telephone switching office (MTSO). 
  • Mobile Tower Switching Office
    • Our mobile has not enough signal powers to directly call a caller residing in another city hence it sends signals to a mobile tower. 
    • The mobile tower then sends signals to MTSO. MTSO check our sim data in its database to find the cell in which the phone is present  and send a signal to another city MTSO. Then MTSO sends signals to mobile through the mobile tower.
    • MTSO is normally located in the central cell of a cluster and is generally connected to the Public Switched Telephone Network (PSTN).
  • Public Switched Telephone Network - PSTN
    • A public switched telephone network is a combination of telephone networks used worldwide, including telephone lines, fiber optic cables, switching centers, cellular networks, satellites and cable systems. 
    • PSTN is a century-old worldwide connected telephone network and lets users make landline telephone calls.
  • Common security concerns
    • Traffic monitoring
    • Location tracking
    • DDOS attacks
    • Access fraud
    • Stolen phones
    • Subscription fraud
  • Security measures
    • Network Traffic Monitor and Analysis
    • Encrypted communication
    • Velocity Checking
    • A subscriber usage-pattern database 
    • Customer call analysis 
    • Geographic dispersion checking 
    • Extensive proprietary antifraud algorithms 
  • Below site explains in detail about different fraud and solutions in cellular network
    • https://csrc.nist.gov/csrc/media/publications/conference-paper/1997/10/10/proceedings-of-the-20th-nissc-1997/documents/031.pdf

Wi-Fi 

  • Wi-Fi is the family of wireless network protocols, commonly used for local network access. Wi-Fi stands for Wireless Fidelity.
  • Wi-Fi uses radio frequency.
  • An internet connection shared with multiple devices within certain range via wifi router. This router is connected directly to the internet modem and acts as a hub to broadcast the internet signal to all your Wi-Fi enabled devices.
  • Wifi repeaters, is used to extend the length of the existing network.
  • Common security concerns
    • Data capture: Need to Encrypt data
    • On-path attack: Always Monitor data
    • Denial of service: Monitor unwanted traffic which is calling the frequency interference
  • Security measures
    • Wired Equivalent Privacy (WEP) encryption was designed to protect against casual snooping but it is no longer considered secure. Later WPS - Wi-Fi Protected Setup, WPA - Wi-Fi Protected Access and WAP2 were introduced, which is also not secure now.
    • In 2018, WPA3 was announced as a replacement for WPA2, increasing security and enabling secure authentication on the wireless router. 
    • WPA3-Personal
      • WPA3 with shared key where everyone uses same key
    • WPA3-PSK
      • Using WPA3 session key is derived from PSK using SAE (Simultaneous Authentication of Equals) to provide stronger defences against password guessing.
      • This allows the access provider and station peers to authenticate each other as part of the handshake process while using cryptographic tools to prevent an attacker from performing an offline password cracking scheme.
  • As like Cellular networks, W-Fi internet access has become much more embedded in society. 


Li-Fi

  • Li-Fi stands for Light Fidelity.
  • LiFi technology will allow us to connect to the internet using light from lamps, streetlights or LED televisions. 
  • Li-Fi uses infrared light or via LED to transmit data.
  • In addition to being cheaper, safer and faster than wifi, it does not need a router and just requires to point your mobile or tablet towards a light bulb to surf the web.
  • Common security concerns
    • Jamming
    • Spoofing
    • Data modification
    • Inability to work without light
  • Security measures
    • LiFi technology is widely considered to be generally more secure than WiFi. 
    • Plenty of security features can be embedded in LiFi systems in order to make them more secure as light cannot pass through walls like radio waves and it carries more volume of data at a time.
    • Encryption
    • Monitoring


Bluetooth

  • Bluetooth wireless technology is a short range communications technology intended to replace the cables connecting portable unit and maintaining high levels of security. 
  • Bluetooth uses a spread-spectrum, frequency-hopping, full-duplex signal.
  • An antenna-equipped chip in each device that wants to communicate sends and receives signals at a specific frequency range defined for short-range communication.
  • For Bluetooth devices to communicate, they pair with each other to form a personal area network (PAN), also known as a piconet.
  • This process is done through discovery, with one device making itself discoverable by the other device. 
  • Bluetooth is a common mobile connectivity method because it has low power consumption requirements and a short-range signal. 
  • Point to Point
    • One-to-one connection, conversation between two devices
  • Point to Multi Point
    • One of the most popular communication methods 802.11 wireless.
    • Multipoint doesn’t necessarily mean that you can stream media from two devices at a time.
  • Common security concerns
    • Bluejacking
      • Sending of unsolicited messages to another device via Bluetooth
    • Bluesnarfing
      • Access a bluetooth enabled device to access data from the device.
      • Security has been patched in the latest devices
  • Security measures
    • Update with latest patches
    • Monitoring
    • Encryption


RFID

  • Radio Frequency Identification (RFID) refers to a wireless system comprised of two components: tags and readers. The reader is a device that has one or more antennas that emit radio waves and receive signals back from the RFID tag.
  • Uses radar technology
    • Radio energy is transmitted
    • Bidirectional communication
  • There are two types of RFID tags: active and passive tags. An active tag can broadcast a signal over a larger distance because it contains a power source. A passive tag, on the other hand, isn’t powered but is activated by a signal sent from the reader.
  • Common security concerns
    • Data capture
    • Spook reader
    • Signal jamming
    • Decrypt communication
  • Security measures
    • Cryptography is primary
    • Blocker tags, prevent unauthorized readers
  • Commonly used as geofencing security measure, radio-frequency identification (RFID) to define a geographic perimeter, when the device enters of exits alerts are sent.


NFC

  • Near-field communication (NFC) is a set of standards for contactless communication between devices. NFC chips in mobile devices generate electromagnetic fields. This allows a device to communicate with other devices. 
  • NFC is extension of RFID technology.
  • Near-field communication transmits data through electromagnetic radio fields to enable two devices to communicate with each other. To work, both devices must contain NFC chips, as transactions take place within a very short distance. 
  • NFC-enabled devices must be either physically touching or within a few centimetres of each other for data transfer to occur.
  • NFC began in the payment-card industry and is evolving to include applications in numerous industries worldwide.
  • The NFC standard has three modes of operation: 
    • Peer-to-peer mode: Information shared exchanged between two mobile devices directly.
    • Read/write mode: An active device receives data from a passive device. 
    • Card emulation: The device is used as a contactless credit card. It emulates a payment card or other physical card in card readers, magnetic-stripe readers, and contactless card readers used to make payments directly from your mobile device.
  • Common security concerns
    • Remote capture
    • Frequency jamming
    • On path attack
    • Lost of NFC device control - digital pickpocketing
  • Security measures
    • Encryption
    • Always patch up-to-date
    • Turn-off when not in use 


Mobile Networks

Mobile networking has evolved significantly since the introduction of the first-generation (1G) mobile network in the 1980s. G refers to Generation. Each Generation is defined as a set of telephone network standards, which details the technological implementation of a particular mobile phone system.
Credits: https://www.techindulge.com/technovation/1g-2g-3g-4g-and-5g-wireless-phone-technology-explained-meaning-and-differences/

1G

  • 1G is the first generation of wireless cellular technology. 
  • 1G supports voice only calls.  
  • 1G is analog technology
  • The maximum speed of 1G is 2.4 Kbps.

2G

  • Changed from analogue (1G) to Digital (2G).
  • Ability to send SMS (Short Message Service) and plain text-based messages.
  • GSM and CDMA was introduced during this period.
  • The maximum speed of 2G with General Packet Radio Service (GPRS) is 50 Kbps. The max theoretical speed is 384 Kbps with Enhanced Data Rates for GSM Evolution (EDGE).
  • Before making the major leap from 2G to 3G wireless networks, the lesser-known 2.5G and 2.75G were interim standards that bridged the gap to make data transmission.

3G

  • Enabled web browsing, email, video downloading, picture sharing and so on.
  • The maximum speed of 3G was around 2 Mbps for non-moving devices and 384 Kbps in moving vehicles. 

4G

  • Applications include amended mobile web access, IP telephony, gaming services, high-definition mobile TV, video conferencing, 3D television, and cloud computing. 
  • The max speed of a 4G network when the device is moving is 100 Mbps. The speed is 1 Gbps for low-mobility communication such as when the caller is stationary or walking.

5G

  • 5G promises significantly faster data rates, higher connection density, much lower latency, and energy savings, among other improvements.
  • 5G offers data transfer rates of up to 20 Gbps, it allows users to download ultra-high-definition videos and access the internet at lightning-fast speeds. 
  • Also offers lower latency, better network coverage, and improved call quality.

Wireless Carriers

Wireless carrier means the cellular technology company that provides mobile telecommunication services for a Supported Device.

CMDA

  • CDMA stands for Code Division Multiple Access. 
  • It is handset-specific.
  • CDMA is not very common, and it is available in comparatively fewer carriers and countries. These devices are exclusive to Canada, Japan, and the United States.
  • The CDMA technology does not support any such feature. It cannot transmit voice and data simultaneously.
  • CDMA is faster, provides better security and has comes with built in encryption.

GSM

  • GSM (Global System for Mobile) standard in Finland was launched by AT&T. Every device uses SIM (Subscribers Identity Module) to communicate with the provider network.
  • GSM is highly available and globally used. Over 80% of the entire world’s mobile networks use it.
  • It uses the Time division multiple access (TDMA) and Frequency division multiple access (FDMA).
  • GSM supports the transmission of both voice and data at once.
  • GSM is slower, less secure and no default encryption compared to CDMA.

LTE

  • Long-Term Evolution (LTE) is used for faster data transfer and higher capacity. Different variations of LTE networks exist across carriers that use different frequencies. For example Sprint, T- Mobile, Verizon, and AT&T all have their own bands of LTE.
  • LTE moves large packets of data to an internet protocol system (IPS). Old ways of moving data used Code-division multiple access (CDMA) and the Global System for Mobile Communications (GSM), and those methods moved only small amounts of data.
  • LTE and 4G simply evolved together, with LTE is industry standard that describes the particular type of the forward edge of the fourth generation’s advancement.
  • 4G LTE functionality has two key preconditions: a network that supports the ITU-R (ITU Radiocommunication Sector) standard speeds and a device powerful enough to match and handle the speeds of that network. 
  • GSM and CDMA all switched to LTE as global 4G standard. As CDMA and GSM, are inefficient uses of the airwaves.

SATCOM

  • For users who lack traditional landline or cellular coverage, satellite communication (SATCOM) is an option for mobile device use. 
  • SATCOM uses an artificial satellite for telecommunication that transmits radio signals. 
  • It can cover far more distance and wider areas than most other radio technologies. 
  • Because satellite phones do not rely on phone transmission lines or cellular towers, they function in remote locations. 
  • Most satellite phones have limited connectivity to the Internet, and data rates tend to be slow, but they come with GPS capabilities to indicate the user’s position in real time. 


Mobile Management

  • Mobile Device Management
    • Managing mobile device access and usage in an organization is a security challenge to achieve. 
    • Manage the company owned or user owned mobile devices centrally. 
    • Set policies on apps, camera, data, access and so on.
    • Access control such as force screen locks, multi factor authentication, remote wipe, geofencing will be part of MDM.
  • Mobile Content Management
    • Secure the content present in the mobile device is role of Mobile Content Management - MCM. 
    • Monitoring and restriction on the file sharing, online content viewing and uploading.
    • Centrally manage the data in cloud using solutions such as Microsoft Office 365.
    • Any data which sent or receives from mobile devices should go through DLP - Data Loss Prevention - preventing any sensitive information leakage.
    • Data on the device needs to be encrypted.
    • Restrict and block external or removable drives.
  • Mobile Application Management
    • Not all applications are secure some are malicious, managing mobile apps is quite tough.  
    • Any new application installed should be managed through Mobile Application Management - MAM and only allowed apps could be installed. 
    • Not all the applications dangerous but still are not required for the business, such as games and social media apps - these applications would be denied for installation.
  • Unified Endpoint Management - UEM
    • Evolution of MDM, manages mobile and non mobiles.
    • End users can use different types of devices, and it could be blended together.
    • Applications can be used across different platforms.


Mobile Protection Measures

  • Remote wipe
    • Managed by MDM, removes all the data from the device whenever required usually during theft.
    • Make sure essential data is properly backed up.
  • Geolocation
    • Location tracking system
    • Used during commute, cross border alerts, find phone
  • Geofencing
    • Restrict mobile feature or features when the device is present in particular location.
    • Authenticate and allow login when the device is located in particular area.
  • Screen lock
    • Locking the mobile using PIN, Passcode, Pattern and Biometerics.
    • Auto lock after configured time.
    • Erase data after few invalid entries.
  • Push notification services
    • Information popup service on the screen.
    • Receives notifications even when the phone is idle or using different application.
  • Passwords and Pins and Biometrics
    • Mobile devices can have multi authentication based on the apps used.
    • Password rotation, reset, complexity policy handled through MDM.
  • Context aware authentication
    • Switch to multi level authentication during abnormalities observed or different pattern followed.
    • Example, access through different location, connected to different wifi, paired with bluetooth.
  • Containerization
    • Segment the storage for business use, to avoid data leak.
    • Easy to manage offboarding where the corporate data is deleted retaining the personal data.
  • Full device encryption
    • Encryption ensures even after the theft data is lost but still secure.
    • Managed and keys are rotated through MDM.
  • MicroSD Hardware Security Module - HSM
    • Provides security services - encryption, key generation, key rotation, digital signatures, store keys securely and encrypted.
  • SEAndriod
    • Security Enhancements for Andriod, built using SELinux (Security Enabled Linux).
    • Centralized policy management. 
  • Always have the device patched with the latest security fixes.


Mobile Deployment Models

Enterprise mobile device deployment models would fall in any of the below agreements:

  • BYOD is Bring Your Own Device
    • Employee owns the device
    • Difficult to secure, could be managed by MDM using containterization.
  • COPE is Company Owned/Personally Enabled
    • Company buys the device, used for both professional and personal.
    • Organization has full control on the device, including monitoring.
  • CYOD is Choose Your Own Device
    • Similar to COPE, corporate own device but user's choice of mobile device.
    • Gets tricky in terms of security for certain models.
  • COBO is Company Owned/Business Only
    • Company buys the device and used only officially with restriction.
  • VMI is Virtual Mobile Infrastructure
    • Apps and Data are separated from the mobile device.
    • Data is stored securely centralized, risk is minimized.


Conclusion
Mobile security along with other general security measures will help the enterprise and the individual to be secure.


Credits and References

  • https://www.uscybersecurity.net/wp-content/uploads/2019/02/Mobile-Security.jpg
  • https://whatsag.com/mobile-technology/the-difference-between-a-cell-tower-and-a-base-station.php
  • https://www.youtube.com/watch?v=1JZG9x_VOwA
  • https://www.baeldung.com/cs/mobile-networking-generations
  • https://www.lifewire.com/1g-vs-2g-vs-2-5g-vs-3g-vs-4g-578681
  • https://www.weboost.com/blog/what-is-4g-lte-and-how-does-it-work#:~:text=LTE%20moves%20large%20packets%20of,it%20helps%20streamline%20your%20service.
  • https://www.uctel.co.uk/blog/4g-vs-lte-understanding-the-difference-between-4g-and-lte#:~:text=So%20what%27s%20the%20difference%20between,compared%20to%20the%20fourth%20generation.
  • https://csrc.nist.gov/csrc/media/publications/conference-paper/1997/10/10/proceedings-of-the-20th-nissc-1997/documents/031.pdf
  • https://www.investopedia.com/terms/n/near-field-communication-nfc.asp
  • https://www.youtube.com/watch?v=UOGZbq4t_g8

Scarcity Brings Efficiency: Python RAM Optimization

  In today’s world, with the abundance of RAM available, we rarely think about optimizing our code. But sooner or later, we hit the limits a...