Thursday, 18 August 2022

Role of Data in Statistics

 


Statistics

Statistics as per wikipedia, is the discipline that concerns the collection, organization, analysis, interpretation, and presentation of data.


Statistics Types

  • Descriptive : What ? Eg: Mean, Mode, Etc
  • Diagnostics : Why ? Is all called as Inferential. Eg: Hypothesis, Conjectures 
  • Predictive : What will happen ? Eg: Regression Analysis
  • Prescriptive : What should we do to make this happen ? Eg: Decision Making

Statistics Terms 

  • Population
    • The entire universe of the data being studied. Though technical very challenging to acquire all the data for study.
    • Eg: Entire nation's people vote opinions
  • Parameter:
    • Characteristics of a population
    • Greek letters are used to represent
    • Eg: for Average μ
  • Sample
    • Generally part or subset of the population
  • Statistics
    • Characteristics of sample
    • Here the english letters are used for eg: x̄ for sample mean, though sometimes can be used to greek letters having ^ over it eg: ^μ
    • Eg: Avg people of each state go for voting from the sample data
  • Variable
    • Question that is being research (ideally sample)
    • Eg: Avg people from Assam go for voting
  • Data
    • Actual value the variable takes
    • Collection of observations
    • Helps us to find the relationships between two or more events
    • Help us predict future
    • Inputs are called as features
    • Outputs are called as labels or response
  • Let see with an example for the story so far.
    • Conducted survey over 10000 citizens to know their voting options
      • Population: Entire nation's citizens
      • Sample: 10000 citizens
      • Statistics : Avg of citizens going to vote from the 10000 citizens
      • Variable : Will vote or not
      • Data: Actually they have voted or not

Data Types

  • Independent 
    • Usually we can have this in X axis
  • Dependent or also called as Target Data/Label Data/Class Variable
    • Usually we can have this in Y axis
  • Quantitative
    • Numeric :
      • Discrete:
        • Which is countable
        • Output or labels, are defined
        • Eg: number of children
      • Continuous:
        • Keeps data getting finer and finer, infinite
        • Which is measurable
        • Output or labels, are not defined
        • Eg: price of home, stock price, age, time, height, weight, salary
        • Types :
          • Ratio
            • Proportions can be taken by comparing
            • Zero is absolute
            • Eg: Salary as 0, means no salary
            • Note: so we don't need to worry on data transformation
            • values have scale and a true zero point
            • eg: level of measurement describes the time needed to complete a project
          • Interval
            • Proportions cannot be taken
            • Zero is not absolute
            • Eg: Temperature : 0 Degrees, where it means here freezing point, but there are values below this also, we can transform this to Kelvin
            • Eg: Time, 00:00 means, not no time, its 12:00am
            • Eg: Longitude and Latitude, Equator
            • Note: Data can be transform here
    • Categorical :
      • Eg: types of car, colour of cage, gender
      • Types:
        • Binary
          • Symmetric : equally important, example Female or Male in interview equal importance irrespective of imbalance of data.
          • Asymmetric : Higher weightage to either of the value, 1 or 0, not equally. Important. Eg: Positive or Negative for disease test, so we can give more weightage Positive results
          • Symmetric: equal importance
            • Female
            • Male
          • Asymmetric: significance given
            • Positive
            • Negative
        • One hot encoding
          • Eg: name, gender
          • Example :
        • Nominal
          • Predetermined categories
          • Cannot if sorted / put data in particular order
          • Eg:
            • Red
            • Blue
            • Green
          • The data further could also be represented as below:

          • Ordinal
            • Can be sorted / put data in particular order but space between the data do not matter
              • Eg: Top 3 runner in a marathon
              • 1 - 15hrs
              • 2 - 17hrs
              • 3 - 25hrs
              • Here the different from 15 - 17 which is 2hrs and 17 - 25 which is 8hrs this is known as space and it doesn't impacts
            • Lacks scale
            • Label encoding (for conversion)
              • Eg:
                • Good
                • Better
                • Best
              • Eg: star rating, height

          • Data
            • Can put data in order and space between data is important
            • Zero does not have any meaning
            • Eg: speed of bike
              • 20km/hr
              • 70km/hr
              • 120km/hr
              • The difference between 20 - 70 which is 50 and 70 - 120 which is 50 again has importance
          • Proportion data
            • Can put data in order and space between data is important
            • Zero does have any meaning
            • Eg: No of bikes each student has
              • 5 or 0 each has value

    • Qualitative
      • Text
      • Images
      • Audio
      • Video

    Data Summary

    • Frequency Table
      • No of times a value occurs in data
      • Types
        • Relative Frequency
        • Cumulative Relative Frequency
    • Probability
      • Chances of a value occurring in an experiment
    • Example: Students marks

    Data

    Frequency

    Relative Frequency (rf)

    Cumulative Relative Frequency (crf)

    90

    300

    60

    60

    95

    100

    20

    80

    100

    100

    20

    100

    • Sample Summary from the data
      • % of students who obtained >= 90 marks in exam
        • 20% + 20% = 40%
      • % of students who scored 90marks or less 
        • 60%
      • % of students who scored more than 90% and less than 100%
        • 20%
      • Average Marks
        • Avg = ∑x * frequency / ∑frequency (total number)
        • Eg:
          • 90*300 = 27000
          • 95*100 = 9500
          • 100*100 = 10000
          • ∑x * frequency =  46500
          • ∑frequency = 500
          • avg = 46500/500 = 93

    Data Visualization

    • Easy to see the data in pictorial way
    • It is very important

    Credits: https://www.researchgate.net/publication/349139148/figure/fig6/AS:1005489378840612@1616738757855/Heat-map-depicting-the-distribution-of-the-population-density-in-the-250-m-250-m.png

    Credits: https://neilpatel.com/wp-content/uploads/2021/03/Data-Visualization_Featured-Image-1.png

    Credits: Anonymous


    Common Symbols and Syntax

    • x power of 5 : x*x*x*x*x; eg : 2 power 3 = 2*2*2 = 8
    • x^5: same as x power of 5
    • x power of -2 : 1 / (x*x); eg: 2 power -3 = 1/(2*2*2) = 1/8 = 0.125
    • x power of (1/n) : 1/2 is actual square root : 
      • 8 power of (1/3) = 2
      • cube root of 8 = 2
    • x! : x * (x-1) * (x-2) * ... 1 : eg: 3! = 3*2*1 = 6
    • Ʃ power n from x=1 of x : x= 1+2+3+..+n
    • Ʃ power n from x=1 of ( x sub i ) : ( x sub i ) = x1 + x2 + ..+xn
      • eg: x = {5,2,1}, n=3, 
      • Ʃ power 3 from x=1 = 5*1+2*1+1*1 = 8

    Measures of Centre
    • Arithmetic Mean : 
      • The arithmetic mean, often simply called the average, is a measure used to find the central tendency of a set of numbers. It's calculated by adding up all the numbers in a dataset and then dividing the sum by the total count of numbers.
      • ẋ: ( Ʃ power n from x=1 of ( x sub i ) ) / n
      • Eg: x = {5,2,1}, n=3, 
        • ( Ʃ power 3 from x=1 of ( x sub i ) ) / 3 = (5+2+1)/3 = 8/3 
      • Mean can be influenced by the outliers
        • Below are the ways to handle outliers
          • Do nothing
          • Trim outlier
          • Wensorized
            • winsorized mean is an averaging method that involves replacing the smallest and largest values of a data set with the observations closest to them, it mitigates the effects of outliers by replacing them with less extreme values - here it either provides for floor or ceiling values
      • Formula:
        • µ: mean of population
        • ẋ: mean of sample
    • Geometric Mean
      • The geometric mean is another measure of central tendency, but it's specifically used when dealing with growth rates or ratios, such as investment returns over multiple periods. It's calculated by taking the nth root of the product of n numbers, where n is the total count of values.
      • Arithmetic mean does not account for the compounding effect of returns over time, especially in the presence of negative returns. The geometric mean, on the other hand, accounts for this compounding effect and provides a more accurate measure of the average growth rate.
      • Geometric Mean = (x₁ * x₂ * x₃ * ... * xₙ)^(1/n)
        • If any one number is with -ve the entire data will be -ve, from the formula its power of ^ 1/n, which will error if we do under root of -ve number.
        • Hence all these numbers has to be more than 0 i.e. >0.
        • In that case +1 is hence added to all the values
        • And finally before root, do -1 to the overall values.
      • Eg:
        • We have invested 100 dollars after 1st year it returned 200 dollars and 2nd year it returned 100 dollars
        • Here in the 1st year the return is 100% and second year the returns are -50%.
        • Lets try to find average return:
          • First lets try with Arithmetic Mean = 1 + (-0.5%) / 2 = 0.25 = 25%
            • Note: if you see we haven't got 25% return, but there the AM seems to have this stating it wrongly.
          • Now using Geometric Mean (here percentage is converted)
            • (((1+1) * (1-0.5)) - 1) ^ 1/2
            • ((2 * 0.5) - 1) ^ 1/2
            • (1 - 1) ^ 1/2
            • 0
          • Here if you see the actual returns are 0 and the GM average stats its correctly.
    • Weighted Mean
      • Arithmetic mean is a special weighted mean as here all weights are equal.
      • For example:
        • (1+2+3+4+5) / 5 = 3
        • (1*1/5)+(2*1/5)+(3*1/5)+(4*1/5)+(5*1/5) = 0.2+0.4+0.6+0.8+1 = 3
      • In weighted mean the 1/n (eg: 1/5) is different for each value, ideally used in portfolio management.
      • Let's say you have three exams with different weights:  
        • Input
          • Exam 1: 20% weight, score of 80 
          • Exam 2: 30% weight, score of 75 
          • Exam 3: 50% weight, score of 90 
        • To calculate the weighted mean, you first multiply each score by its respective weight, then sum up these products, and finally divide by the sum of the weights.  
          • Multiply each score by its weight: 
            • Exam 1: 0.20 * 80 = 16 
            • Exam 2: 0.30 * 75 = 22.5 
            • Exam 3: 0.50 * 90 = 45 
          • Sum up these products: 16 + 22.5 + 45 = 83.5  
          • Divide by the sum of the weights: 
            • Total Weight = 0.20 + 0.30 + 0.50 = 1  
            • Weighted Mean = 83.5 / 1 = 83.5  
        • So, the weighted mean score is 83.5. This means that when taking into account the weights assigned to each exam, the overall average score is 83.5.
    • Harmonic Mean
      • The harmonic mean is a type of average used to determine the central tendency of a set of numbers. 
      • It's calculated by dividing the number of values in the dataset by the sum of the reciprocals of the values. 
      • It is the inverse of the arithmetic mean of the reciprocals.
      • If you have  n values  x1, x2 , . . . , xn, 
        • harmonic mean (H) is:  n / ∑xi
      • One common application of the harmonic mean is in situations where rates are involved, such as calculating average speeds or average rates of return.
      • It's particularly useful when dealing with ratios or rates, where outliers can heavily skew the arithmetic or geometric means.
      • Example: 1, 2, 3, 4, 5, 25
        • 6 / (1/1 + 1/2 + 1/3 + 1/4 ​ + 1/5 ​ + 1/25)
        • 6 / 2.32
        • 2.58
      • Compilation of all the means:
        • AM: 1+2+3+4+5+25 = 6.67
        • HM: 2.58
        • GM: 3.80
        • AM * HM = 6.67 * 2.58 = 17.2086
        • GM^2 = 14.44
      • Another analogy, Arithmetic Mean * Harmonic Mean = Geometric Mean ^ 2
      • Median :
        • Middle value of the ordered data set
        • ** must be sorted
        • Is much closer to the values in the series
        • Median takes in account relative position of data
        • Median : (n+1)/2 : row number
      • Mode :
        • Most frequently occurring data value
        • Sample modes
          • unimode -> only one most occurred value
          • bimode -> two different values but occurred with same frequency
          • trimode -> three different values but occurred with same frequency
        • Only mode is used to represent nominal series


      Measures of Spread/Dispersal

      • The spread of data.
      • If the mean is the same for two sets of data, but each has a different spread of values. As mean is expected returns, the risk is how far the value from the mean or deviated from the mean.
      • Variability of data is risk.
      • Range:
        • Diff of the highest value and lowest value of data : max-min
        • The problem with range is only taking into account two extreme values into its calculation. But we want to consider all the data points to understand variability of data.
      • MAD - Mean Absolute Deviation
        • Calculate mean and evaluate the difference of each data points against mean.
          • x - x̄
        • As the sum of all the deviation will be always be zero, take the abs (x - x̄) there the sum of all and divide by n.
        • MAD = ∑( | x - x̄ | ) / n
        • This says on a average the spread from mean. 
        • Here the problem is penalization of data away from mean is not visible.
      • Variance:
        • Sum of square distances from each point of the mean
        • The reason its squared, because of two reasons
          • The sum values would less result to 0
          • The more away from the mean the square will penalize and make it further away. The higher the difference the higher the square would take away from mean.
        • Formula
          • population (σ2)

          • sample (s power 2)

        • Bessel's Corrections
          • Why n-1? Loosing the degree of freedom
          • Law of large numbers, as the sample size increase the sample mean approaches population mean.
          • Population mean can take any value, as any value added will be change the population mean based on the data.
          • Sample mean is destined to be population mean, it cannot be anything it want to be. It can take any values until n-1, but not n as it needs to be adjusted to the population mena. Hence it lose the freedom by 1.
      • Standard Deviation:
        • Square root of the variance
        • Same units of the sample
        • On the average, how far are my values from the mean, how reliable are my central values
        • Formula
          • population : σ= square root of (σ2) 
          • sample : s = square root of (s power 2)
        • Now, why we would have n-1
          • When having n :
            • efficient estimate
            • consistent estimate
          • When having n-1
            • unbiased estimate
      • Zscore
        • Zscore, are the ways comparing values across different sets of data where the means and standard deviations are different.
          • Formula : z = ( x – µ ) / ס
          • µ – mean
          • ס – standard deviation
      • Downside Deviation:
        • Standard deviation is the measure of volatility, and spread of data around the centre.
        • Here the downside deviation is used ideally by risk managers where the moving downside movement is risk for the investor.
        • Share Price is 100, if this goes to 120 it is not of concern but if it goes to 80 it becomes risk.
        • Formula:
          • DD = square root (∑((x - D) ^ 2) / n - 1)
        • Eg:
          • If the minimum returns needed is 3% then instead of x̄ it is D and from the example the value of D is 3.
          • Input
            • January: +2% 
            • February: +3% 
            • March: -1% 
            • April: -2% 
            • May: +1% 
            • June: +4% 
            • July: -3% 
            • August: -2% 
            • September: +1% 
            • October: +2%
            • November: +6%
            • December: +1%
          • Calculation
            • Find mean, (2+3-1-2+1+4-3-2+1+2+6+1) / 12 = 1
            • Lets take the months which are greater than equal to 3% and find the difference from the mean and square it.
              • February: +3% = (3-1)^2 = 4
              • June: +4% = (4-1)^2 = 9
              • November: +6% = 25
            • Sum of squared deviations
              • 4 + 9 + 25 = 38
            • Downside Deviation
              • Square root ( 38 / 12 )
              • 1.78%

      Measures of Location

      • Locates the data in the series.
      • Percentile
        • Divides the data into 100 equal parts
      • Decilies 
        • Divides the data into 10 equal parts
      • Quantiles
        • Divides the data into 5 equal parts 
      • Quartile
        • Divides the data into 4 equal parts 
        • Here the values that split the data into equal chunks are known as quartiles, as they split the data into quarters. 
        • Describe data, through quartiles and inter quartile range (IQR)
        • Every data point is considered, as its not aggregated (median is used)
        • How its calculated
          • Divide series into half (median)
            • 0 21 21 31 41 51 | 61 71 81 91 91 200
          • Divide each sub series into further half (median)
            • 0 21 21 | 31 41 51 | 61 71 81 | 91 91 200
          • And now this becomes quartiles
            • 1st quartile : 26
            • 2nd quartile : 56
            • 3rd quartile : 86
          • Quartile ranges would be never the same size in the real world
          • IQR – tells about middle range, from 25 percentile to 75 percentile
          • Outliers
            • Common practice is set to a “fence”, that is 1.5 times width of the IQR
            • Anything outside the fence is considered as outlliers
            • This is calculated as percentiles
          • Formulas
            • IQR = Q3 - Q1
            • Low outliers: anything less than Q 1 − 1.5(IQR)
            • High outliers: anything greater than Q 3 + 1.5(IQR)
      • Example :


      Bivariate Data

      • Compare two variables
      • As said earlier,
        • x axis : independent variable
        • y axis : dependent variable
      • Scatter plot, can be used to find the correlation between two variable
        • But does’t prove causality – while is cause and affect
        • As called to be spurious correlation
        • We may need more statistical analysis is determine causality
        • Reasons of spurious correlation
          • By chance
          • Linked to common variable to another variable, here when one variable and another variable are related to the third variable
          • When relationship emerges by dividing by some common variable
      • Covariance:
        • Compare two variable, to compare their variances
        • How far the each variable’s mean fall
        • If the values are more away from zero it means there are very well correlated
        • Very difficult to interpret covariance, though it shows either positive or negative co-relationship.
        • Formula :
      Credits: https://www.wallstreetmojo.com/
      • Correlation:
        • Correlation shows the linear relationship and can be explained by line for linear data.
        • Formula:
          • cor (x, y) = cov(x, y) / sx sy
            • cov(x, y) = covariance of x and y
            • sx = standard deviation of x
            • sy = standard deviation of y
        • Relation between two variables
          • Status of relationship
          • Direction of relationship
          • Magnitude of relationship
        • Positive Correlation
          • increase x will increase y
        • Negative/Inverse Correlation
          • increase x will decrease y
        • Near Zero
          • No linear relationship
      • Pearson Correlation Coefficient:
        • Compare two variables, with normalization, as not like covariance as its values are taken from variance which is squared
        • Values would now fall between +1 and -1, where
          • 1 = total positive linear correlation
          • 0 = no linear correlation
          • -1 = total negative linear correlation
        • Formula :
      Credits: https://www.wallstreetmojo.com/


      Credits and References

      https://www.bestproxyreviews.com/wp-content/uploads/2020/09/Statistical-Analysis-Methods.jpg


      No comments:

      Post a Comment

      Scarcity Brings Efficiency: Python RAM Optimization

        In today’s world, with the abundance of RAM available, we rarely think about optimizing our code. But sooner or later, we hit the limits a...