Thursday, 24 March 2022

FastAPI



Introduction


FastAPI is a modern and as the name is says fast web framework to build APIs in Python.

Author of this framework is Sebastián Ramírez. Available from Jan 2019 (based on release notes).



Features


  • Performance: Built over ASGI - (Asynchronous Server Gateway Interface) instead of WSGI - (Web Server Gateway Interface)
  • Fast to code: Increase the speed to build the APIs
  • Documentation:  Auto generates the documentation while developing the API 


Fig1: Auto generated document by FastAPI


  • Data validation: Built in data validation that can detect invalid datatype during the run and returns the reason for bad input in JSON format. Pydantic is used for data validation.
  • Based on open standards: Uses OpenAPI for API creation, including declarations of path operations, parameters, body requests, security, etc.
  • Security: Supports all the security standards defined in the OpenAPI standards, such as HTTP Basic, OAuth2, etc.
  • Support: Has small but prompt community to support. Additionally the user documentation is very much detailed.
  • Benchmarks: Comparing to other frameworks here is the performance benchmark result publishing in https://www.techempower.com/benchmarks/ 


Source: https://www.techempower.com/benchmarks/

Fig2: Benchmark results on best response per second





Installation


pip install fastapi uvicorn


# or

poetry add fastapi uvicorn

pipenv install fastapi uvicorn

conda install fastapi uvicorn -c conda-forge


Note: FastAPI does not have a built-in development server, so an ASGI server like Uvicorn or Daphne is required. We will look further using uvicorn.



Usage


Hello World:


Lets create basic hello world program, where the server will return string by calling the root endpoint.


Code: fastapi_runner.py


import uvicorn

from fastapi import FastAPI


app = FastAPI()


@app.get("/")

def home():

    return {"Hello": "World"}


if __name__ == "__main__":

    uvicorn.run("fastapi_runner:app")



Output:


Url: http://127.0.0.1:8000/ 



Fig3: Response from root “/“ endpoint running using FastAPI



Url: http://127.0.0.1:8000/docs

 

Fig4: FastAPI documentation autogenerated for the root endpoint replying the HelloWorld



Create an item:


Let's create post method, which expects gets body with list of attributes. Also we will see if pass invalid input the validation is internally handled by FastAPI.


Code:


import uvicorn

from fastapi import FastAPI

from pydantic import BaseModel

from typing import Optional, Set


app = FastAPI()


class Item(BaseModel):

    name: str

    description: Optional[str] = None

    price: float

    tax: Optional[float] = None

    tags: Set[str] = []


@app.post("/items/", response_model=Item, summary="Create an item")

async def create_item(item: Item):

    """

    Create an item with all the information:


    - **name**: each item must have a name

    - **description**: a long description

    - **price**: required

    - **tax**: if the item doesn't have tax, you can omit this

    - **tags**: a set of unique tag strings for this item

    \f

    :param item: User input.

    """

    return item


if __name__ == "__main__":

    uvicorn.run("fastapi_runner:app")




Documentation:

We could see the auto generator document from the summary and function comments. 



Fig5: FastAPI POST documentation


Invalid Input:

Passing string instead of float in “price” variable.


{

  "name": “first item",

  "description": “first item description",

  "price": "!2",

  "tax": 123,

  "tags": ["#fastapi", "#python"]

}


Error Response: 


Fig6: FastAPI POST invalid error response



Async Example:


Below we will see the usage of query parameter and async & await functions in FastAPI.


Code:

import uvicorn

from fastapi import FastAPI

from pydantic import BaseModel

from typing import Optional, Set

import asyncio


app = FastAPI()


async def goto_sleep(text, sleeptime):

    print(f"before sleep {text}")

    await asyncio.sleep(sleeptime)

    print(f"after sleep {text}")

    return f"{text} {sleeptime}"


@app.get("/slow")

async def slowresponse(st: int = 5):

    print("below slowresponse starts")

    sleptfor_one, sleptfor_two = await asyncio.gather(

        goto_sleep("one", st),

        goto_sleep("two", st),

    )

    print("after slowresponse {} {}".format(sleptfor_one, sleptfor_two))

    

    return {"slow": "completed"}


if __name__ == "__main__":

    uvicorn.run("fastapi_runner:app")


Input:


From the source code, we are expecting the value for query parameter “st” and also have default value set as 5 taken from the “slowresponse” function argument. 



Fig7: FastAPI GET query documentation



Output:


Here we could see the output of calling the “/slow” endpoint. 


Fig8: FastAPI GET response documentation


Logs output: 



Fig9: FastAPI async calls logs



FastAPI vs Flask


Features

FastAPI

Flask

User Documentation

Natively yes

Not natively supported, need to use flask-swagger or such external module 

SGI

Uses ASGI

Uses WSGI

Data validatiaon

Natively yes

Not natively supported

Beginner friendly

Yes

Yes

Concurrent programming

Natively yes using async and await

Not natively supported

Testing of APIs

Yes using fastapi.testclient.TestClient class

Yes using test_client() function

HTTP Methods

Need to create separate decorator for each method

Eg: @app.get(“/“), @app.post(“/”)

Can combine multiple methods in single call.

Eg: @app.route(“/“, methods=[“GET”, “POST”]

Templates

Not natively supported, need to install Jinja.

But also having HTMLResponse to support html response

Natively yes, by importing render_template

CORS

Natively yes

Not natively supported

Authentication

Natively yes, using fastapi.security

Not natively supported, server third party modules are available


Table1: FastAPI vs Flask




Cons

  • Relatively new
  • Small community compared to other frameworks
  • Not a con but from the blogs mostly seems to be famous in ML community


Conclusion

FastAPI is natively new, it is save lots of time in building the Web API easily and with user friendly documentation.


Reference

Thursday, 10 March 2022

Simple Linear Regression Using Python



Machine Learning is the scientific study of algorithms and statistical models that computer systems use in order to perform a specific task effectively without using explicit instructions. 

Supervised is a type of Machine Learning in which machine learns from the labelled data and from the learnings predict the output. Linear Regression is one of the popular model comes under supervised learning. Basically regression algorithms used for the prediction of continuous variables such as house price, stock price, weather forecast, etc.


Fig 1: Supervised Learning

Credits: https://www.javatpoint.com/supervised-machine-learning

Through this blog, we will learn more about Simple Linear Regression, we will understand the concept with the help of an example, learn more about its implementation in Python, and more. 


Description

  • Regression means it searches for the relationship among the variables. Lets say sales of the ice-cream is determined by the weather of the city. If the humidity increases, ice-cream sales increases.

  • Linear Regression is used for finding the linear relationship between target which is the single output variable (y) can be calculated from a linear combination of the multiple input variables (x).

  • In Simple Linear Regression is used to find the relationship between two variables. One is the input variable (x) and other is the output variable (y).

  • For the input and output there are other synonyms as well, knowing this will help you to talk in the same language.

    • Input variable(s) which is represented as (x) is also known as independent variables, predictors

    • Output variable(s) which is represented as (y) is also known as dependent variables, responses.

  • Now we will see what is the simplest algorithm used for prediction 


Learning Model – Ordinary Least Squares (OLS)

  • Most common method

  • Can have more than one input

  • The Ordinary Least Squares procedure seeks to minimise the sum of the squared residuals

  • Each time, we calculate the distance from each data point to the regression line, square it, and sum all of the squared errors together

  • And we will take the parameters of least sum of all squared errors

  • Lets learn through an example.


Example

  • Here we want to predict the price of the wine given the age of the wine.


Input 

  • Here,

    • Output variable (y) is price of the wine

    • And the input variable (x) is age of the wine

    • Note: There could be more than one input variable

  • To understand better we will run through an example with values in it. Below is the input:

Age – in years

Price – in rupees

0

75

1.2

100

2.1

220

3.4

300

4.1

440

5.6

500

7

700



Visualisation

  • Below we see the graphical representation of the data

Source Code:

# Visualising the Training set 

plt.scatter(X, y, color = 'red')

plt.title('Age vs Price of Wine (Training set)')

plt.xlabel('Price of Wine')

plt.ylabel('Age of Wine')

plt.show()

Output:

Fig 2: Data Visualisation

Analysis

  • We could easily see there is a relationship between these two variables, in technical terms this suggests these variables are correlated.

  • The correlated coefficient - R determines how strong the correlation is.

    • There are several types of correlated coefficient but the most popular is Pearson's Correlation

    • R ranges from +1 to -1, further from 0 the stronger the correlation

      • 1 indicates a strong positive relationship.

      • -1 indicates a strong negative relationship.

      • 0 indicates no relationship at all.

    • Graphical viewing a correlation of -1, 0 and +1

Fig 3: Correlation Visualisation

Credits: https://www.statisticshowto.com/probability-and-statistics/correlation-coefficient-formula/#Pearson

  • Pearson’s correlation coefficient formula.

Fig 4: Correlation coefficient formala

Credits: https://www.statisticshowto.com/probability-and-statistics/correlation-coefficient-formula/#Pearson

  • The two variables divided by the product of the standard deviation of each data sample. It is the normalization of the covariance between the two variables to give an interpretable score. 

  • Lets switch back to our scenario of age and price of wine, whether is there a correlation, and if yes what would the R value.

    Source Code:

from scipy.stats import pearsonr

# calculate Pearson's correlation

corr, _ = pearsonr(list(X.flat), y)

print('Pearsons correlation: %.3f' % corr)

    Output:

            Pearsons correlation: 0.985

  • The R value is 0.985 which is very much near to 1, and this is strongly positively correlated

  • Note: In simple terms the shape of the data is important to know to select the model for prediction. So always draw the plot to get the sense of data.



Linear Regression Equation

  • Now that we have seen our data is a good use case for linear regression, lets look into the formula of it.

  • Linear equation

    • y = B0 + B1*x + e

  • Here,

    • y = predicted variable
    • B0 = intercept
      • B0 is the intercept, the predicted value of y when the x is 0. In our example you could see when x is 0 the value of y is 75, also it makes sense since the wine is new we cannot serve it for free right.
    • B1 = coefficient
      • B1 is the regression coefficient – how much we expect y to change as x increases. Again going back to our example, as the number of years increase there is proportion of wine price also increase. We would need to identify how much it increases.
    • e = random error
      • e is unexplained by X component of Y. The error could be reduced by adding more independent variables or by adding relevant independent variables.
  • When we have multiple input variables

    • y = B0 + B1*x1 + B2*x2 + ... + BN*xN

  • The motive of the linear regression algorithm is to find the best values for B0, B1, .. , BN


Model Train Validate Test Split

  • Before train our model, we would need to do Train and Validation split before training our model. 

  • Train: Set of input and output data for the model to learn.

  • Validate: Set of input and output data used for evaluation of the model, to see if the model is not overfitted and reacting good to unseen data. Also used for fine tuning hyper parameters, feature selection, thresholds cut-off selection, and so on.

  • Test: Final set of unseen data to evaluate the final model

    Fig 5: Randomly split the input data into train, valid, and test set. 

    Credits: https://towardsdatascience.com/how-to-split-data-into-three-sets-train-validation-and-test-and-why-e50d22d3e54c


OLS Calculation

  • In OLS method, we have to choose the values of B0 and B1 such that, the total sum of squares of the difference between the calculated and observed values of y, is minimised.

  • Before we write source code in python, lets understand how the OLS works if we had to do it manually.

  • Step 1: Calculate the below data from the below formulas i.e. Sxx – Sum of squares of x, Syy – Sum of squares of y and Sxy – Sum of products of x and y

    • Formulas

      • ̅x – Mean of x

      • ̅y – Mean of y

      • Sxx – Sum of squares of ( x - ̅x )

      • Syy – Sum of squares of ( y - ̅y )

      • Sxy – Sum of products of ( x - ̅x ) and ( y - ̅y )

    • Calculation


Age of wine

Price of wine

x - ̅x

y - ̅y

( x - ̅x ) ²

( y - ̅y ) ²

( x - ̅x )* ( y - ̅y )

1

0

75

-3.34

-258.57

11.15

66859.18

864.37

2

1.2

100

-2.14

-233.57

4.57

54555.61

500.51

3

2.1

220

-1.24

-113.57

1.53

12898.47

141.15

4

3.4

300

0.06

-33.57

0

1127.04

-1.91

5

4.1

440

0.75

106.43

0.56

11327.04

80.58

6

5.6

500

2.26

166.43

5.11

27698.47

375.65

7

7

700

3.66

366.43

13.4

134269.9

1340.08

Sum

23.4

2335

0 (will always be zeros)

0 (will always be zero)

36.36

is

Sxx

308735.71

is

Syy

3300.42

is

Sxy

Average

3.34

is

̅x

333.57

is

̅y 








  • Step 2: Find the residual sum of squares, it is written as Se or RSS

    • Formulas

      • y is the observed value

      • Å· is the estimated value based on our regression equation

        • Note: The caret in Å· is affectionately called a hat, so we call this parameter estimate y-hat

      • y – Å· is called the residual and is written as e 

    • Calculation


      Age of wine

      x

      Price of wine

      y

      Predicted price of wine

      Å· = B0 + B1x

      Å·

      Residuals (e_

      y -  Å·

      Squared residuals

      ( y -  Å· )^2 

      1

      0

      75

      B0 + B1 * 0

      75 - (B0 + B1 * 0) 

      [ 75 - (B0 + B1 * 0) ] ^ 2

      2

      1.2

      100

      B0 + B1 * 1.2

      100 - (B0 + B1 * 1.2)

      [ 100 - (B0 + B1 * 1.2) ] ^ 2

      3

      2.1

      220

      B0 + B1 * 2.1

      220 - (B0 + B1 * 2.1)

      [ 220 - (B0 + B1 * 2.1) ] ^ 2

      4

      3.4

      300

      B0 + B1 * 3.4

      300 - (B0 + B1 * 3.4)

      [ 300 - (B0 + B1 * 3.4) ] ^ 2

      5

      4.1

      440

      B0 + B1 * 4.1

      440 - (B0 + B1 * 4.1)

      [ 440 - (B0 + B1 * 4.1) ] ^ 2

      6

      5.6

      500

      B0 + B1 * 5.6

      500 - (B0 + B1 * 5.6)

      [ 500 - (B0 + B1 * 5.6) ] ^ 2

      7

      7

      700

      B0 + B1 * 7

      700 - (B0 + B1 * 7)

      [ 700 - (B0 + B1 * 7) ] ^ 2

      Sum

      23.4

      2335

      7B0 + 16.4B1

      2335 - (7B0 + 23.4B1)

      Se

      Average

      3.34

      is

      ̅x

      333.57

      is

      ̅y 

      B0 + 3.34B1

      is

      B0 + B1̅x

      333.57 - (B0 + 3.34B1)

      is

      ̅y - (B0 + ̅xB1)

      Se / 7

    • Note: Below formulas how the derivation has been arrived, is out of scope for t his blog. Hint: Used Calculus to differentiate y = (B0 + B1x)^n with respect to x. Calculus is the mathematical study of continuous change, and differentiate calculas cuts something into small pieces to find how it changes.

  • Step 4: Find the regression equation

    • To find the regression equation, we have the B0 and B1 formulas

      • B0 = ̅y - ̅x B1

      • B1 = Sxy / Sxx

    • Lets plug in the values we calculated in Step 1

      • B1 = Sxy / Sxx

        • 3300.43/36.36

        • 90.77

      • B0 = ̅y - ̅x B1

        • 333.57 – (3.34 * 90.77)

        • 30.40

    • The regression equation is

      • y = 30.40 + 90.77x

Note: The values above are rounded, if you run through program there would be slight variations. Also here we should have normalised the value of the price of wine, to see non biased results.


  • So far we have seen the linear equation, just a recap

    • The relationship between the residuals and the slope B1(alias a) and intercept B0 (alias b) is always:

      • B1 : 

        • sum of products of x and y / sum of squares of x

        • Sxy / Sxx

      • B0 :

        • mean of y – mean of x * B1

        • ̅y - ̅x B1



Cost Function

  • Helps us to find the best possible values of B0, B1, .., BN

  • Average is taken, by all the residuals

  • Also known as Mean Squared Error(MSE) function

    Fig 6: MSE Formula

  • We can also use Root Mean Squared Error (RMSE) function as cost function

    Fig 7: RMSE Formula

  • RMSE is more sensitive to outliers. But when the outliers are exponentially rare (like in a bell shaped curve), the RMSE performs very well and is generally preferred.



Correlation Coefficient (R)

  • Now that we have the regression equation, we would need to see how accurate our predicted results are using this equation.

  • Our expectation, shown in the Fig2, the line should be closest to the dots. That is the predicted values should be closer to the observed values(dots).

  • To help to understand the accuracy we can use Correlation Coefficient, we use R to represent the accuracy of a regression equation.

  • This is the same function used before to check the correlation of the data to make sure linear regression model is suitable candidate for this data.

    • Formula:

      • R = sum of products y and ̅y / root of sum of squares of y and sum of squares of ̅y

    • Data

      • B1

        • Sxy / Sxx

        • 3300.43/36.36

        • 90.77

      • B0

        • ̅y - ̅x B1

        • 333.57 - (3.34 * 90.77)

        • 30.40


          Age of wine

          x

          Price of wine

          y

          Predicted price of wine

          Å· = B0 + B1x


          Å· = 30.40 + 90.77x 

          Residuals (e =

          y -  Å·

          Squared residuals

          ( y -  Å· )^2 

          1

          0

          75

          30.11

          44.89

          2014.8

          2

          1.2

          100

          139.05

          -39.05

          1524.68

          3

          2.1

          220

          220.75

          -0.75

          0.56

          4

          3.4

          300

          338.76

          -38.76

          1502.24

          5

          4.1

          440

          402.3

          37.7

          1421.04

          6

          5.6

          500

          538.47

          -38.47

          1479.97

          7

          7

          700

          665.56

          33.44

          1186.15

          Sum

          23.4

          2335



          Average

          3.34

          is

          ̅x

          333.57

          is

          ̅y 



          Se / 7 = 9129.42

  • To check how much variance is explained by our regression equation we use R^2. And it is also called as Coefficient of Determination.



Model Training

  • Its really simple to write code in python, as we have all the functions already written for us to use.

  • Module called sklearn we will be using for the task.

  • Source Code:

from sklearn.linear_model import LinearRegression 

lm = LinearRegression() 

lm.fit(X_train, y_train) 

y_pred = lm.predict(X_test)

  • Here,

    • X_train, is training input

    • y_train, is training expected output

    • X_test, is testing input

    • y_test, is testing actual output

    • y_pred, is testing predicted output

  • Methods used,

    • LinearRegression() : LinearRegression fits a linear model with coefficients B = (B1, …, BN) to minimize the residual sum of squares between the observed targets in the dataset, and the targets predicted by the linear approximation.

    • fit() : method will fit the model to the input training instances

    • predict() : perform predictions on the testing instances, based on the learned parameters during fit


Model Output

  • Output Source Code

    • y_pred

      • array([ 30.11355599, 139.04715128, 220.74734774, 338.75874263,
       402.30333988, 538.47033399, 665.55952849])

  • print("Coeffient(B1) {} and Intercept(B0) {}".format(lm.coef_, lm.intercept_))

    • Coeffient(B1) [90.77799607] and Intercept(B0) 30.11355599214147


Predictions

  • Now that we have trained our model lets see how its predictions have come.

    Fig8: Prediction Vs Actual

  • Source Code

plt.title('Prediction of Price of Wine based on the given Age')

plt.xlabel("Age of wine") #x label

plt.ylabel("Price of wine") #y label



plt.scatter(X_test, y_test, c ="pink",

marker ="s", 

edgecolor ="green",

s = 100,

label = "Actual")

plt.scatter(X_test, y_pred, c ="yellow",

marker ="^", 

edgecolor ="red",

s = 80,

label = "Prediction")

plt.legend()

  • The model has really performed well, but as you see we were not 100% correct in our predictions, so how confident we are about our model's prediction. For this to say we need to know about Confidence Interval (seen above).

  • Side Note: Generally speaking we actually don't want our model to be 100%, because it would be considered as overfitted, and will not generalise well. (Not that with 90% accuracy model will generalise.)

    To say in simple terms assume model as a student, and it is going for English Exam. We have given the model few questions to learn from, coincidently the same questions appeared in exam. Model got excited and did well and scored 100. 

    But will always that would not be the case right, assume there were some questions bit reordered. As the model had mugged up the answers, it couldn't understand and failed. 

    So what we learn from here is instead of giving few questions give the model the whole book (i.e. more data). You see its hard to mug up whole book, model will start learning and finding patterns. And this will be more generalise for unseen similar data, which would perform well in production.


Assumption

  • Input and Output variables should be numeric
  • Outliers should be removed or replaced
  • Remove Collinearity between input variables (X), as Linear regression will over-fit your data when you have highly correlated input variables, there should be collinearity with input and output but not within input variables
  • Rescale Inputs, as Linear regression will often make more reliable predictions if you rescale input variables using standardization or normalization
  • Relationship between independent and dependent variables should be linear
  • The error term must be additive.


Advantages

  • Predictions are fast.
  • Good if inference is needed for the prediction.


Drawbacks

  • Assumption, the relationship between the independent and dependent variable is linear.
  • Affected by the outliers.



References


Original Blog Posted in OSFY

https://www.opensourceforu.com/2022/04/an-introduction-to-low-code-and-no-code-test-automation-tools/

For further research and updates maintaining the blog here.


Scarcity Brings Efficiency: Python RAM Optimization

  In today’s world, with the abundance of RAM available, we rarely think about optimizing our code. But sooner or later, we hit the limits a...