Thursday, 9 June 2022

AWS Cloud Security Services - Infrastructure Protection

Cloud Security

Cloud eases the hardware and application management, provides easy accessibility. But there are several assumptions and reservations about switching to cloud due to security.


Well if we talk about attacks it's every where, whether its cloud or on premisses, either external or internal. Hence we need to be ready always.


Cloud security, is a collection of security measures designed to protect cloud-based infrastructure, applications, and data from the external and internal threats.


Cloud security is employed in cloud environments to protect a company's data from different security concerns such as distributed denial of service (DDoS) attacks, malware, hackers, unauthorized user access or use and many other security threats.


A reliable cloud service provider (CSP) can put your mind at ease and keep your data safe with highly secure cloud services.

  • Resource Drain: Lowers the burden from the developers in monitoring the attacks
  • Less Expertise Dependency: We need the expertise in security for any org, but the load would be shared and dependency as well. 
  • Eases Multiple Integrations: Having multiple applications and integration with security standards and monitoring itself is whole new project.
  • Avoid Expensive tools: Analysis and prevention of the attacks would require tools and that would add up the cost

Way to approach cloud security is different for every organization and can be dependent on several variables but following the best practices is a good start. Putting in place adequate countermeasures to defend against modern-day cyberattacks.


Both CSP and Cloud Adopter are equally responsible in providing security.



Categorization
  • Identity and Access Management
    • AWS Identity and Access Management IAM
    • AWS Single Sign On
    • Amazon Cognito
    • AWS Directory Service
    • AWS Resource Access Manager
    • AWS Organization
  • Detection
    • AWS Security Hub
    • Amazon GuardDuty
    • Amazon Inspector
    • AWS Config
    • AWS CloudTrail
    • AWS IoT Device Defender
    • AWS Detective
  • Data Protection
    • AWS Macie
    • AWS Key Management Service
    • AWS CloudHSM
    • AWS Certificate Manager
    • AWS Secrets Manager
  • Infrastructure Protection
    • AWS Security Groups
    • AWS NACL
    • AWS WAF
    • AWS Shield
    • AWS Network Firewall
    • AWS Firewall Manager


Infrastructure Protection

Infrastructure protection is one on the key pillar of cloud services. We will only see today using AWS services to protect our infrastructure. 

Security Groups

Security groups are acting as a “firewall” on EC2 instances. Using security group, we can control both the incoming and outgoing traffic. They are stateful, which means any traffic in is allowed to go out, can go back in. So we do not need to write explicit outbound rules, as said if it is allowed in, the traffic will be allowed out as well. Supports only allow rules. 

It can reference by CIDR and security group id, there can be more than one rule assigned to an endpoint. If no rules are applied the default will be assigned, which will deny all the inbound and allow all the outbound. Evaluates all the rules before deciding whether to allow traffic. 

Security groups comes with no cost, one can have as many security groups. As you could see this is joint responsibility of the cloud consumer, the cloud provides the feature to block or allow requests. But as the user, we need to write appropriate rules.

Default: inbound denied, outbound all allowed.
Creating security groups does not involves cost.

Below are few use cases mentioned from the aws docs.

  • Web server rules
  • Database server rules
  • Rules to connect to instances from your computer
  • Rules to connect to instances from an instance with the same security group
  • Rules for ping/ICMP
  • DNS server rules
  • Amazon EFS rules
  • Elastic Load Balancing rules
  • VPC peering rules


Fig: Sample rules, to allow ssh and http from everywhere 



Network ACL - NACL



NACLs are acting as a “firewall” at the subnet level. We can associate a single NACL to multiple subnets, but only one NACL can be assigned a subnet.

These stateless as the inbound and outbound rules apply for all traffic. For example, if the output traffic is allowed our request would reach out, but if expect a response, sorry it would need to have the outbound rules specified. Supports all both allow and deny rules for both inbound and outbound traffic. Here we can specify only reference a CIDR range (no hostname). 

Evaluations rules in number order when deciding to allow traffic, lowest numbered rule gets the first preference. If the lowest rule denies, the traffic will be allowed, even if the higher rule allows. The rules starts from 1 to the highest of 32766.  

Default: allow all inbound, allow all outbound
New NACL: denies all inbound, denies all outbound

Fig: Sample rules, which allows and denies the outbound traffic


WAF



AWS WAF is a web application firewall that helps protect your web applications or APIs against common web exploits and bots that may affect availability, compromise security, or consume excessive resources. This is usually at the layer 7 which is application layer.


AWS WAF gives you control over how traffic reaches your applications by enabling you to create security rules that control bot traffic and block common attack patterns, such as SQL injection or cross-site scripting, limits the number of calls to the server, limits the request size.


We can get started quickly using Managed Rules for AWS WAF, a pre-configured set of rules managed by AWS or AWS Marketplace Sellers to address issues like the OWASP Top 10 security risks and automated bots requests which exploit the application. 


Deploy AWS WAF on Amazon CloudFront, the Application Load Balancer, Amazon API Gateway and AWS AppSync. 


Pricing is based on how many rules we deploy and how many web requests our application receives. Irrespective of the pricing using WAF is definitely essential:

  • Frictionless setup, deploy without changing your existing architecture
  • Low operation overload, ready to use, built in set of rules and as well rules set available in AWS marketplace
  • Bot control, protecting against automated bots using bot controlled rules. This has been designed to help stop common and pervasive bot traffic on your application. AWS Threat Researchers continuously examine the traffic to identify and categorise bots. Combine the natively available mitigation techniques on AWS WAF. Also provides insights and visibility into bot traffic through CloudWatch.
  • Customizable security, highly flexible rule engine that can inspect requests with single milliseconds latency
  • Advanced automation, API driven architecture and fast rule propagation allows you to detect and respond to the threats in real time


Fig: First page of WAF


Shield


AWS Shield is a managed Distributed Denial of Service (DDoS) protection service that safeguards applications running on AWS.

A DDos is a cyber-attack in which the perpetrator seeks to make a machine or network resource unavailable to its intended users by temporarily or indefinitely disrupting services of a host connected to a network.  Denial of service is typically accomplished by flooding the targeted machine or resource with superfluous requests in an attempt to overload systems and prevent some or all legitimate requests from being fulfilled.

There are two tiers of AWS Shield:
  • AWS Shield Standard
  • AWS Shield Advanced
For all AWS customers by default AWS Shield Standard is enabled for free. This provides always-on detection and defends against most common, frequently occurring network and L3/L2 transport layer DDoS attacks that target your web site or applications. 

AWS Shield Advanced provides more advanced DDoS detection, near real-time visibility into attacks. Supports integration as of now with AWS WAF a web application firewall to EC2, Elastic IP, ELB, Amazon CloudFront, AWS Global Accelerator and Amazon Route 53 charges. It protects whatever resources associated with it and provides protection for network layer (layer 3), transport layer (layer 4), and application layer (layer 7) attacks. Costing $3000 per month per organization.

AWS Shield Advanced, provides access to Shield Response Team which is 24/7. And cost protection during the attacks. A team of specialized security engineers dedicated to provide the support during the DDOS attacks. They will help us in shield onboarding as well. Proactive engagement during the DDOS attacks, attack analysis, writing custom WAF rules for mitigations, fighting against the bots and mitigation strategies are also part of their responsibilities. AWS Shield Advanced, provides AWS WAF with no additional charges. 

Need to subscribe to AWS Shield Advanced for each AWS account that you want to protect. If you want to subscribe multiple accounts, it is recommended to use AWS Firewall Manager. Caveat, Firewall Manager doesn't support Amazon Route 53 or AWS Global Accelerator, but it supports the other resource types that can be protected by Shield Advanced. It provides supports globally.

Well one might wonder this seems to be an expensive affair, yes I can agree to this. But when it comes to business having the application resilient, audit complaint is most important. Imagine if Swiggy is not able to identify and serve potential customers' it would be a great loss for the organization. For instance during an attack, Shield Advanced promotes your network ACL to the AWS border, which can process multiple terabytes of traffic. Your network ACL is able to provide protection for your resource well beyond your network's typical capacity. It means provides cost protection during the attacks.

The story doesn't ends here, load balancer's mapped with auto scaling is usually the strategy. Imagine the costs it would shoot up for while handing these mock calls. Hiring a security engineer and his/her backup will just add up the cost so much. Compliance and regulatory also goes on toss with these attacks.  As said this is worthy investment for an application. 


Fig: Subscribe to Shied Advanced


Network Firewall

Network Firewall protects at L2/L3 which is Network/Transport layer of the VPC. Based on the security rules written, all the traffic following into and out of the network is managed and monitored. This can be setup with just a few clicks and scales automatically based on the network traffic.


Provide protections from common network threats. 

  • URL filtering on outbound flows
  • Pattern matching on packet data beyond IP/Port/Protocol 
  • Ability to alert on specific vulnerabilities for protocols beyond HTTP/S
  • Stateful Inspection
  • Intrusion prevention and detection

Using Network Firewall we have below merits:

  • AWS Network Firewall’s flexible rules engine lets you define firewall rules that give you fine-grained control over network traffic
  • Import rules already written in common open source rule formats 
  • Enable integrations with managed intelligence feeds sourced by AWS Partners

Pricing

  • pay an hourly rate for each firewall endpoint.
  • pay for the amount of traffic, billed by the gigabyte.
Below architecture explains now Network Firewall works, lets dig deep into it:
Credits: https://aws.amazon.com/blogs/aws/aws-network-firewall-new-managed-firewall-service-in-vpc/
Fig: Explains the Network Firewall Architecture

From the above we observe below:
  • AWS Network Firewall Manager act on VPC Level, for an Availability Zone.
  • Firewall endpoint is in the public subnet, which acts as fence for all the requests coming in. The firewall endpoint insects the incoming and outgoing packets based on the rules configured.
  • Rules engines,  holds the firewall policy which holds the collection of stateful and stateless rule groups.
    • Stateful: Stateful firewalls are capable of monitoring and detecting states of all traffic on a network to track and defend based on traffic patterns and flows. This is An intelligent system, stateful firewalls base future filtering decisions on the cumulative sum of past and present findings.
    • Stateless: Stateless firewalls, however, only focus on individual packets, using preset rules to filter traffic. This provides faster performance.
  • Rule groups inspects the packets and based on perform the configured action.
  • Actions, below are four types of action can take place with the rules make and we can the default action as well applying for all the packets
    • Pass: All the packet to reach its destination
    • Drop: Block the packet to proceed further
    • Forward to stateful rules: Proceed with stateful inspection
    • Custom action: Sends the metric to CloudWatch with value specified in the configuration by us
This is the simple illustration of how Network Firewall secures. 



Firewall Manager


Firewall Manager helps in centrally configure and manage security rules across all accounts. This brings consistency and enforces protections as said across all the accounts, even as new applications are created. This provides single view compliance posture centrally across all the AWS accounts.


AWS Firewall Manager currently handles six types of protection policies - AWS WAF, AWS Shield, Amazon VPC security groups, AWS Network Firewall, Amazon Route 53 Resolver DNS Firewall and Palo Alto Cloud Next-generation firewalls.  


Prerequisites of using Firewall Manager are below, well also you need to run this from your administration account: 

  • AWS Organization: To manage all accounts
  • Firewall Administration: To deploy AWS WAF rules across
  • AWS Config: To detect newly created resources


Firewall Manager creates below impact:

  • Simplify management of firewall rules across all accounts
  • Easily deploy managed rules across accounts
  • Centrally deploy protections for VPCs
  • Audit any existing security group in the VPC
  • Control traffic leaving and entering network
  • Protection policies are priced with a monthly fee 100$ per policy per region

Credits: https://aws.amazon.com/blogs/aws/aws-firewall-manager-central-management-for-your-web-application-portfolio/
Fig: After applying the rules validation the compliance of all the accounts


Summary

Let's recap with a small story, assume a mango seller wanting to digitize my business. I have created a web application running on ec2 instance.  

  • Allow SSH for connecting from my home to the EC2 instances and HTTP from anywhere – Security Group
    • Security Groups to protect Amazon Elastic Compute Cloud (Amazon EC2) instances
    • Here we are allowing requests from everywhere on HTTP to hit my website
    • And allowing SSH only from IP so that I only can get into the machine
As my business is doing good, but I have one competitor, As his business is getting impacted because if me, he wants to bring my application down. 

  • Block the competitor – NACL
    • Network ACLs to protect Amazon Virtual Private Cloud (VPC) subnets
    • We can block specify IPs entering our subnet, as in NACL we have both Allow and Deny
    • In Security Groups I cannot deny
    • NACL has evaluations rules in number order when deciding to allow traffic, will lowest numbered rule
Now my business is doing good, but at random times getting heavy requests at same time causing hinderance to my system. Well this looks fishy I looked towards AWS how it can solve my problem.

  • Block the HTTPS calls from users sending more than 1000 requests – WAF
    • AWS Web Application Firewall (WAF) to protect web applications running on Amazon CloudFront, Application Load Balancer (ALB), App Sync  or Amazon API Gateway

As my business now become stable I have scaled my applications DB to private subnet, one more private subnet for the application servers and created public subnet only for the load balancer. Again am getting spike requests from different countries where absolutely I don't trade.

  • Block the requests for the VPC geowise – Network Firewall
    • AWS Network Firewall, a high availability, managed network firewall service for your virtual private cloud (VPC).
We have our applications running, but suddenly costs of the hardware spikes but that's not reflecting in the sales. I want to avoid these unexpected attacks and I don't want to invest in the security engineer for this.

  • Protect my system from DDOS attacks – DDOS
    • AWS Shield to protect against Distributed Denial of Service (DDoS) attacks.
Fantastic, I was called as best mango app. Now am opening branch in different country. Wow! But all the security lessons I don't want to loose. I have created another one application for other country but the security underline remains the same. I want the common rules should be same across my org.

  • Ensure all security is aligned across accounts – Firewall Manager
  • Even with accident changes, have been configured such as to revert the changes.

Let's visualise our story, please take a close look into all the security services applied:
Credits: https://www.youtube.com/watch?v=T3kqljTLR50



Conclusion

Today we have seen few of the security services which are available now current date. We have not covered the access management, data security and detection aspect of the security. 


Security comes with cost and we cannot be 100% secure, its evolving process. And cloud provided and consumer have equal contribution in bringing protection and resilience to the application.


Security is not easy, we learn from experience and research. 



References

https://docs.aws.amazon.com/

https://www.prplbx.com/resources/blog/aws-overview/

https://medium.com/nerd-for-tech/aws-series-2-deep-dive-aws-security-layer-network-web-apps-a629f60631ef

https://www.youtube.com/watch?v=T3kqljTLR50

https://medium.com/binbash-inc/aws-network-firewall-using-aws-firewall-manager-with-terraform-part-2-b402dffecfb0

https://cloud.in28minutes.com/aws-certification-security-groups-vs-nacl-comparison

https://www.cdw.com/content/cdw/en/articles/security/stateful-versus-stateless-firewalls.html





Thursday, 2 June 2022

Is Your Database Meeting Your Needs?





Introduction

    Data is new oil this is heard everywhere and also said to be heart of the organization. Well I totally agree to this, every organization makes decisions either long term or short term based on the available data.    

    Database is where data resides, this makes choosing right database very much important and challenging as well. There are different types of database SQL vs NoSQL, centralised vs distributed, in memory vs persistence storage, etc and each has its own merits and demerits in various terms including how reliable it is. 

    Today we will see three concepts ACID, BASE and CAP each which defines the various dimensions such as consistency, guaranteed transaction, stable distributed system, and so on.  There is no best concept, each carries its own merits and challenges. Let's get deep dive into how each concept helps in assuring data consistency or data availability or transaction reliability and how each concept is different from another.



ACID

Credits: https://morpheusdata.com/cloud-blog/when-do-you-need-acid-compliance/


    For any database, let say SQL/NoSQL/GraphDB ACID is an important first principles when it comes to handling a transaction.

    Before we get further, let's know about Transaction. It is collection of queries, which is considered as one task. So we need to run all or none at all. If something in between fails either intentional or non-intentional, everything needs to be rolled back to the start just before the transaction begin.

E.g. XYZ Transferring 100rs to ABC
  • Transaction Begin
  • Read Balance from XYZ
  • Proceed only if XYZ > 100rs
  • XYZ Debit 100rs
  • ABC Credit 100rs
  • Transaction Commit 
    Here in the above example we have 100rs transaction happening if anything fails in between, XYZ should continue holding back the 100rs, rather its debit from him but not credited to ABC. That is not good right.

    There is another flavour of transaction where it would be read only but still require to be in transaction. We call it Read-only transaction. Well one might think why we ever need it, let's see with an e.g. its requirement:
  • Without Transaction:
    • Count of Number of Deals booked -> Job 1
    • 1 Deal got cancelled -> Job 2
    • Count of Dealers along with count of Deals booked -> Continuation of Job 1
    • Here you see there would conflict in the answers produced by Job 1
  • With Transaction
    • Transaction Begin -> Job 1
    • Count of Number of Deals booked -> Job 1
    • Count of Dealers along with Deal booked -> Continuation of Job 1
    • Transaction Commit
    In this case, all the reads within this transaction gets consistent data from the database even during interruption from another jobs, the current transaction would not be affected. Usually used while building huge report or taking snapshot.


Atomic

    Atomicity represents as single unit, it should not break. All the queries in a transaction must succeed, else nothing occurs - all should be rollback. There are two kinds of issue causing failure, during a transaction.
  • Query failure, such as duplicate primary key insert or update invalid data and so on
  • Database failure, such as hardware crash, power issue and so on
    In case of query failure, we still have the control and performing the rollback. But in case of database failure such as hardware crash or crash while performing commit. These are the serious issues which needs to be handled sensitively.

    Some databases write in the disk directly and in the end mark the transaction as success, here the commit would be faster but involves bad data risk. To avoid the bad data, some databases copy the old data before applying changes we call them read-copy-update, well here hard disk space grows rapidly though. 

    While some databases maintain changes in the RAM and only during success write in the database. Here the rollback is faster but holds issues such as RAM overload and RAM crash. As you know there is no free tea, we need to sacrifice on fast write or fast rollback here in these two scenarios.

    Typically several systems implement atomicity by journaling. The system synchronises the logs as necessary after changes have successfully taken place. After crash recovery, ignores incomplete entries. Although different databases implementations vary depending on many factors such as
  • Concurrency issues, allowing the threads of users working on same data using lock mechanism. Having time gaps to avoid dead locks.
  • Control the number of operations included in a single transaction. Here both the throughput and the possibility of deadlock is controlled. The more the operations included in a transaction, that more likely it is that a transaction will block other operations and deadlock will occur.
  • Start databases only when stale transactions are clear, well in some databases they start immediately and meanwhile crashed changes will be fixed in parallel.
    Although the principle of atomicity - i.e. complete success or complete failure - remains.


Consistency

    Consistency of the data in persistent disk, as well consistency means data is in a consistent state when a transaction starts and when it ends. The progress made from one valid state to another, taking care of auto trigger changes, ensuring constraints and so on. For example:
  • When the data transfers from A to B, it should verify B exists and values after transaction completes should be valid without any corruption
  • Another example from Hussien's course, where he explains having two tables
    • Picture: Picture ID, Number of Likes to the Picture
    • User_Liked_Picture: UserID, Liked Picture ID
    • Here there should be consistency across both the tables where count of likes of the picture should match with users who like it. If there is a mismatch we can call this scenario as inconsistence referential integrity.
    Consistency in reads in also equally important, it would be little simple in single db where you write and read from the same. But having distributed databases makes it complicated, if a write is committed and we tend to read from it, well if its same database it would be prompt but in distributed the write needs to be reached from where it is asked for read. But if we get at times stale information, this is where we have concept called Eventual Consistency. It says it's not consistent but will provide required information after sync. Early results of eventual consistency data queries may not have the most recent updates because it takes time for updates to reach replicas across a database cluster.
    
    At times one would need strong consistency, well we need to have synchronous synchronisation though it would slower but strong vs asynchronous synchronization which would be eventual but faster.


Isolation

    Usually many transaction run in parallel. Isolation guarantees the intermediate state of a single transaction should be invisible to other transactions, leaving the database in the same state that would have been obtained if these transactions were executed sequentially. As a result, these transactions that run concurrently appears to be executed in the serialized or sequential order.

    For example, in an application that transfers funds from one account to another, the isolation property ensures that another transaction see the transferred funds in one account to the other, without leaving both the accounts with credit, as well nor debits in both.

    There is problem which isolation property faces due to concurrency, called as read phenomena.
Read Phenomena(Problems): When a transaction reads data that another transaction might have changed. Below are few phenomenas:
  • Dirty reads: Reading the data which is not fully committed
  • Non repeatable reads: Reading twice the same data. Not necessary using the same query, could be different as well but reading the same data. For example read the sum of picture likes and users liked the picture -- here the values would have changed due to other transaction, resulting the mismatch between sum of likes from user table vs total number of likes from pictures table. 
  • Phantom reads:  Re-reading the non-existence data, such as reading the sum of likes but then new record got inserted by another transaction and you have not read it yet.
  • Lost updates / Serialization anamoly: Read the written data, but unable to find it due to other transaction deleted it or modified.
Isolation levels define the degree to which a transaction must be isolated from the data modifications made by any other transaction in the database system. A transaction isolation level is defined by the above phenomena.
  • Read Uncommitted Isolation: Allowing to read uncommitted changes across transactions, this will be fast as nothing is needed to be maintained. But down fall there would be several dirty reads happening.
  • Read Committed Isolation: Most popular and default isolation level, while running a transaction can and only see the committed changes by other transaction.
  • Repeatable Read Isolation: Transaction when reads a row will remember it and keep it unchanged which would be been affected by other transactions. This remembered value remains the same as read before hence when re-read happens it would not be affected. The caveat here is this is expensive, but at the cost of valid reads.
  • Serializable Isolation: All the transaction run here serially, losing the concept of concurrency. This would be the slowest of all.
  • Snapshot Isolation: Each transaction takes the snapshot of the committed changes before the start of the transaction and here the new inserted rows causing Phantom reads will not occur as it gets read from the snapshot. If the number of concurrent transactions increases it would costly in terms of space holding the N number of snapshots.
Credits: https://alxibra.medium.com/isolation-level-in-rails-847edc9e347d



Durability

    All the transaction successfully information needs to be stored in persistent storage (nonvolatile system storage), even in the event of a system failure such as hardware crash or power failure, we can still recover the committed changes it should not be lost.

    Some databases to make write faster they tend to write in to RAM and in backend batch wise snapshot to the disk, well yes it would be faster but less durable.

    A realtime example ensuring durability is, in an application that transfers fund from one account to another, the durability property ensures that the changes made to each account has made persistent changes and irreversible.

Durability Techniques:
    Below are few techniques ensuring high durability:
  • WAL - Write ahead log
    • Write in the direct database are expensive and also risky when performing uncommitted changes, during failure need to revert everything.
    • WAL records the changes in the log in stable storage and later transferred to the db.
    • Well here there is always an issue with OS Cache, which would write in RAM and transfer in the hard disk batch wise, in DB we need to record every item in persistent database. Hence calling fsync OS command resolves this, as it forces writes to always to the disk.
    • Checkpoint is a point where the program writes all the changes specified in the WAL to the db, and clears the logs.
  • Snapshot
    • Generate a complete snapshot of the data set in the memory within the specified time interval, and then save it in the hard disk.
    • The save operation could be synchronous or asynchronous.
    • But it has primary two issues with this technique:
      • The last in memory data may be lost. Because there is a time interval for generating snapshots, it will cause the last data to fail to generate a snapshot.
      • Time-consuming and performance-consuming, because the data in the memory needs to be completely snapshotted and written to the hard disk.
  • AOF - Append Only File
    • Write any write operations to the log file, only append files but not rewrite files
    • AOF file is essentially a redo log, through which database state can be restored
    • When the data is restored, the file is read to rebuild the data, that is, if database is restarted. The write command will be executed from front to back according to the content of the log file to complete the data recovery work.
    • As the number of commands executed increases, the size of AOF file increases, which can lead to several problems, such as increased disk usage and slow restart loading. But there are several techniques to control this issue.

So far we have seen ACID which has an impact while selecting a database where loads to transactions are involved. Mostly we have transactions running in the relational databases, which ensures the consistency.


BASE

    In distributed environment in general BASE concept is applied, as here the ground rule is availability over consistency. BASE model provides high availability as the primary principle.    

    NoSQL in general follows BASE.  NoSQL, the database which stores data in documents such as key-value pair, wide-column, pure documents, graph rather the traditional relational model easy to query. NoSQL enables rapid changes, believes in zero downtime, supporting large users from globally distributed.
    
    As said, NoSQL databases usually it adheres to BASE. But we have few NoSQL databases that holds certain degrees of ACID compliance.

    Below are the three properties of BASE.

Basic Availability

    The system is guaranteed to be available for all the users all the time, even in the event of failure. Ideally the system should always available, should reach the replica in case of failure, should not lock up and the data should be available while the operation is in progress. Immediate consistency is not the priority, the key is availability assurance.

Soft-state

    As we are not consistent, the stored values may change due to the eventual consistency model. Here we could get stale information which was before sync. It becomes developer's responsibility to make sure of the consistency.

Eventual consistency

    The fact currently the data might be stale, but the system state is gradually replicated to all the nodes. Data reads are still possible, but it might not reflect the reality.

From the above principles it's quite clear, BASE favours availability over consistency.



CAP

Credit: https://medium.com/@sumitsethia94/consistency-or-availability-of-databases-do-you-really-understand-cap-8ecf2b3bb099

    Now let's look into CAP theorem, it is key to understand when your database needs to be distributed. It can be applied to both NoSQL and SQL databases.

    Some systems needs to be live all-time 24*7*365, well it is hard to have them running in single server and single data centre, as there could be hardware crash, power failure or any natural disaster in the city. As a result we always need to anticipate and plan for system failures and design our system running from different servers and different locations all the time to users better. Additionally data should be reliable along with highly available.

    The CAP theorem is also called Brewer's Theorem, because it was first advanced by Professor Eric A. Brewer during a talk he gave on distributed computing in 2000. Two years later, MIT professors Seth Gilbert and Nancy Lynch published a proof of "Brewer's Conjecture".

Consistency

    This consistency is different from ACID property's consistency. Here a consistent view of data at each node means all the nodes always have the same and most recent view of the data. If not yet the most recent or if it is not consistent, error must be returned. Any write to the database should be immediately replicated to the replicas. Though this will be most challenging to achieve, syncing globally at same time is hard. There are several techniques such as change log capturing, querying based on the timestamps, acknowledge from maximum nodes and so on to overcome and provide the fastest replication across the nodes (no here I didn't want to call this eventual consistency).     
    Also at times consistency goes on toss when we read from the non-updated cache. Example of consistency would be maintaining all the changes done to your bank balance should be same all the times in every system where it is distributed and saved. And every retrieval should give the same balance from wherever it is checked.

Availability

    Availability of data at each node means the system always responds to requests. Here we would need to get the data but it may not be the most recent. Getting the response is most important, it should not give error as the response. Caching can be used here, to respond back even though it might be stale. Well yes this might sound contradicting with consistency, but ideally in availability getting response is very important. Taking YouTube as an example, we would need videos may not have all the videos or comments count up to date.      

Partition Tolerance

    Tolerance to network partitions means systems remain online if network problems occur. During power shutdowns, or natural disaster, the partial system might be down, the system is still tolerance to accept and processes the user's requests. Ideally here replication across nodes and network helps. For example, if I hold account in Chennai bank branch but there is power shutdown in the city, it should not stop me from accessing the money from other or same location.

CAP Gaurantees

    In CAP theorem, guarantees only two properties at a time, we cannot achieve all three together. We need to tradeoff based on our requirement. From few examples let's see why it could be only 2 not all 3. Here it talks that either the system can be consistent or available not both for the distributed design.

AP - Availability and Partition Tolerance
Here the application needs to be provide high levels of availability. Eg:
  • You have 1 master node and 2 replica nodes
  • Write to the master, get the success immediately after write access
  • Asynchronous replication takes place in the backend
  • 1 of the replica lost it connection with the master for time being
  • Now the disconnected replica is contacted it gives the response but its stale
  • As per our requirement it should be available we are good here, but not consistent
  • YouTube, here needs to be highly available

CP - Consistency and Partition Tolerance
Here the application needs to be all time consistent. Eg:
  • You have 1 master node and 2 replica nodes
  • Write to the master, get the success only if all the nodes or maximum nodes get the replica applied successfully
  • Here the write is blocked until it gets success message
  • The blocking hence is making the system slow and not available
  • The replica set nodes in the system send a heartbeat (ping) to every other node to keep track if other replicas or primary nodes are alive or dead. If no heartbeat is received within N seconds, then that node is marked as inaccessible
  • Yes we can call it partition tolerance, if one replica goes down the other can still serve the purpose
  • As the commit happens in the end, the consistency is maintained
  • Each node also maintains the operations logs, for the fallback server coming up and getting synced like other nodes
  • Bank systems, needs to be highly consistent

CA - Consistency and Availability
    CA database enable consistency and availability across all nodes. RDBMS databases are CA guarantee. As these systems are usually in single system hence we can have both Consistent and Available.
    CA database can't deliver fault tolerance. Hence these cannot be used as distributed systems. 

Conclusion

    Today we saw  needs to be ACID properties which helps in determining how to select database based on each of its parameter when performing transaction. As ACID is mostly used in RDMS - SQL because of its nature, it needs to be highly consistent. For NoSQL which is distributed nature, uses BASE which needs to be highly available. ACID is usually used in the banking applications and BASE in general used in social network applications.
    In the end we have CAP Theorem states that is impossible to have both the worlds together which is high availability or high consistency. There are many database build each specialises in either of the parameters, all that we need to understand our requirement and see which best matches with the available databases.
    What is your requirement? Dose the database meets your database meets your needs?

References

https://alxibra.medium.com/isolation-level-in-rails-847edc9e347d
https://www.udemy.com/course/database-engines-crash-course/learn/lecture/28927902?start=0#overview
https://phoenixnap.com/kb/acid-vs-base

Scarcity Brings Efficiency: Python RAM Optimization

  In today’s world, with the abundance of RAM available, we rarely think about optimizing our code. But sooner or later, we hit the limits a...