Monday, 8 December 2025

Scarcity Brings Efficiency: Python RAM Optimization


 

In today’s world, with the abundance of RAM available, we rarely think about optimizing our code. But sooner or later, we hit the limits and realize it a bit too late. This blog explores how we can make better use of RAM by applying various Python optimization techniques.

Memory Usage Stats

Let's begin with how to measure the amount of RAM used by a python program.  As there are always more than one ways to perform an action and each has its merits.

Using sys.getsizeof:

This is a built-in function from the sys module that returns the shallow size of a single Python object in bytes.

  • What it measures: Only the memory directly attributed to the object itself (e.g., the list structure), excluding the size of any referenced objects (e.g., the integers or strings inside a list).
  • Pros:
    • Simple, fast and lightweight.
    • Useful for quick comparisons of basic object overhead (e.g., empty list vs. small list).
  • Cons/Limitations:
    • Shallow only as it underestimates containers like lists, dicts, or custom objects with references values in them.
    • As this is not recursive, it doesn't help with complex structures or total memory footprint.
  • When to use: For basic, one-off checks on simple objects or understanding Python's object overhead. Not suitable for profiling real memory usage in applications.
import sys

def get_object_size_mb(obj):
"""
Returns the size of a Python object in megabytes (MB).
"""
size_bytes = sys.getsizeof(obj)
size_mb = size_bytes / (1024 * 1024) # Convert bytes to MB
return size_mb

# create huge list
a = [i for i in range(5_000_000)]
size_of_list_mb = get_object_size_mb(a)
print(f"Size python object is: {size_of_list_mb:.2f} MB")

##OUTPUT
% /usr/bin/python3 getsizeof_sample.py Size python object is: 38.35 MB


Using tracemalloc:

A built-in standard library module (since Python 3.4) for tracing Python memory allocations.

  • What it measures: Tracks memory blocks allocated by Python, providing detailed statistics (size, count, traceback) grouped by line, file, or traceback. Take snapshots and compare them to find leaks or differences.
  • Pros:
    • No external dependencies.
    • Precise for Python-allocated memory (lower overhead than some tools).
    • Excellent for debugging leaks: shows exactly where allocations happen.
    • Supports peak/current traced memory and object-specific tracebacks.
  • Cons:
    • Only tracks Python allocations (misses C extensions like NumPy unless they use Python's allocator).
    • Requires starting tracing early (ideally at program start).
    • More low-level; needs code to take/compute snapshots.
  • When to use: Investigating memory leaks easily, finding allocation hotspots, or comparing memory before/after code sections. Great for detailed, programmatic analysis in production-like debugging.
import tracemalloc

tracemalloc.start()

# create huge list
a = [i for i in range(5_000_000)]

current, peak = tracemalloc.get_traced_memory()
print(f"Current memory: {current / 10**6:.2f} MB; Peak: {peak / 10**6:.2f} MB")

tracemalloc.stop()

##OUTPUT
% /usr/bin/python3 tracemalloc_sample.py
Current memory: 180.21 MB; Peak: 180.21 MB


Using memory_profiler:

A third-party package (pip install memory_profiler) for line-by-line memory profiling.

  • What it measures: Memory usage (typically Resident Set Size - is the portion of a process that is actually living in RAM right now) after each line in decorated functions, showing increments. Uses psutil as the default backend (optional switch to tracemalloc for more precise Python-only allocations)
  • Pros:
    • Easy line-by-line view (similar to line_profiler for time).
    • Works cross-platform (Windows/Linux/Mac).
    • Plots memory over time with mprof.
    • Backend flexible (can switch between psutil and tracemalloc).
  • Cons:
    • External dependency (and often needs psutil).
    • Higher overhead due to frequent sampling of process memory.
    • RSS is coarse and may overestimate memory (includes shared libs, OS allocations). 
    • Less precise than tracemalloc for pinpointing pure Python object allocations.
  • When to use: Quick, high-level overview of memory growth per line in functions/scripts. Ideal for exploratory profiling for scripts or functions where you only need to know which lines are causing memory jumps, not deep allocation traces.
from memory_profiler import memory_usage

def measure_ram(func):
def wrapper(*args, **kwargs):
mem_before = memory_usage()[0]
result = func(*args, **kwargs)
mem_after = memory_usage()[0]
print(f"RAM used: {mem_after - mem_before:.2f} MB")
return result
return wrapper

@measure_ram
def sample_func():
a = [i for i in range(5_000_000)]

sample_func()

##OUTPUT
% /usr/bin/python3 memory_profiler_sample.py
RAM used: 3.60 MB

Memory Usage Techniques


Effective Data Structure Usage

Choosing the right data structure is one of the most important ways to reduce RAM usage and improve performance. Different structures store data differently and have different memory footprints.

Array Vs List: Using a regular list to store many numbers is much less efficient in RAM than using an array object. More memory allocations have to occur, which each take time, calculations also occur on larger objects, which will be less cache friendly, and more RAM is used overall, so less RAM is available to other programs.

from memory_profiler_sample import measure_ram #Note: imported function is from above sample
from array import array

# Create many integers
N = 5_000_000

@measure_ram
def sample_list():
print("List Creation")
data_list = list(range(N))

sample_list()

@measure_ram
def sample_arr():
print("Array Creation")
data_array = array('i', range(N))

sample_arr()

##OUTPUT
% /usr/bin/python3 list_VS_arr_sample.py
List Creation
RAM used: 41.45 MB
Array Creation
RAM used: 19.16 MB

Tuple Vs List: Tuples use less memory as there is it cannot be resized and no extra capacity reserved.  Faster and more cache-friendly. Similar behaviour goes with set vs dict.

Choose the Smallest Numeric Type: When working with large datasets in Python, NumPy arrays are far more memory-efficient than pure Python lists or Pandas Series (which often default to higher-precision types). The key is to select the smallest dtype that can safely represent your data without overflow or loss of precision. Here are few options int64 uses 8 bytes,  int32 uses 4 bytes, int16 uses 2 bytes.
 
Use Sparse Data Structures: Use libraries like, scipy.sparse or pandas.SparseArray, when the data holds most values as zero, instead storing all zeros does not makes effective use of RAM. Hence in the sparse matrices store only non-zero entries.


Bytes versus Unicode

Python distinguishes between text (human-readable characters) and binary data (raw bytes - images, files, network packets, compressed data, encrypted data). Understanding the difference is essential for file handling, networking, APIs, and encoding issues. Unicode strings use more memory because it requires to store character metadata.  Bytes are compact as it uses only 1 byte per element.

import sys

def get_object_size_mb(obj):
"""
Returns the size of a Python object in megabytes (MB).
"""
size_bytes = sys.getsizeof(obj)
size_mb = size_bytes / (1024 * 1024) # Convert bytes to MB
return size_mb

# Create N string
N = 500_000_000

s = "hello 😊"

# Unicode string (each character may take multiple bytes)
def sample_text():
print("Text Creation")
data_text = s*N
print(f"Size is: {get_object_size_mb(data_text):.2f} MB")

sample_text()

# Bytes version (UTF-8 encoded)
def sample_bytes():
print("Bytes Creation")
data_text = s*N
data_bytes = data_text.encode("utf-8")
print(f"Size is: {get_object_size_mb(data_bytes):.2f} MB")

sample_bytes()



##OUTPUT
% /usr/bin/python3 text_VS_unicode_sample.py
Text Creation
Size is: 13351.44 MB
Bytes Creation
Size is: 4768.37 MB



Huge Text data store

Handling massive text data (logs, documents, transcripts, crawled data, chat history, etc.) requires storage formats and structures that are memory-efficient, compressed, and scalable.

Use Compression (gzip, bz2, zstd): Making use of compression helps in text compresses 10–30× smaller, as storing raw text wastes huge amounts of space.

mmap: Map a file (or part of it) directly into memory and treat gigabytes of data as if it were a giant string or byte array. Super fast for large files and low memory usage. Giving a bytes-like object that can be terabytes in size, but uses almost no RAM until you touch the pages, resulting no RAM explosion.

DAWG (Directed Acyclic Word Graph): A DAWG is a highly compressed data structure used to store large sets of strings efficiently. It is especially powerful for dictionaries, NLP vocabularies, autocomplete engines, and prefix-based searches. 

import dawg

words = ["cat", "car", "cart", "dog", "doing"]

d = dawg.CompletionDAWG(words)

print(d.keys("ca")) # prefix search: ['cat', 'car', 'cart']
print("dog" in d) # True

Tries: Refers to a tree-based data structure also known as a prefix tree or digital tree. It is used to efficiently store and retrieve a dynamic collection of strings by sharing common prefixes. Prefix-based string data structures play a crucial role in text indexing, NLP, search engines, auto-complete systems and large dictionary storage.
Below are few commonly used structures are:
  • Marisa-Trie (Matching Algorithm with Recursively Implemented StorAge)
    • Static compressed trie using Louds (Level-Order Unary Degree Sequence) representation. 
    • Extremely small memory footprint (often 5–10% of original text). 
    • Supports predictive search (common prefix search) and exact-match very efficiently. 
    • Once built, the structure is read-only. Hence not suitable for frequent insertions/deletions.
    • Python module: marisa-tri.
  • Double-Array Trie (DAT)
    • The Double-Array Trie (DAT) is a compact Trie representation using two arrays: 
      • BASE array
        • determines offset for child transitions
        • base[i]: starting index for children of node i
      • CHECK array
        • stores the parent index to validate transitions
        • check[i]: parent node of node i (for collision detection)
    • Very fast transitions (just array indexing). 
    • Supports dynamic insertion/deletion (though slower than lookup).
    • Python module: datrie
  • HAT-Trie (Hash-Array mapped Trie)
    • A HAT-Trie combines:  
      • A Trie for prefix organization
      • Hash buckets for fast leaf-node lookups
    • Starts with a hash table at the top levels, then “bursts” into small trie containers when buckets fill. 
    • Extremely cache-efficient because burst containers are small and contiguous. 
    • Supports insertion, deletion, prefix search, and can store values/frequencies.
    • Python module: hat-trie


General Tips

Avoiding Copies: Avoid making unnecessary copies of the objects, which will occupy RAM with redundant data. Instead use references, or views from NumPy or use iterators. 
Note: For strings Python reuses identical immutable strings, this avoids storing the same string many times.

Profiling: Like shown above use the profiling methods to check memory usage while programming. These will help to detect early issues, like excessive allocations, memory leaks, inefficient structures.

Use Generators Instead of Storing Everything in Memory: A generator yields one item at a time as there here we would not need to store entire dataset in RAM.

Numpy: While working with numeric data, opt in to use numpy arrays as they offers many fast algorithms. 

Bitarray: If you're code requires lots of bit strings, numpy and bitarray packages both provide efficient representation of bits packed into bytes.

Micro Python: Will be interesting project for working with embedded sytems, its a tiny memory footprint allows developers to write high-level, readable Python code while still running on hardware traditionally restricted to low-level C/C++. It brings much of the core Python 3 functionality to devices with extremely limited resources—often with as little as 16–256 KB of RAM and flash storage.

Unit Test Cases: In the aim of optimization one must not forget the fundamentals - the purpose of the code. Hence make sure to have unit test suite in place before you make algorithmic changes.


Conclusion

Just to recap, profiling helps you see the problem; effective structures and techniques help you fix it. Together, they form a solid foundation for writing memory-efficient, high-performance Python programs. The general tips practices help improve performance, reduce memory pressure, and make applications more scalable—especially when working with large datasets or resource-constrained environments.


References:

High Performance Python - O'Reilly Book
LLMs for making the topics presentable

No comments:

Post a Comment

Scarcity Brings Efficiency: Python RAM Optimization

  In today’s world, with the abundance of RAM available, we rarely think about optimizing our code. But sooner or later, we hit the limits a...