Bird
Raised Fist0
NumPydata~5 mins

Working with large files efficiently in NumPy

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Introduction

Large files can be too big to load all at once. Working efficiently helps save memory and time.

You have a huge dataset that does not fit into your computer's memory.
You want to process data in parts instead of loading everything at once.
You need to read or write large numerical data quickly.
You want to avoid your program crashing due to memory overload.
You want to speed up data analysis by handling data in chunks.
Syntax
NumPy
import numpy as np

# Load part of a large file using memory mapping
array = np.memmap('filename.dat', dtype='float32', mode='r', shape=(1000, 1000))

np.memmap lets you access small parts of a big file without loading it all.

You specify the data type, mode (read or write), and shape of the data.

Examples
This opens a large file as if it were an array, but only loads parts when needed.
NumPy
import numpy as np

# Memory-map a large binary file for reading
data = np.memmap('data.bin', dtype='float64', mode='r', shape=(5000, 5000))
This creates a big file and writes numbers from 0 to 9999 without loading all in memory.
NumPy
import numpy as np

# Create a new memory-mapped file for writing
mmap_array = np.memmap('newfile.dat', dtype='int32', mode='w+', shape=(10000,))
mmap_array[:] = np.arange(10000)
For text files like CSV, read in smaller parts and convert to numpy arrays for processing.
NumPy
import numpy as np
import pandas as pd

# Read a large CSV file in chunks using pandas and convert to numpy arrays
chunks = pd.read_csv('large.csv', chunksize=10000)
for chunk in chunks:
    arr = chunk.to_numpy()
    # process arr here
Sample Program

This program creates a large file with 1 million numbers from 0 to 1. Then it reads only 10 numbers from the file without loading all data into memory.

NumPy
import numpy as np

# Create a large memory-mapped file and write data
filename = 'large_data.dat'
size = 1000000  # 1 million elements

# Create file with zeros
mmap_array = np.memmap(filename, dtype='float32', mode='w+', shape=(size,))
mmap_array[:] = np.linspace(0, 1, size)

# Flush changes to disk
mmap_array.flush()

# Now read only a slice without loading entire file
mmap_read = np.memmap(filename, dtype='float32', mode='r', shape=(size,))
slice_data = mmap_read[100000:100010]

print(slice_data)
OutputSuccess
Important Notes

Memory mapping works best with binary files, not text files.

Always specify the correct data type and shape to avoid errors.

Use chunk reading for large text files like CSVs, then convert to numpy arrays.

Summary

Use np.memmap to work with large binary files without loading all data.

Read or write data in parts to save memory and speed up processing.

For large text files, read in chunks and convert to numpy arrays for analysis.

Practice

(1/5)
1. What is the main advantage of using np.memmap when working with large binary files?
easy
A. It allows accessing data on disk without loading the entire file into memory.
B. It automatically compresses the file to save disk space.
C. It converts binary files into text files for easier reading.
D. It loads the entire file into memory for faster processing.

Solution

  1. Step 1: Understand np.memmap functionality

    np.memmap creates a memory-map to an array stored in a binary file on disk, allowing access without loading all data into RAM.
  2. Step 2: Compare options with this behavior

    Only It allows accessing data on disk without loading the entire file into memory. correctly describes this behavior. Options B, C, and D describe unrelated or incorrect features.
  3. Final Answer:

    It allows accessing data on disk without loading the entire file into memory. -> Option A
  4. Quick Check:

    np.memmap = Access data on disk [OK]
Hint: Remember: memmap reads from disk, not full memory load [OK]
Common Mistakes:
  • Thinking memmap loads entire file into memory
  • Confusing memmap with file compression
  • Assuming memmap converts file formats
2. Which of the following is the correct syntax to create a memory-mapped array from a binary file named data.bin with dtype float32 and shape (1000, 1000)?
easy
A. np.memmap('data.bin', dtype='float64', mode='r', shape=(1000, 1000))
B. np.memmap('data.bin', dtype='int32', mode='w', shape=(1000, 1000))
C. np.memmap('data.bin', dtype='float32', mode='r+', shape=(1000, 1000))
D. np.memmap('data.bin', dtype='float32', mode='rw', shape=(1000, 1000))

Solution

  1. Step 1: Identify correct dtype and mode

    The question asks for dtype 'float32' and a mode that allows reading and writing, which is 'r+'.
  2. Step 2: Check each option

    np.memmap('data.bin', dtype='float32', mode='r+', shape=(1000, 1000)) matches dtype 'float32' and mode 'r+'. np.memmap('data.bin', dtype='int32', mode='w', shape=(1000, 1000)) has wrong dtype 'int32' and mode 'w' (write only). np.memmap('data.bin', dtype='float64', mode='r', shape=(1000, 1000)) has wrong dtype 'float64' and mode 'r' (read only). np.memmap('data.bin', dtype='float32', mode='rw', shape=(1000, 1000)) uses invalid mode 'rw'.
  3. Final Answer:

    np.memmap('data.bin', dtype='float32', mode='r+', shape=(1000, 1000)) -> Option C
  4. Quick Check:

    Correct dtype and mode = np.memmap('data.bin', dtype='float32', mode='r+', shape=(1000, 1000)) [OK]
Hint: Use mode 'r+' for read/write memmap [OK]
Common Mistakes:
  • Using wrong dtype for the file data
  • Using invalid mode like 'rw'
  • Confusing read-only 'r' with read/write 'r+'
3. Consider the following code snippet:
import numpy as np
filename = 'largefile.dat'
# Create memmap
mmap = np.memmap(filename, dtype='int32', mode='r', shape=(4, 4))
print(mmap[2, 3])

If the file contains a 4x4 array with values from 0 to 15 in row-major order, what will be the output?
medium
A. 12
B. 14
C. 15
D. 11

Solution

  1. Step 1: Understand data layout

    The file stores values 0 to 15 in a 4x4 array in row-major order: [[0,1,2,3],[4,5,6,7],[8,9,10,11],[12,13,14,15]]
  2. Step 2: Find value at position (2, 3)

    Row 2 (0-based) is [8,9,10,11]. Index 3 in this row is 11.
  3. Final Answer:

    11 -> Option D
  4. Quick Check:

    Value at (2,3) = 11 [OK]
Hint: Remember zero-based indexing for arrays [OK]
Common Mistakes:
  • Confusing row and column indices
  • Using 1-based indexing instead of 0-based
  • Mixing up row-major and column-major order
4. You try to create a memmap with this code:
mmap = np.memmap('data.bin', dtype='float32', mode='r+', shape=(1000, 1000))

but get an error: ValueError: cannot mmap an empty file. What is the likely cause and how to fix it?
medium
A. The dtype 'float32' is invalid; use 'float64' instead.
B. The file 'data.bin' is empty; initialize it with correct size before memmap.
C. The mode 'r+' is read-only; use 'w+' to write.
D. The shape (1000, 1000) is too large; reduce it to (100, 100).

Solution

  1. Step 1: Understand error cause

    The error means the file exists but has zero bytes, so memmap cannot map it with the given shape and dtype.
  2. Step 2: Fix by initializing file size

    To fix, create or resize the file to hold the required data (1000*1000*4 bytes for float32) before memmap.
  3. Final Answer:

    The file 'data.bin' is empty; initialize it with correct size before memmap. -> Option B
  4. Quick Check:

    Empty file causes mmap error [OK]
Hint: Ensure file size matches array size before memmap [OK]
Common Mistakes:
  • Changing dtype without fixing file size
  • Using wrong mode without file content
  • Reducing shape without reason
5. You have a very large text file with 1 billion numbers separated by spaces. You want to analyze the data using numpy but cannot load all at once. Which approach is best to process this file efficiently?
hard
A. Read the file in chunks, convert each chunk to numpy arrays, and process incrementally.
B. Use np.memmap directly on the text file to access numbers.
C. Read the entire file into memory as a string, then convert to numpy array.
D. Convert the text file to CSV and load with pandas without chunking.

Solution

  1. Step 1: Understand file type and memory limits

    The file is a large text file, not binary. np.memmap works only with binary files, so Use np.memmap directly on the text file to access numbers. is invalid.
  2. Step 2: Choose efficient reading method

    Reading entire file at once (Read the entire file into memory as a string, then convert to numpy array.) is memory-heavy. Converting to CSV and loading without chunking (Convert the text file to CSV and load with pandas without chunking.) also risks memory overload. Reading in chunks and processing incrementally (Read the file in chunks, convert each chunk to numpy arrays, and process incrementally.) is memory efficient and practical.
  3. Final Answer:

    Read the file in chunks, convert each chunk to numpy arrays, and process incrementally. -> Option A
  4. Quick Check:

    Chunk reading for large text files = Read the file in chunks, convert each chunk to numpy arrays, and process incrementally. [OK]
Hint: Process large text files in chunks, not all at once [OK]
Common Mistakes:
  • Trying to memmap text files
  • Loading entire large file into memory
  • Ignoring memory limits when converting formats