Story Opening

The out-of-memory crash was embarrassing, mostly because Arjun knew better. In Java he would never Files.readAllLines() a 40 GB file; he’d use Files.lines() and a lazy Stream, or a BufferedReader in a try-with-resources block.

He just hadn’t known what the Python equivalents were. Priya showed him in four lines:

from pathlib import Path
Path("txns.csv").write_text("id,amount\nT1,120.0\nT2,45.5\nT3,3000.0\n") # tiny sample
with open("txns.csv", encoding="utf-8") as f: # try-with-resources
next(f) # skip the header line
total = sum(float(line.split(",")[1]) for line in f) # streams line by line
print(total) # -> 3165.5

A file object is an iterator that yields one line at a time. The generator expression pulls lines through lazily. At no point does more than one line live in memory. That one pattern — lazy iteration — runs through the entire Python data stack, from file reading to PyTorch’s DataLoader.


Java → Python: The Quick Map

JavaPython
Iterable<T> / Iterator<T>Iterable (__iter__) / iterator (__next__)
hasNext() + next()next() until StopIteration is raised
Lazy Stream pipelineChained generators / generator expressions
Custom SpliteratorA generator function with yield
Stream.limit, concat, iterateitertools.islice, chain, count
try-with-resources / AutoCloseablewith / context managers (__enter__, __exit__)
Checked exceptionsNone — all exceptions are unchecked
Files.lines(path)open(path) — the file object iterates lines
java.nio.file.Pathpathlib.Path

Deep Dive: The Iterator Protocol

Every for loop in Python runs on two tiny methods:

  • An iterable has __iter__(), which returns an iterator. Lists, dicts, strings, files, ranges are iterables.
  • An iterator has __next__(), which returns the next item or raises StopIteration when exhausted. (Iterators also have __iter__ returning themselves.)

for x in xs: is shorthand for this:

amounts = [120.0, 45.5, 3000.0]
it = iter(amounts) # calls amounts.__iter__()
while True:
try:
x = next(it) # calls it.__next__()
except StopIteration: # end of data is signalled by an exception, not hasNext()
break
print(x)
# next() accepts a default instead of raising:
print(next(iter([]), "empty")) # -> empty

Iterables vs iterators: the one-shot trap

A list can be iterated many times — each for asks for a fresh iterator. An iterator (a generator, a file, a map object, a zip) can be consumed once. Java Streams behave the same way, but Java at least throws IllegalStateException on reuse. Python silently gives you nothing:

squares = (x * x for x in range(4)) # generator: an ITERATOR
print(sum(squares)) # -> 14
print(sum(squares)) # -> 0 (already exhausted — no error!)
# Same trap with map/zip/filter objects:
pairs = zip(["a", "b"], [1, 2])
print(list(pairs)) # -> [('a', 1), ('b', 2)]
print(list(pairs)) # -> []
# If you need multiple passes, materialise once:
data = list(x * x for x in range(4))
print(sum(data), max(data)) # -> 14 9

Gotcha — This bites hardest in functions that iterate their argument twice (e.g. compute a mean, then a variance). Pass a list and it works; pass a generator and the second pass sees nothing, producing a silently wrong answer. Either document that you need a sequence, or call list() at the top.


Generators: Iterators You Write Like Functions

Writing an iterator class by hand (__iter__ + __next__ + state fields) is tedious — exactly like implementing Iterator<T> in Java. A generator function does it for you: any function containing yield returns a generator object when called. Each next() runs the body until the next yield, then pauses with all local state preserved.

def countdown(n):
print("starting") # runs on the FIRST next(), not when countdown() is called
while n > 0:
yield n # hand back a value and pause here
n -= 1
print("done") # runs when the loop ends, just before StopIteration
gen = countdown(3) # nothing printed yet: the body hasn't started
print(type(gen).__name__) # -> generator
print(next(gen)) # prints "starting", then the value 3
print(list(gen)) # prints "done" after consuming 2 and 1 -> [2, 1]

Why it matters: memory

import sys
eager = [i * 2 for i in range(1_000_000)] # one million ints materialised
lazy = (i * 2 for i in range(1_000_000)) # a recipe for producing them
print(sys.getsizeof(eager) > 8_000_000) # -> True (~8 MB just for the pointers)
print(sys.getsizeof(lazy) < 500) # -> True (a couple of hundred bytes, regardless of size)
print(sum(lazy) == sum(eager)) # -> True

Infinite sequences

Generators can be infinite, because they only compute what’s asked for:

from itertools import islice
def transaction_ids(prefix="TXN"):
n = 1
while True: # never ends — that's fine
yield f"{prefix}-{n:06d}"
n += 1
print(list(islice(transaction_ids(), 3))) # -> ['TXN-000001', 'TXN-000002', 'TXN-000003']

Deep Dive: Generator Pipelines

The real power appears when you chain generators into stages — each stage pulls items from the previous one, one at a time. It is a Java Stream pipeline built from plain functions, and it processes a 40 GB file with constant memory.

import csv
from pathlib import Path
from collections import defaultdict
# --- create a small sample file so the example is runnable ------------------
Path("txns.csv").write_text(
"txn_id,merchant,amount,country\n"
"T1,AcmeMart,120.0,IN\n"
"T2,ZipFuel,not_a_number,IN\n"
"T3,AcmeMart,15000.0,US\n"
"T4,BookNook,80.0,IN\n"
"T5,ZipFuel,22000.0,SG\n",
encoding="utf-8",
)
# --- Stage 1: source. Yields dict rows lazily. ------------------------------
def read_rows(path):
with open(path, newline="", encoding="utf-8") as f: # file stays open while iterating
yield from csv.DictReader(f) # delegate to another iterator
# --- Stage 2: parse + clean. Drops bad rows instead of crashing. ------------
def parse(rows):
for row in rows:
try:
row["amount"] = float(row["amount"])
except ValueError:
print(f"skipping bad row {row['txn_id']}")
continue
yield row
# --- Stage 3: filter. -------------------------------------------------------
def foreign_only(rows, home="IN"):
return (r for r in rows if r["country"] != home) # generator expression stage
# --- Stage 4: terminal operation (the 'collect'). ---------------------------
pipeline = foreign_only(parse(read_rows("txns.csv"))) # NOTHING has executed yet
totals = defaultdict(float)
for row in pipeline: # pulling drives every stage
totals[row["merchant"]] += row["amount"]
print(dict(totals))
# skipping bad row T2
# {'AcmeMart': 15000.0, 'ZipFuel': 22000.0}
graph LR F[(txns.csv)] -->|line| R[read_rows] R -->|dict| P[parse] P -->|clean dict| FO[foreign_only] FO -->|foreign row| T[for loop
terminal] T -.->|next| FO FO -.->|next| P P -.->|next| R

Each next() request travels up the chain; each item travels down. Exactly one row is in flight at a time.

Tip — yield from iterable delegates to a sub-iterator: it yields every item of iterable in turn. It replaces for x in iterable: yield x and is how you compose generators.


itertools: The Stream Operators You Were Missing

from itertools import islice, chain, batched, accumulate, takewhile, groupby, pairwise
amounts = [120, 45, 3000, 80, 15000, 60]
print(list(islice(amounts, 2, 5))) # -> [3000, 80, 15000] (skip 2, limit 3)
print(list(chain([1, 2], (3, 4), range(5, 7)))) # -> [1, 2, 3, 4, 5, 6] (concat)
print(list(batched(amounts, 4))) # -> [(120, 45, 3000, 80), (15000, 60)] (3.12+)
print(list(accumulate(amounts))) # -> [120, 165, 3165, 3245, 18245, 18305] (running total)
print(list(takewhile(lambda a: a < 1000, amounts))) # -> [120, 45]
print(list(pairwise([10, 15, 12]))) # -> [(10, 15), (15, 12)] (sliding pairs)
# groupby groups CONSECUTIVE keys — sort first for SQL-like grouping.
events = sorted([("IN", 10), ("US", 5), ("IN", 7)], key=lambda e: e[0])
print({k: sum(v for _, v in grp) for k, grp in groupby(events, key=lambda e: e[0])}) # -> {'IN': 17, 'US': 5}

Tip — batched is your micro-batcher. Sending rows to a model or an API in chunks of 512 is one line: for chunk in batched(rows, 512): score(chunk). On Python < 3.12, use islice in a loop.


Exceptions: Unchecked, Hierarchical, and Used for Control Flow

Python has no checked exceptions — every exception behaves like a RuntimeException. The structure is familiar, with two additions: else and exception chaining with from.

def parse_amount(raw: str) -> float:
try:
value = float(raw)
except ValueError as e:
# 'raise ... from e' chains the cause (like new X(msg, cause) in Java).
raise InvalidTransaction(f"bad amount {raw!r}") from e
else:
# runs only if the try block raised NOTHING — keeps the try block minimal
if value < 0:
raise InvalidTransaction("negative amount")
return value
finally:
pass # always runs: cleanup goes here (but prefer 'with' — see below)
class SentinelError(Exception): # project base exception
"""Base class for all Sentinel errors."""
class InvalidTransaction(SentinelError): # specific subtype
pass
for raw in ["120.5", "abc", "-3"]:
try:
print(parse_amount(raw))
except InvalidTransaction as e:
cause = type(e.__cause__).__name__ if e.__cause__ else None
print(f"rejected: {e} (cause: {cause})")
# 120.5
# rejected: bad amount 'abc' (cause: ValueError)
# rejected: negative amount (cause: None)

EAFP vs LBYL

Java culture is mostly LBYL — Look Before You Leap: check map.containsKey(k) then map.get(k). Python culture prefers EAFP — Easier to Ask Forgiveness than Permission: just try it and handle the exception. Exceptions are cheap enough in Python, and EAFP avoids race conditions between the check and the action.

config = {"threshold": "0.8"}
# LBYL (works, but two lookups and a check-then-act gap)
if "threshold" in config:
threshold = float(config["threshold"])
# EAFP (idiomatic)
try:
timeout = float(config["timeout"])
except KeyError:
timeout = 30.0
print(threshold, timeout) # -> 0.8 30.0

Gotcha — never write a bare except:. It catches everything, including KeyboardInterrupt (Ctrl-C) and SystemExit. Even except Exception: should be rare and should log the error. Catch the narrowest exception that you can actually handle.

Tip — Common built-in exceptions map neatly: ValueError ≈ IllegalArgumentException, TypeError ≈ ClassCastException, KeyError/IndexError ≈ NoSuchElementException/IndexOutOfBoundsException, AttributeError ≈ NullPointerException (usually you called a method on None), NotImplementedError ≈ UnsupportedOperationException.


Context Managers: with Is try-with-resources

with guarantees cleanup. Any object with __enter__ and __exit__ works, exactly as any AutoCloseable works in try-with-resources.

from pathlib import Path
path = Path("scores.txt")
# File automatically closed when the block exits — even if an exception is raised.
with path.open("w", encoding="utf-8") as f:
f.write("T1,0.91\nT2,0.12\n")
print(f.closed) # -> True
# Multiple resources in one statement (parenthesised form, 3.10+):
with (
open("scores.txt", encoding="utf-8") as src,
open("high.txt", "w", encoding="utf-8") as dst,
):
for line in src:
if float(line.split(",")[1]) > 0.5:
dst.write(line)
print(Path("high.txt").read_text().strip()) # -> T1,0.91

Writing your own: the class way and the generator way

import time
from contextlib import contextmanager
class Timer:
"""Class-based context manager: __enter__ / __exit__."""
def __enter__(self):
self.start = time.perf_counter()
return self # bound to the 'as' target
def __exit__(self, exc_type, exc, tb):
self.elapsed = time.perf_counter() - self.start
return False # False = don't swallow exceptions
@contextmanager
def model_mode(state: dict, mode: str):
"""Generator-based: code before 'yield' is __enter__, after it is __exit__."""
previous = state["mode"]
state["mode"] = mode
try:
yield state # the body of the 'with' block runs here
finally:
state["mode"] = previous # restored even if the block raised
with Timer() as t:
sum(range(100_000))
print(t.elapsed > 0) # -> True
model = {"mode": "train"}
with model_mode(model, "eval"):
print(model["mode"]) # -> eval
print(model["mode"]) # -> train

That second pattern — temporarily switch a mode and guarantee it’s restored — is exactly what torch.no_grad() does in Part 10.

from contextlib import suppress
# suppress: "ignore this specific exception" — cleaner than try/except/pass.
cache = {}
with suppress(KeyError):
del cache["missing"]
print("still running") # -> still running

Files the Modern Way: pathlib, csv, json

import csv
import json
from pathlib import Path
data_dir = Path("data") / "raw" # '/' joins paths — OS-independent
data_dir.mkdir(parents=True, exist_ok=True)
# JSON: dicts and lists map directly to objects and arrays.
config = {"model": "logreg", "threshold": 0.8, "features": ["amount", "hour"]}
cfg_path = data_dir / "config.json"
cfg_path.write_text(json.dumps(config, indent=2), encoding="utf-8")
loaded = json.loads(cfg_path.read_text(encoding="utf-8"))
print(loaded["features"]) # -> ['amount', 'hour']
# CSV writing with DictWriter (newline="" avoids blank lines on Windows).
rows = [{"txn_id": "T1", "amount": 120.0}, {"txn_id": "T2", "amount": 45.5}]
csv_path = data_dir / "txns.csv"
with csv_path.open("w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["txn_id", "amount"])
writer.writeheader()
writer.writerows(rows)
print(sorted(p.name for p in data_dir.glob("*.*"))) # -> ['config.json', 'txns.csv']
print(csv_path.suffix, csv_path.stem) # -> .csv txns

Tip — always pass encoding="utf-8". Python’s default file encoding is platform-dependent (historically cp1252 on Windows), which breaks on merchant names like “Café”. Python 3.15 makes UTF-8 the default everywhere; until you’re on it, be explicit.

Tip — For real tabular data you’ll use pandas.read_csv(path, chunksize=100_000), which gives you an iterator of DataFrames — the same lazy idea, vectorised (Part 8). Plain csv remains useful for quick scripts and for streaming transforms.


Tips, Tricks & Gotchas

Tip — enumerate, zip, map, filter, reversed, dict.items() are lazy too. They return iterators or views, not lists. Wrap them in list() when you need to print or reuse them.

Gotcha — generators delay errors. A bug inside a generator doesn’t surface when you create the pipeline, only when something consumes it. If a pipeline “does nothing”, check that something actually iterates it.

Gotcha — closing over files in generators. If you return a generator from inside a with open(...) block in a regular function, the file closes when the function returns, and iteration fails with ValueError: I/O operation on closed file. Put the with inside the generator function, as read_rows does above.

Tip — sum, min, max, any, all, sorted, "".join, dict(), set() all accept any iterable, so you can feed them a generator expression directly without building a list.


Key Takeaways

ConceptRemember
Iterator protocoliter() gives an iterator; next() until StopIteration
One-shot iteratorsGenerators, files, zip, map exhaust silently — materialise if reused
Generatorsyield pauses the function; constant memory; can be infinite
PipelinesChain generator stages like Stream operations; yield from to delegate
itertoolsislice, chain, batched, accumulate, groupby (sort first)
ExceptionsAll unchecked; raise ... from for causes; else for the success path
EAFPTry and handle, rather than check then act
Context managerswith = try-with-resources; write with a class or @contextmanager

Story Closing

The rewritten loader streamed the 40 GB export through a four-stage generator pipeline in constant memory, logging and skipping the 0.3% of rows with corrupt amounts. Arjun’s laptop fan barely noticed.

It wasn’t fast, though. Forty minutes for a single pass. And when Arjun opened the shared features.py module to plug his loader in, he hit a different kind of problem: a function called build(data, cfg, mode=None) with no documentation and no types. What was data? A list? A DataFrame? What keys did cfg need?

He missed his compiler.

In Part 6, Arjun brings types back — type hints, mypy, Pydantic — and sets up the tooling that makes Python feel as safe as Maven and JUnit.


This is Part 5 of a 10-part series: “Python for Java Developers: From Streams to Tensors.”