What’s one thing you learned? What’s still confusing?
Decorators: Functions That Wrap Functions
Every @app.get(...) in FastAPI and every @pytest.fixture is a decorator: a function that takes a function and hands back a better one.
Python's Memory Model: Objects, References & Identity
The #1 source of "why is my function mutating my list?!" bugs is misunderstanding Python's reference model.
How CPython Works: Bytecode, Objects, the Cycle Collector and the GIL
Read the bytecode CPython runs with dis, explain what a Python object is, find reference cycles, and say what the GIL and the free-threaded build change.
Interactive Labs for This Track
Loop Visualizer
You're a factory robot repeating the same task on an assembly line — watch how loops automate repetitive work
List Slicing
You have a playlist of 50 songs — grab just tracks 10 through 20 with a single slice expression
Sorting Algorithms
You're organizing a library of 10,000 books — which sorting method is fastest?
Ask questions, share insights
Dataset class to hold your training rows. Then len(ds) fails with a TypeError, ds[0] fails, a for loop over it fails, and print(ds) shows a memory address. A plain list does all of this without being asked.__getitem__ and __len__, truthiness, __iter__, __call__, and the rule that gives a method its self. It ends with a short pointer to with. The NumPy and pandas lesson that follows is about objects built exactly this way.print(x) calls __repr__, x + y calls __add__, x == y calls __eq__, x < y calls __lt__ and hash(x) calls __hash__, always on the class and never on one instance. An operator method returns NotImplemented when it cannot handle its operand, and a class that defines __eq__ defines __hash__ from the same fields. A reflected method such as __radd__ handles 0 + x, which is how sum() starts. With that much in hand you can read this lesson on its own, and every Vector or Dataset below is written out in full where it is used.__getitem__ and __len__Dataset that holds training rows should behave like a list of them. Here it is with none of the methods, and what Python says about each attempt.class Dataset:
def __init__(self, samples):
self.samples = list(samples)
ds = Dataset([((3, 4), 0), ((1, 2), 1)])
print(ds) # <__main__.Dataset object at 0x...>
for label, fn in [("len(ds)", lambda: len(ds)),
("ds[0]", lambda: ds[0]),
("iterate", lambda: list(ds))]:
try:
fn()
except TypeError as err:
print(f"{label}: TypeError: {err}")len(ds): TypeError: object of type 'Dataset' has no len(), then 'Dataset' object is not subscriptable, then 'Dataset' object is not iterable. Each message names the capability that is missing. The rest of the lesson supplies them one at a time.__getitem__, which receives whatever you put between the square brackets. Because it delegates to a list, negative indexes and slices come for free.class Dataset:
def __init__(self, samples):
self.samples = list(samples)
def __getitem__(self, index):
return self.samples[index]
rows = [((5.1, 3.5), 0), ((4.9, 3.0), 0), ((6.3, 3.3), 1)]
ds = Dataset(rows)
print(ds[0]) # ((5.1, 3.5), 0)
print(ds[-1]) # ((6.3, 3.3), 1)
print(ds[0:2]) # [((5.1, 3.5), 0), ((4.9, 3.0), 0)]
for features, label in ds:
print(label, features)
print(list(ds) == rows) # True
print(((4.9, 3.0), 0) in ds) # True
try:
len(ds)
except TypeError as err:
print("TypeError:", err) # object of type 'Dataset' has no len()list(ds) and in all worked, though the class never defined __iter__ or __contains__. When Python needs to loop over an object that has no __iter__, it falls back to calling __getitem__ with 0, 1, 2 and so on until an IndexError is raised. The same fallback powers in, which loops and compares each item with ==. Only len() is still missing, since nothing can be inferred about the length.Dataset defines only __init__ and __getitem__ (delegating to a list). What does for x in ds: do?
You can watch the fallback happen with a class that announces each call.
class Spy:
def __getitem__(self, i):
print("asked for", i)
if i >= 3:
raise IndexError
return i * 10
print(list(Spy()))asked for 0 through asked for 3, then [0, 10, 20]. Index 3 raised IndexError, so the loop stopped without ever using that value.__len__ and make slicing return a Dataset rather than a bare list, so that ds[:2] is still something you can call len() on. A __repr__ makes the printout readable.class Dataset:
def __init__(self, samples):
self.samples = list(samples)
def __repr__(self):
return f"Dataset({len(self)} samples)"
def __len__(self):
return len(self.samples)
def __getitem__(self, index):
if isinstance(index, slice):
return Dataset(self.samples[index])
return self.samples[index]
ds = Dataset([((5.1, 3.5), 0), ((4.9, 3.0), 0), ((6.3, 3.3), 1)])
print(ds) # Dataset(3 samples)
print(len(ds)) # 3
print(ds[:2]) # Dataset(2 samples)
print(ds[-1]) # ((6.3, 3.3), 1)len() insists on a non-negative integer, so __len__ must return one. Returning a float raises TypeError.__bool__ and __len__if x:, while x:, not x, and and or all need to decide whether an object counts as true. Python asks for x.__bool__() first. If the class has no __bool__, it asks for x.__len__() and treats 0 as false. If neither exists, every instance is true.__len__, an empty dataset looks like a full one, and a guard such as if not ds: raise ValueError("no data") never fires.class NoLen:
pass
class WithLen:
def __init__(self, items):
self.items = list(items)
def __len__(self):
return len(self.items)
print(bool(NoLen())) # True
print(bool(WithLen([]))) # False
print(bool(WithLen([1]))) # True
class Vector:
def __init__(self, *values):
self.values = tuple(values)
def __bool__(self):
return any(self.values)
print(bool(Vector(0, 0))) # False
print(bool(Vector(0, 1))) # TrueDataset its truthiness the moment you wrote __len__. __bool__ is for the cases where "empty" does not mean "length zero". Here a vector of all zeros is false, like 0 and 0.0. Use it only when that matches what a reader would expect, since surprising truthiness hides bugs.A class defines __len__ returning 0 and nothing else. What does if obj: do?
__iter____getitem__ fallback works, but __iter__ says what you mean and is the one every for loop looks for first. It must return an iterator: an object with a __next__ method that eventually raises StopIteration. You rarely write one by hand. Return the iterator of an underlying list, or write __iter__ as a generator function, which the earlier lesson on generators covered.class Dataset:
def __init__(self, samples):
self.samples = list(samples)
def __len__(self):
return len(self.samples)
def __getitem__(self, index):
return self.samples[index]
def __iter__(self):
return iter(self.samples) # a fresh iterator every time
def batches(self, size):
for start in range(0, len(self), size):
yield self.samples[start:start + size]
ds = Dataset([((5.1, 3.5), 0), ((4.9, 3.0), 0), ((6.3, 3.3), 1), ((5.8, 2.7), 1)])
print(sum(label for _, label in ds)) # 2
print(len(list(ds)), len(list(ds))) # 4 4
print([len(batch) for batch in ds.batches(3)]) # [3, 1]__iter__ returns a new iterator each call, the dataset can be looped over every epoch. That is the difference between a collection and a one-shot stream. The batches method is an ordinary generator method. It is how a training loop pulls mini-batches out of the dataset.torch.utils.data.Dataset: define __len__ and __getitem__, and a DataLoader can batch and shuffle it. The loop below is a plain-Python version of what a DataLoader does, using only those two methods. It never calls __iter__. It builds a list of indexes with len(dataset), optionally shuffles it, and asks __getitem__ for each one.import random
class Dataset:
def __init__(self, samples):
self.samples = list(samples)
def __len__(self):
return len(self.samples)
def __getitem__(self, index):
return self.samples[index]
def loader(dataset, batch_size, shuffle=False, seed=0):
order = list(range(len(dataset))) # uses __len__
if shuffle:
random.Random(seed).shuffle(order)
for start in range(0, len(order), batch_size):
yield [dataset[i] for i in order[start:start + batch_size]] # uses __getitem__
ds = Dataset([(x, x * x) for x in range(5)])
for batch in loader(ds, 2):
print(batch) # [(0, 0), (1, 1)], then [(2, 4), (3, 9)], then [(4, 16)]
seen = [pair for batch in loader(ds, 2, shuffle=True) for pair in batch]
print(len(seen), sorted(seen) == ds.samples) # 5 Trueshuffle=True the order inside the batches changes, but every sample still appears exactly once. In a real training script you would write DataLoader(ds, batch_size=2, shuffle=True) and the dataset class would stay the same.__call__A trained model is an object that holds weights and is used like a function. Written as a plain class, the call fails.
class LinearModel:
def __init__(self, weights, bias):
self.weights = weights
self.bias = bias
model = LinearModel([0.5, -1.0], 2.0)
try:
model([3.0, 4.0])
except TypeError as err:
print("TypeError:", err) # 'LinearModel' object is not callable__call__ fixes it. The weights now live on the instance, between calls, and the call site still reads model(x).class LinearModel:
def __init__(self, weights, bias):
self.weights = weights
self.bias = bias
def __call__(self, features):
return sum(w * x for w, x in zip(self.weights, features)) + self.bias
def __repr__(self):
return f"LinearModel(weights={self.weights}, bias={self.bias})"
rows = [((3.0, 4.0), 1), ((1.0, 1.0), 0)]
model = LinearModel([0.5, -1.0], 2.0)
print(model) # LinearModel(weights=[0.5, -1.0], bias=2.0)
print(model((3.0, 4.0))) # -0.5
print([round(model(x), 2) for x, _ in rows]) # [-0.5, 1.5]
print(callable(model), callable(len), callable(5)) # True True Falsenn.Module defines __call__, so model(batch) runs your forward method, wrapped with the hooks the framework needs. Anything with __call__ also works wherever Python wants a function: map, sorted(key=...), a callback. A function cannot remember anything between calls unless you use a closure. An object with __call__ remembers whatever you store on it.__get__ is called a descriptor. When you read an attribute through an instance, Python finds the name on the class and, if the object has __get__, calls __get__(instance, owner) and uses what it returns instead of the object itself. Reading through the class passes None as the instance.class Loud:
def __get__(self, obj, objtype=None):
print("__get__ called with", type(obj).__name__, objtype.__name__)
return 42
class Box:
size = Loud() # stored on the class, not on any instance
b = Box()
print(b.size) # __get__ called with Box Box, then 42
print(Box.size) # __get__ called with NoneType Box, then 42def inside a class body makes an ordinary function, and that function's __get__ is what turns m.predict into a bound method with m already filled in as self. There is no special self syntax. It is __get__ at work.class Model:
def predict(self, x):
return x * 2
m = Model()
plain = Model.__dict__["predict"] # the raw function, nothing bound yet
bound = plain.__get__(m, Model) # what m.predict does behind the scenes
print(type(plain).__name__, type(bound).__name__) # function method
print(bound.__self__ is m, bound(3)) # True 6@CountCalls replaces the function with an instance of a class that has __call__ but no __get__. An instance of your own class is not a function, so reading Model().predict returns the CountCalls object unchanged, no self is supplied, and 3 lands in the self slot.from functools import update_wrapper
class CountCalls:
def __init__(self, func):
update_wrapper(self, func)
self.func = func
self.calls = 0
def __call__(self, *args, **kwargs):
self.calls += 1
return self.func(*args, **kwargs)
class Model:
@CountCalls
def predict(self, x):
return x * 2
try:
Model().predict(3)
except TypeError as err:
print("TypeError:", err) # TypeError: Model.predict() missing 1 required positional argument: 'x'CountCalls its own __get__. Reached through an instance, it returns MethodType(self, obj), a bound method that calls the CountCalls object with obj as the first argument. Reached through the class, obj is None and it returns itself, so Model.predict.calls still works.from functools import update_wrapper
from types import MethodType
class CountCalls:
def __init__(self, func):
update_wrapper(self, func)
self.func = func
self.calls = 0
def __call__(self, *args, **kwargs):
self.calls += 1
return self.func(*args, **kwargs)
def __get__(self, obj, objtype=None):
if obj is None: # reached through the class
return self
return MethodType(self, obj) # reached through an instance: bind it
class Model:
@CountCalls
def predict(self, x):
return x * 2
print(Model().predict(3), Model.predict.calls) # 6 1Model. A function decorator that returns a plain function needs none of this, because functions already have __get__.with statement: a pointerwith follows the same pattern, with two methods instead of one. __enter__ runs when the block starts and its return value is what as binds. __exit__ runs when the block ends, whether normally or through an exception.class Stage:
def __init__(self, name):
self.name = name
def __enter__(self):
print("start", self.name)
return self
def __exit__(self, exc_type, exc, tb):
print("end", self.name, "failed" if exc_type else "ok")
return False # let any exception continue upward
try:
with Stage("train"):
raise RuntimeError("out of memory")
except RuntimeError as err:
print("caught:", err)start train and end train failed before the error reaches your except. Without these two methods, Python says 'Stage' object does not support the context manager protocol (a TypeError since Python 3.11, and Python 3.14 adds a short hint in brackets to the message). The lesson on generators and context managers builds the same thing with @contextmanager, and the lesson on advanced OOP has the class form in full.Vector, a Dataset with __len__, __getitem__ and __repr__, and a few lines of output. Work through the challenges in order and run after each one.Tests · After the challenges, a + b has 3 samples, a[0:1] is a Dataset, the membership test is True, and sum works on a list of datasets.
| Syntax | Method | Notes |
|---|---|---|
len(x) | __len__ | Must return a non-negative integer. |
x[i], x[a:b] | __getitem__ | Gives iteration and in through the sequence fallback. |
for item in x | __iter__ | Return a fresh iterator each time for a reusable collection. |
item in x |
repr, print, +, *, ==, hash and < are in Part 1. Assignment follows the same pattern as the rows above: x[i] = v calls __setitem__ and del x[i] calls __delitem__. NumPy arrays and pandas objects define these methods too, which is why len(arr), arr[0] and arr[1:3] behave as they do on a list. One difference to remember: bool(arr) on a NumPy array with more than one element raises an error because the answer is ambiguous, unlike the __len__ rule used here.What does this print? class Bag: def __init__(self, items): self.items = items def __len__(self): return len(self.items) print(bool(Bag([])), bool(Bag([0])))
Next: NumPy and pandas, whose arrays and data frames are built from exactly these methods.
__contains__Falls back to iteration and == when missing. |
bool(x), if x: | __bool__, then __len__ | Otherwise always true. |
f(x) | __call__ | State on the instance, call syntax outside. |
with x as v: | __enter__, __exit__ | Cleanup that runs on errors too. |
m.predict | function __get__ | Fills in self. A decorator class used on a method needs its own __get__. |