What’s one thing you learned? What’s still confusing?
NumPy: Arrays for ML Data
Build and reshape NumPy arrays, centre a feature matrix with broadcasting, select rows with boolean masks, and draw repeatable random numbers.
Pandas Essentials: Tables for ML Data
Read a CSV or Parquet file into a pandas DataFrame, select and filter rows, handle missing values, and summarise groups with groupby.
Pandas Advanced: merge, pivot, apply, and Time Series
The difference between a junior and a senior data engineer is .merge() semantics — knowing why your row count exploded after a "harmless" join.
Interactive Labs for This Track
Loop Visualizer
You're a factory robot repeating the same task on an assembly line — watch how loops automate repetitive work
List Slicing
You have a playlist of 50 songs — grab just tracks 10 through 20 with a single slice expression
Sorting Algorithms
You're organizing a library of 10,000 books — which sorting method is fastest?
Ask questions, share insights
This lesson continues from Regular Expressions: Match, Capture and Parse. It applies the same tools to two jobs, cleaning raw text for an NLP pipeline (a chain of steps that prepares text for a language model) and validating and extracting fields, then adds lookahead and the reason some patterns are dangerously slow. The next lesson, on modules and packages, shows how to organise your own code into importable files.
The last lesson in five lines. If you only skimmed it, read these first.
re.search(p, s) finds a pattern anywhere and returns a match object or None. re.fullmatch(p, s) succeeds only when the pattern covers the whole string.(?P<name>...) names it, groupdict() returns every named group, and \1 inside a pattern repeats what group 1 matched.re.IGNORECASE, re.MULTILINE (^ and $ work per line), re.DOTALL and re.VERBOSE..* takes as much as it can. A lazy .*?, or a class such as [^>]*, stops sooner.r"...", so Python leaves the backslashes for the regex engine.import re
m = re.search(r"(?P<key>\w+)=(?P<value>\S+)", "loss=0.43 lr=0.01")
print(m.groupdict())
print(bool(re.fullmatch(r"\d+", "42 ")))
print(re.findall(r"<.*?>", "<b>hi</b>"))
# Output:
# {'key': 'loss', 'value': '0.43'}
# False
# ['<b>', '</b>']re.sub calls, each doing one job. Doing them one at a time and printing the result is the easiest way to debug a chain.import re
raw = "<p>Loved it!! Best movie EVER, see https://example.com/r?id=7 :) the the end. 5 stars</p>"
step = re.sub(r"<[^>]+>", " ", raw)
step = re.sub(r"https?://\S+", " ", step)
step = step.lower()
step = re.sub(r"\d+", "0", step)
step = re.sub(r"[^a-z0' ]", " ", step)
step = re.sub(r"\b(\w+) \1\b", r"\1", step)
step = re.sub(r"\s+", " ", step).strip()
print(step)
print(step.split())
# Output:
# loved it best movie ever see the end 0 stars
# ['loved', 'it', 'best', 'movie', 'ever', 'see', 'the', 'end', '0', 'stars']0, anything that is not a letter, 0, an apostrophe or a space becomes a space, a backreference collapses a repeated word, and finally runs of whitespace collapse to one space. A backreference step, the one from the Groups section of the last lesson, collapses a repeated word: it is why the the became the. The number step is why 5 stars ends as 0 stars.Order matters: the link pattern needs the colon and slashes that the punctuation step would delete. And a regex strip is good enough for tidying text but is not an HTML parser, so nested or malformed markup needs a real one.
You can also tokenize with a regex directly. A pattern that finds words and keeps contractions together does it in one line.
import re
print(re.findall(r"[a-z']+", "i can't believe it's 50% off"))
# Output:
# ['i', "can't", 'believe', "it's", 'off'][a-z']+ accepts only letters and apostrophes. The pattern is a decision about what a token is.run-20260314-0042: the word run, an eight digit date, then a four digit counter. To validate a whole string, use re.fullmatch.import re
RUN_ID = re.compile(r"run-\d{8}-\d{4}")
for s in ["run-20260314-0042", "run-20260314-42", "xrun-20260314-0042", "run-20260314-0042 "]:
print(repr(s), bool(RUN_ID.fullmatch(s)))
# Output:
# 'run-20260314-0042' True
# 'run-20260314-42' False
# 'xrun-20260314-0042' False
# 'run-20260314-0042 ' Falsesearch and name the parts you want.import re
m = re.search(r"run-(?P<day>\d{8})-(?P<seq>\d{4})", "see run-20260314-0042 now")
print(m.groupdict())
# Output:
# {'day': '20260314', 'seq': '0042'}Email addresses look like a natural regex job, and they show where regex stops being the right tool. A short pattern catches the usual typos.
import re
EMAIL = re.compile(r"[\w.+-]+@[\w-]+(?:\.[\w-]+)+")
tests = ["ada@example.com", "ada.l+ml@mail.example.co.uk", "ada@", "ada example@x.com",
"a@b..com", "ada@example", "nobody@nowhere.invalid", '"ada l"@example.com']
for e in tests:
print(e, bool(EMAIL.fullmatch(e)))
print(re.findall(r"[\w.+-]+@[\w-]+(?:\.[\w-]+)+", "Mail ada@example.com or grace@mail.org."))
# Output:
# ada@example.com True
# ada.l+ml@mail.example.co.uk True
# ada@ False
# ada example@x.com False
# a@b..com False
# ada@example False
# nobody@nowhere.invalid True
# "ada l"@example.com False
# ['ada@example.com', 'grace@mail.org']nobody@nowhere.invalid, an address that cannot receive mail, and it rejects "ada l"@example.com, which the email standard allows. A regex checks the shape of the text. It cannot tell you whether anyone reads that mailbox, and a pattern that tried to follow the whole standard would be enormous. The usual approach is a loose shape check at input time, then a confirmation message, because only a delivered message proves an address works.(?=...) means "followed by" and (?!...) means "not followed by". A lookbehind (?<=...) checks what comes before.import re
t = "latency 120 ms, size 4 MB, latency 85 ms"
print(re.findall(r"\d+ ms", t))
print(re.findall(r"\d+(?= ms)", t))
print(re.findall(r"(?<=latency )\d+", t))
# Output:
# ['120 ms', '85 ms']
# ['120', '85']
# ['120', '85']re, a lookbehind must have a fixed width, so (?<=a|bc)x raises re.error: look-behind requires fixed-width pattern. If you find yourself nesting several lookarounds, a group with search is usually easier to read.re is a backtracking engine. At each choice it tries one path, and if the rest of the pattern fails it goes back and tries another. That is why a single failed match can cost far more than a successful one. When two quantifiers can claim the same characters, the number of paths explodes.(a+)+b. A run of a characters can be divided between the inner a+ and the outer + in many ways, and when there is no b at the end the engine must try them all before giving up.import re
import time
pattern = re.compile(r"(a+)+b")
for n in [12, 14, 16, 18, 20]:
text = "a" * n + "c"
start = time.perf_counter()
pattern.search(text)
ms = (time.perf_counter() - start) * 1000
print(n, f"{ms:.1f} ms")a about doubles the time. The input is tiny and boring. The cost lives entirely in the pattern's ambiguity.a+b fails on the same text in a few microseconds, because there is only one way to match.Which of these patterns is the most dangerous on text that nearly matches but fails at the end?
[^>]* over .*. And never run a regex that a user typed against your own server, because a hostile pattern can freeze it. Some engines, such as Google's RE2, avoid backtracking altogether and match in time proportional to the text length, at the price of dropping features like backreferences. Backtracking blow-ups have caused real outages, including Stack Overflow in 2016 and Cloudflare in 2019, both described in the companies' published post-mortems.This should report True only for a run ID that is exactly run, eight digits, a dash and four digits. It reports True for an ID with a stray newline and for an ID with text in front. Fix the check.
[True, False, False, False]
Tests · Prints valid ids: 3 and mean latency: 168.33 (the latencies 120, 85 and 300). The message with the malformed ID run-2026-0001 is skipped, and the 0.43 in the first message is not collected because it is not followed by ms.
| Job or pattern | What it does |
|---|---|
re.sub(r"<[^>]*>", " ", s) | replace every tag with a space |
re.sub(r"https?://\S+", " ", s) | remove links |
re.sub(r"\s+", " ", s).strip() | collapse runs of whitespace |
re.sub(r"\b(\w+) \1\b", r"\1", s) | collapse one repeated word |
re.fullmatch(p, s) | validate a whole string |
re.search(p, s) with named groups | pull a field out of free text |
(?=...) |
csv module for CSV files, even when quoted commas tempt you toward a regex. The tokenizer below shows what happens to the cleaned text next.Interactive Lab
See how text gets split into tokens, the step that comes after the cleaning in this lesson
re.sub calls, and the order matters because each step sees only what the previous step left. To validate a whole string use re.fullmatch, and to take an ID out of free text use re.search with named groups. A regex checks shape, never meaning: it cannot tell you whether an email is deliverable or whether markup is well formed.A lookahead or a lookbehind checks the text around a match without including it. Nested quantifiers over the same characters can backtrack catastrophically, with the time roughly doubling per extra character, so prefer specific classes and never run a pattern that a user typed on your own server.
The pattern (a+)+b took about 1.4 seconds on a string of 24 a's followed by c, on one machine. Roughly how long would you expect 26 a's to take on that machine?
(?!...)| followed by, not followed by |
(?<=...) (?<!...) | preceded by, not preceded by (fixed width only) |
(a+)+b | the shape to avoid: a quantifier on a group that can match the same characters in several ways |