Strings, Part 1: Slicing, Methods and Formatting
After this lesson, you will be able to:
- Understand why strings are immutable and what that means for how Python handles them
- Index and slice strings with positive and negative indices, step sizes, and reversal
- Use the full suite of essential string methods for case, search, split, join, and padding
- Master f-string format specifications: alignment, precision, thousands, percentages, padding
- Use the .format() method for string interpolation and understand when you will encounter it in older code
- Write multi-line strings, raw strings, and byte literals correctly
- Explain how Python handles text from every language in the world, and the difference between str and bytes
- Avoid the O(n²) string concatenation trap and use join() for efficient string building
- Use the re module for basic pattern matching in text
Before You Start
#Strings are Immutable Sequences
# Creating strings -- four equivalent ways
s1 = 'single quotes'
s2 = "double quotes"
s3 = '''triple single quotes (can span lines)'''
s4 = """triple double quotes (can span lines)"""
# Strings are sequences -- you can iterate
word = "Python"
for char in word:
print(char, end=" ") # P y t h o n
print()
print(len(word)) # 6
print(word[0]) # 'P'
print(word[-1]) # 'n'
# But you CANNOT change a character
try:
word[0] = "J" # TypeError!
except TypeError as e:
print(f"Error: {e}") # 'str' object does not support item assignmentTry it! Type"hello"[0]in the Python REPL. What letter do you get? Now try"hello"[-1]. Python counts from 0, and negative numbers count from the end!
.split() to .encode() with before/after visualization.# Every "modification" returns a NEW string object
name = "alice"
upper = name.upper() # new object
print(name) # "alice" -- original unchanged!
print(upper) # "ALICE" -- new object
# To actually change the variable, reassign it
name = name.upper()
print(name) # "ALICE"
# Strings can be used as dict keys (because they're immutable and hashable)
word_count = {
"machine": 42,
"learning": 38,
"data": 55,
}
print(word_count["machine"]) # 42#String Indexing and Slicing
String indexing and slicing use the exact same syntax as lists. Master this once, use it everywhere.
s = "Python"
# 0123456 (positive indices)
# -6-5-4-3-2-1 (negative indices)
# Single character access
print(s[0]) # 'P' -- first character
print(s[1]) # 'y'
print(s[5]) # 'n' -- last character (index = len-1)
print(s[-1]) # 'n' -- last character (negative: from end)
print(s[-2]) # 'o' -- second to last
print(s[-6]) # 'P' -- first character via negative index# Slicing: s[start:stop:step]
# start is inclusive, stop is exclusive
s = "Hello, World!"
print(s[0:5]) # "Hello" -- index 0 to 4
print(s[7:12]) # "World" -- index 7 to 11
print(s[:5]) # "Hello" -- from start to 4
print(s[7:]) # "World!" -- from 7 to end
print(s[:]) # "Hello, World!" -- full copy
# Step
print(s[::2]) # "Hlo ol!" -- every other character
print(s[::3]) # "H,rd" -- every third character
print(s[::-1]) # "!dlroW ,olleH" -- reversed!
# Negative start/stop
print(s[-6:]) # "orld!" -- last 6 characters
print(s[-6:-1]) # "orld" -- last 6, excluding very lastWhat does 'abcdef'[1:-1:2] return?
# Practical slicing patterns
filename = "report_2024_final.pdf"
print(filename[:-4]) # "report_2024_final" -- strip extension
print(filename[-3:]) # "pdf" -- get extension
email = "alice@example.com"
at_pos = email.index("@")
username = email[:at_pos]
domain = email[at_pos+1:]
print(username) # "alice"
print(domain) # "example.com"
# Palindrome check using reversal
def is_palindrome(s):
s = s.lower().replace(" ", "")
return s == s[::-1]
print(is_palindrome("racecar")) # True
print(is_palindrome("A man a plan a canal Panama")) # True (after cleanup)
print(is_palindrome("hello")) # FalseHitIndexError: string index out of rangeorTypeError: string indices must be integers? Asking fors[len(s)]or using a float/string index are the usual culprits. See the error decoder for fixes — slicing (s[a:b]) never raises on out-of-range, only indexing does.
#Essential String Methods
Python's string methods are one of the most useful parts of the standard library. They're called on the string with dot notation and always return a new string (because strings are immutable).
#Case Methods
text = "hello, World! this is PYTHON."
print(text.upper()) # "HELLO, WORLD! THIS IS PYTHON."
print(text.lower()) # "hello, world! this is python."
print(text.title()) # "Hello, World! This Is Python."
print(text.capitalize()) # "Hello, world! this is python." (only first char up)
print(text.swapcase()) # "HELLO, wORLD! THIS IS python."
# Making case consistent before comparing
words = ["Machine", "LEARNING", "machine", "MACHINE"]
normalized = [w.lower() for w in words]
unique = set(normalized)
print(unique) # {'machine', 'learning'} -- deduplication works correctly#Testing / Checking Methods
# These return True or False
print("hello123".isalnum()) # True -- all alphanumeric
print("hello".isalpha()) # True -- all letters
print("12345".isdigit()) # True -- all digits
print(" ".isspace()) # True -- all whitespace
print("Hello World".istitle()) # True -- title case
# Starts/ends with
url = "https://api.example.com/v1/models"
print(url.startswith("https://")) # True
print(url.endswith("/models")) # True
# Multiple prefixes/suffixes (pass a tuple)
filename = "data.csv"
print(filename.endswith((".csv", ".tsv", ".parquet"))) # True
# Checking substring membership (using 'in' operator)
sentence = "learning python is fascinating"
print("learning" in sentence) # True
print("deep" in sentence) # False
print("learning" not in sentence) # False#Searching Methods
text = "the cat sat on the mat"
# find() returns index of first occurrence, -1 if not found (safe)
print(text.find("cat")) # 4
print(text.find("dog")) # -1 -- not found, no error
# index() same as find() but raises ValueError if not found
print(text.index("cat")) # 4
try:
text.index("dog")
except ValueError as e:
print(f"Error: {e}") # substring not found
# rfind() and rindex() search from the RIGHT
print(text.rfind("the")) # 18 -- last occurrence
print(text.rfind("at")) # 20 -- last "at"
# count() counts non-overlapping occurrences
print(text.count("the")) # 2
print(text.count("at")) # 3 -- "cat", "sat", "mat"
print("aaa".count("aa")) # 1 -- non-overlapping!What does name print at the end? name = "alice" name.upper() print(name)
#Modifying Methods
# replace() -- returns new string with substitutions
text = "I love cats. Cats are great. My cat is named Whiskers."
print(text.replace("cat", "dog"))
# "I love dogs. Cats are great. My dog is named Whiskers."
# Note: case-sensitive! "Cats" was not replaced
print(text.replace("cat", "dog", 1)) # replace only first occurrence
# "I love dogs. Cats are great. My cat is named Whiskers."
# strip(), lstrip(), rstrip() -- remove whitespace (or specified chars)
messy = " \t hello world \n "
print(repr(messy.strip())) # 'hello world'
print(repr(messy.lstrip())) # 'hello world \n '
print(repr(messy.rstrip())) # ' \t hello world'
# Strip specific characters
csv_field = ',"Alice",'
print(csv_field.strip(',').strip('"')) # 'Alice'
# Cleaning up user input
dirty_tokens = [" hello ", "\nworld\t", " python "]
clean_tokens = [t.strip() for t in dirty_tokens]
print(clean_tokens) # ['hello', 'world', 'python']#Splitting Methods
# split() -- split on whitespace by default (handles multiple spaces, tabs, newlines)
sentence = " the quick brown fox "
words = sentence.split()
print(words) # ['the', 'quick', 'brown', 'fox'] -- cleaned up!
# split(delimiter) -- split on specific character
csv = "Alice,25,Engineer,New York"
fields = csv.split(",")
print(fields) # ['Alice', '25', 'Engineer', 'New York']
# split with maxsplit limit
text = "key=value=with=equals"
key, rest = text.split("=", 1) # split at most once
print(key) # "key"
print(rest) # "value=with=equals"
# rsplit() -- split from the right
path = "/usr/local/bin/python"
print(path.rsplit("/", 1)) # ['/usr/local/bin', 'python']
# splitlines() -- split on any line ending (\n, \r\n, \r)
multiline = "Line 1\nLine 2\r\nLine 3\rLine 4"
print(multiline.splitlines()) # ['Line 1', 'Line 2', 'Line 3', 'Line 4']
# partition() -- split into exactly 3 parts: before, separator, after
url = "https://example.com/path?query=1"
protocol, sep, rest = url.partition("://")
print(protocol) # "https"
print(rest) # "example.com/path?query=1"What does 'hello'[::-1] produce?
[::-1] means "every character, stepping by -1", which iterates backwards. This is the canonical Python idiom for reversing a string (or any sequence). The general form is [start:stop:step]; negative step reverses direction. Compare with ''.join(reversed('hello')) which also works but is longer. Common variations: s[::2] (every other char), s[1:] (drop first char), s[:-1] (drop last char).#Joining Methods
# join() -- the FAST way to concatenate many strings
words = ["machine", "learning", "is", "fun"]
# join: put separator BETWEEN each item
sentence = " ".join(words)
print(sentence) # "learning python is fun"
# Different separators
csv_line = ",".join(words)
print(csv_line) # "machine,learning,is,fun"
path = "/".join(["usr", "local", "bin", "python"])
print(path) # "usr/local/bin/python"
# No separator
initials = "".join([w[0].upper() for w in words])
print(initials) # "MLIF"
# join is the CORRECT way to build strings in a loop (see CommonMistake)
tokens = ["The", "quick", "brown", "fox"]
result = " ".join(tokens) # ONE allocation, O(n) time
print(result)#Padding Methods
# ljust, rjust, center -- pad to a minimum width
name = "Alice"
print(name.ljust(10)) # "Alice " -- left-aligned, padded right
print(name.rjust(10)) # " Alice" -- right-aligned, padded left
print(name.center(10)) # " Alice " -- centered
print(name.center(10, "-")) # "--Alice---" -- custom fill character
# zfill -- pad numbers with leading zeros
for n in [1, 42, 100, 1234]:
print(str(n).zfill(5))
# 00001
# 00042
# 00100
# 01234
# Practical: formatting a table
headers = ["Name", "Score", "Grade"]
row1 = ["Alice", "95", "A"]
row2 = ["Bob", "78", "C"]
print(f"{headers[0]:<10}{headers[1]:>8}{headers[2]:>8}")
print(f"{row1[0]:<10}{row1[1]:>8}{row1[2]:>8}")
print(f"{row2[0]:<10}{row2[1]:>8}{row2[2]:>8}")
# Name Score Grade
# Alice 95 A
# Bob 78 CWhat does `' hi '.strip()` return?
#The .format() Method: The Classic Approach
.format() was the standard way to interpolate values into strings. You will see it constantly in older codebases, documentation, and Stack Overflow answers.# Basic .format(): {} are placeholders, filled by positional args
print("Hello, {}! You are {} years old.".format("Alice", 25))
# Hello, Alice! You are 25 years old.
# Positional index
print("{0} and {1}, then {0} again".format("first", "second"))
# first and second, then first again
# Named placeholders
print("{name} scored {score:.1f}%".format(name="Bob", score=87.3))
# Bob scored 87.3%
# Format spec works the same as f-strings
pi = 3.14159
print("Pi is {:.3f}".format(pi)) # Pi is 3.142
print("Count: {:,}".format(1234567)) # Count: 1,234,567
print("Accuracy: {:.1%}".format(0.934)) # Accuracy: 93.4%
# Filling a template from a dict
template = "Name: {name}, Age: {age}, City: {city}"
data = {"name": "Alice", "age": 30, "city": "London"}
print(template.format(**data)) # Name: Alice, Age: 30, City: London#F-String Mastery: The Format Spec Mini-Language
f"...") are Python's most powerful string formatting tool. Inside {} you can put any expression, and after a : you can add a format specification to control exactly how the value is displayed.{value:[[fill]align][sign][#][0][width][grouping][.precision][type]}# Basic f-strings
name = "Alice"
score = 95.678
count = 1234567
print(f"Hello, {name}!") # "Hello, Alice!"
print(f"Score: {score}") # "Score: 95.678"
print(f"2 + 2 = {2 + 2}") # "2 + 2 = 4" (expressions work!)
print(f"Upper: {name.upper()}") # "Upper: ALICE" (method calls work!)# FLOAT PRECISION: .Nf controls decimal places
pi = 3.14159265358979
print(f"{pi:.2f}") # "3.14" -- 2 decimal places
print(f"{pi:.4f}") # "3.1416" -- 4 decimal places (rounded)
print(f"{pi:.0f}") # "3" -- no decimals (rounded)
print(f"{pi:10.3f}") # " 3.142" -- width 10, 3 decimals
# THOUSANDS SEPARATOR: , (comma)
big_num = 1234567890
print(f"{big_num:,}") # "1,234,567,890"
print(f"{big_num:,.2f}") # "1,234,567,890.00"
print(f"{3.5e6:,.0f}") # "3,500,000"
# PERCENTAGE: .N% (multiplies by 100, adds %)
accuracy = 0.9734
print(f"{accuracy:.1%}") # "97.3%"
print(f"{accuracy:.0%}") # "97%"
print(f"{0.5:.2%}") # "50.00%"# ALIGNMENT: < left, > right, ^ center
# {value:fill_char align width}
label = "hello"
print(f"{label:<10}") # "hello " left-aligned, width 10
print(f"{label:>10}") # " hello" right-aligned
print(f"{label:^10}") # " hello " centered
print(f"{label:*^10}") # "**hello***" centered with * fill
print(f"{label:-<10}") # "hello-----" left with - fill
print(f"{42:0>8}") # "00000042" zero-pad to width 8
# Numbers right-align by default, strings left-align
for name, score in [("Alice", 95), ("Bob", 78), ("Charlie", 88)]:
print(f"{name:<10} {score:>5}")
# Alice 95
# Bob 78
# Charlie 88# REPR and STR conversion
data = [1, "hello", None, 3.14]
print(f"{data!r}") # repr: [1, 'hello', None, 3.14] (shows quotes)
print(f"{data!s}") # str: [1, 'hello', None, 3.14] (same here)
text = "hello\tworld\n"
print(f"{text!r}") # 'hello\\tworld\\n' (shows escape sequences)
print(f"{text!s}") # hello world (interprets escape sequences)
# DEBUG format (Python 3.8+): variable=value
x = 42
y = [1, 2, 3]
name = "Alice"
print(f"{x=}") # x=42
print(f"{y=}") # y=[1, 2, 3]
print(f"{name=}") # name='Alice'
print(f"{x + y[0]=}") # x + y[0]=43 (expressions too!)# NESTED f-strings and dynamic formatting
# You can put variables inside the format spec!
width = 10
precision = 3
value = 3.14159
print(f"{value:{width}.{precision}f}") # " 3.142"
# Dynamic column widths
col_width = 15
data = [("Alice", 95.5), ("Bob", 78.3), ("Charlie", 88.9)]
for name, score in data:
print(f"{name:{col_width}} {score:.1f}%")
# Integer formats
n = 255
print(f"{n:d}") # "255" -- decimal (default)
print(f"{n:b}") # "11111111" -- binary
print(f"{n:o}") # "377" -- octal
print(f"{n:x}") # "ff" -- hex lowercase
print(f"{n:X}") # "FF" -- hex uppercase
print(f"{n:#x}") # "0xff" -- hex with 0x prefix
print(f"{n:#b}") # "0b11111111" -- binary with 0b prefixWhat does `f'{3.14159:.2f}'` produce?
#Multi-line Strings, Raw Strings, and Bytes
#Multi-line Strings
# Triple quotes span multiple lines
poem = """Roses are red,
Violets are blue,
Python is great,
And so are you."""
print(poem)
print(len(poem.splitlines())) # 4 lines
# Useful for SQL queries, HTML templates, etc.
query = """
SELECT name, score
FROM students
WHERE score >= 90
ORDER BY score DESC
LIMIT 10
"""
# Multi-line in an expression (use backslash to avoid leading newline)
message = (
"Dear Alice,\n"
"Your order has shipped.\n"
"Expected delivery: Monday."
)
# Adjacent string literals are automatically concatenated at compile time!#Raw Strings
# In regular strings, backslash has special meaning (escape sequences)
print("C:\\Users\\Alice\\Documents") # C:\Users\Alice\Documents
print("\t tab \n newline") # actual tab and newline
# Raw strings (r"...") treat backslashes as literal characters
path = r"C:\Users\Alice\Documents"
print(path) # C:\Users\Alice\Documents (no need to double-escape)
print(len(path)) # 25 (backslashes are real chars)
# Critical use: regular expressions (backslashes are common in regex)
import re
# Without raw string: double every backslash
pattern1 = "\\d+\\.\\d+" # matches digits.digits
# With raw string: much cleaner!
pattern2 = r"\d+\.\d+" # same pattern, but readable
text = "Pi is approximately 3.14159"
match1 = re.findall(pattern1, text)
match2 = re.findall(pattern2, text)
print(match1) # ['3.14159']
print(match2) # ['3.14159'] (same result)
# Windows file paths
windows_path = r"C:\Program Files\Python\python.exe"
print(windows_path) # no escaping needed!#Bytes vs Strings
# str is a sequence of Unicode characters (text)
text = "Hello"
print(type(text)) # <class 'str'>
# bytes is a sequence of raw bytes (0-255) -- for network/file I/O
data = b"Hello"
print(type(data)) # <class 'bytes'>
print(data[0]) # 72 (ASCII code for 'H')
# You CANNOT mix str and bytes
try:
result = "Hello" + b" World"
except TypeError as e:
print(f"Error: {e}") # can only concatenate str (not "bytes") to str
# Converting between str and bytes: encode / decode
text = "Hello, World!"
encoded = text.encode("utf-8") # str --> bytes
print(encoded) # b'Hello, World!'
print(type(encoded)) # <class 'bytes'>
decoded = encoded.decode("utf-8") # bytes --> str
print(decoded) # "Hello, World!"
print(type(decoded)) # <class 'str'>What does 'Python'[1:-1] return?
This program is supposed to print the name in uppercase, but it's printing the original lowercase string. Fix it without changing the line that defines name.
Hello, ADA LOVELACE