Python Mini-Projects, Part 1
After this lesson, you will be able to:
- Apply functions, data structures, and control flow together to build a working calculator with history
- Combine dictionaries, OOP, and file I/O to create a persistent contact book
- Use string processing, collections.Counter, and file I/O to build a word frequency analyzer
Before You Start
#How This Lesson Works
For each project, you will get:
- A problem statement. What you are building and why
- Concepts used. Which earlier lessons are relevant
- Starter code — a skeleton with TODOs to guide you
- A full solution — to compare against after you try it yourself
- Challenge extensions — harder tasks for when you want more
#Project 1: Calculator with History
#Problem Statement
This is a classic first project because it ties together user input, functions, conditionals, and data structures in a way that feels like a real application.
#Concepts Used
- Functions. Each operation is a separate function (from the Functions lesson)
- Lists as stacks. History uses
append()andpop()(from the Data Structures lesson) - Control flow — menu loop with if/elif (from the Control Flow lesson)
- String processing — parsing user input
#Starter Code
Try to fill in the TODOs before looking at the solution.
Tests · Run the starter code first — you should see errors where TODOs are. Fill them in, then compare with the solution.
#Challenge Extensions
- Scientific operations. Add
power(a, b),sqrt(a), andmodulo(a, b)functions - Expression parser — instead of separate a, op, b values, parse a string like
"10 + 5"using.split() - Persistent history. Save history to a text file so it survives between runs (use what you learned in the File I/O lesson)
- Redo functionality. Add a separate "redo stack" that stores undone items and lets you redo them
#Project 2: Contact Book
#Problem Statement
This project combines data structures with file I/O and introduces you to the kind of CRUD (Create, Read, Update, Delete) operations that are the foundation of almost every software application.
#Concepts Used
- Dictionaries. Contacts stored as name-to-info mappings (from the Dicts & Hash Tables lesson)
- OOP — the ContactBook as a class with methods (from the Classes & OOP lesson)
- File I/O — saving and loading contacts with JSON (from the File I/O lesson)
- Exception handling — graceful error handling for missing files and invalid input
#Starter Code
Tests · Fill in all TODO methods, then run to test. The output should show contacts being added, searched, deleted, and saved.
#Challenge Extensions
- Update contact. Add an
update_contact(name, phone=None, email=None)method that updates only the provided fields - Group contacts. Add a "group" field (family, work, friends) and a
list_by_group(group)method - Export to CSV. Add a
save_to_csv(filename)method that writes contacts in comma-separated format - Input validation — validate that phone numbers contain only digits and dashes, and emails contain an
@symbol - Favorite contacts. Add a boolean
favoritefield and alist_favorites()method
#Project 3: Word Frequency Analyzer
#Problem Statement
Build a tool that reads text, counts the frequency of each word, filters out common "stop words" (like "the", "a", "is"), and reports the top N most frequent words.
This is the exact foundation of how text processing works in NLP. When you hear about "tokenization" and "term frequency" in machine learning, this is what is happening at the simplest level. Every NLP pipeline starts with exactly this kind of word counting — you are building the first stage of a natural language processing system.
#Concepts Used
- Strings — splitting, lowering, stripping punctuation (from the Strings Deep Dive lesson)
- Dictionaries — manual word counting (from the Dicts & Hash Tables lesson)
- collections.Counter — efficient counting (from the Collections Module lesson)
- File I/O — reading text from files (from the File I/O lesson)
- Functions — clean decomposition into reusable pieces
#Connection to Machine Learning
This is how tokenization works at the simplest level. In NLP, the first step of any text pipeline is breaking text into tokens (words) and counting their frequency. The techniques you build here — lowercasing, removing punctuation, filtering stop words, counting frequencies — are the exact preprocessing steps used before feeding text into models like BERT, GPT, and other transformers. The only difference is that production systems use more sophisticated tokenizers (like BPE or WordPiece), but the principle is identical.
#Starter Code
Tests · Fill in all TODO functions, then run the analysis. You should see a text-based bar chart of the most frequent words in the sample text.
#Challenge Extensions
- Bigrams. Instead of single words, count pairs of consecutive words (e.g., "machine learning" appears together). This is called bigram analysis and is used heavily in NLP
- TF-IDF. If you have multiple documents, compute Term Frequency-Inverse Document Frequency to find words that are important to a specific document but not common across all documents
- Read from file. Modify
analyze_textto accept a filename instead of a string, and read the file contents - Sentiment words. Create a list of positive words and negative words, then count how many of each appear in the text to get a basic sentiment score
- Zipf's Law — plot word rank vs. frequency and observe that it follows Zipf's law (the nth most common word appears roughly 1/n times as often as the most common word)
#What You Have Built
Take a moment to appreciate what you just did. In three projects, you used:
| Concept | Project 1 | Project 2 | Project 3 |
|---|---|---|---|
| Functions | Yes | Yes | Yes |
| Control Flow | Yes | Yes | Yes |
| Lists | Yes | Yes | Yes |
| Dictionaries | Yes | Yes | Yes |
| OOP (Classes) | — | Yes | — |
| File I/O | — | Yes | Yes |
| String Processing | Yes | — | Yes |
| collections.Counter | — | — | Yes |
| json module | — | Yes | — |
| Exception Handling | Yes | Yes | — |
#Design Principles You Practiced
Even though this lesson was about building, you were also practicing software design without realizing it:
#1. Separation of Concerns
clean_text() only cleans text. tokenize() only splits into words. count_words() only counts. This makes code easier to test, debug, and reuse.#2. DRY (Don't Repeat Yourself)
calculate() function in Project 1 routes to specific operation functions instead of repeating arithmetic logic.#3. Encapsulation (OOP)
ContactBook class owns its contacts dict and all the methods that operate on it. Outside code does not need to know how contacts are stored internally. Part 2's quiz game pushes this further with three cooperating classes.#4. Graceful Error Handling
Division by zero, missing files, invalid input — your code handles these cases instead of crashing. This is what separates a script from software.
#5. Pipeline Architecture
In the Calculator project, why is history implemented as a list (stack) rather than a dictionary?