Recruiters spend seconds on a portfolio. In those seconds, two questions get answered: "Can this person ship?" and "Is this real, or AI-cribbed from Cursor?" The 5 projects in this guide are designed to answer both.
The advice that follows rests on one claim: one (1) deep, complete, demonstrably-shipped project that you can point to in 10 seconds does more for you than a long list of shallow ones. Not five projects. Not a Coursera completion certificate. One project, deep, real, and live. You do not have to take that on faith — it follows directly from how screening actually works. A reviewer with a stack of applications and limited time cannot evaluate breadth; they sample your strongest signal and extrapolate. Optimizing for the best single artifact is the rational response to that process. This lesson is how to build that one project — and how to package it so it survives the sample.
Learning Objectives
After this lesson, you will be able to:
Apply the rule of 1: one great project beats five mediocre projects
Build a 10-point portfolio rubric you can score your own work against
Distinguish weak, average, and strong portfolios — with specific examples and the reasoning that ranks them
Write a GitHub README that survives the 30-second hiring-manager scan
Build a portfolio website with the four sections that decide first impressions
Optimize LinkedIn so the headline, About, and Featured sections work together
Use blog posts and demo videos as portfolio multipliers
Avoid the 7 portfolio mistakes that nuke otherwise-strong candidates
The phrase to internalize: depth, not breadth. A typical strong candidate's GitHub has 8-12 repositories — most are small experiments, course exercises, or forks. Exactly one is the "headline" project: a system that took 40-80 hours to build, has a clean README, runs in one click, has tests, has an eval, has a writeup, and has a live demo or video. That one project is what the recruiter forwards, what the hiring manager screenshots in Slack, and what every interview round will return to.
Building five tutorial-grade projects is a coping mechanism. It feels productive. It signals nothing.
Score your portfolio on these 10 dimensions, 1 point each. Aim for 8+. Be honest.
#
Dimension
What "1 point" looks like
1
Real problem
The project solves a problem that exists outside CS coursework. Not MNIST. Not Titanic. Something a user wants.
2
Works end-to-end
A user can hit a URL or run one command and use it. No "intended workflow" diagrams replacing actual function.
3
One-click setup
git clone && docker compose up (or equivalent). README has correct env vars. No hidden assumptions.
4
Clean README
First 5 lines tell me what it does, what stack, and how to try it. Screenshots or GIF in the README.
5
Tests exist and pass in CI
At least 5-10 tests. GitHub Actions runs them. Green checkmark visible.
6
Evaluation methodology
You can quantify how well it works. Accuracy/precision/recall, latency, cost, or a domain-specific metric.
7
Live demo or video
Deployed to a free tier (HF Spaces, Vercel, Modal) OR a 2-minute screen recording showing it in action.
8
Writeup / blog post
1,000-2,000 words explaining the design, trade-offs, what failed, what you learned. Linked from README.
9
Real-world hardening
Logging, error handling, rate limiting, basic security. Not a Jupyter notebook in disguise.
10
One non-trivial technical decision
You made a specific choice (e.g. hybrid search > pure vector, GPT-4o-mini > GPT-4o for cost) and you can defend it.
A 10/10 portfolio is rare and gets interviews almost regardless of resume. A 7/10 portfolio gets interviews with a decent resume. A 4/10 portfolio gets ignored regardless of resume.
These are constructed examples — composites written to illustrate the scoring rubric, not real candidates. Score your own portfolio against the same rubric as you read.
Candidate has 47 GitHub repositories. All public. Top-pinned:
coursera-ml-week1 — sklearn linear regression on synthetic data. README: "My homework for Coursera Andrew Ng Course." No screenshots. No demo.
mnist-pytorch — copy of the official PyTorch MNIST tutorial. README: "MNIST classifier in PyTorch." No deviation from the tutorial. No writeup.
chatbot-langchain — a notebook that calls ChatOpenAI from LangChain and runs agent.run("Hello"). README: 4 lines, no setup instructions, no API key handling.
44 other forks and small one-off scripts.
Score: 3/10 (real problem 0, works end-to-end 0, one-click setup 0, clean README 0, tests 0, eval 0, demo 0, writeup 0, hardening 0, non-trivial decision 0, but each repo has SOMETHING running, so points 1 + 2 are scratched in).
Why it fails: The candidate has demonstrably written code, but every single project is a tutorial reproduction. Zero original work. A hiring manager scanning this learns nothing new about the candidate's judgment, taste, or ability to ship — only that they have completed online courses. Courses are table stakes; they do not differentiate.
The fix is not "add more repos." The fix is to delete 40 of them, pin 5 that show progression, and invest 60 focused hours into a single new project that scores 8+ on the rubric.
Candidate has 12 repositories, 3 pinned. The headline project:
legal-doc-rag — A RAG system over a public corpus of US case law. Built with LangChain, ChromaDB, GPT-4. Streamlit UI. README has setup instructions and a screenshot. Has a requirements.txt. Live demo on Hugging Face Spaces (free tier, slow but functional). One writeup: a 600-word blog post about the build.
Score: 6/10. Real problem (1), works end-to-end (1), one-click setup (1 — pip install works, though no Docker), clean README (1), tests (0 — none), eval (0 — no quantified accuracy), live demo (1), writeup (1 — present but thin), hardening (0 — no error handling, no rate limiting), non-trivial decision (0 — used default settings throughout).
Why it is "average and not strong": this is a respectable project but it is the same project a hundred other candidates have shipped. No evaluation, no thoughtful design choices, no demonstration of senior thinking. The candidate would get filtered out of frontier-lab loops but might make it into a smaller AI startup if their resume is otherwise OK.
Lift from 6 to 9: Add an eval harness with 50 hand-labeled question-answer pairs. Measure accuracy on the eval set. Try 3 retrieval strategies (BM25, cosine vector, hybrid) and report which won. Document the choice. Add 10 unit tests on the chunking and retrieval functions. Add basic error handling for API timeouts. Pin the dependencies. That work is roughly 20 additional hours, and it triples the portfolio's hiring power.
Candidate has 8 repositories, 3 pinned, with one anchor project that takes the entire first screen:
clinical-trial-matcher — A system that matches a patient summary to relevant active clinical trials from clinicaltrials.gov. Hybrid retrieval (BM25 + sentence-transformers vector + a reranker). FastAPI backend, React frontend. Deployed to fly.io with a public demo. README has a GIF, an architecture diagram, links to the writeup, links to a live API doc page.
Eval harness: 200 hand-labeled patient-trial pairs. Reports precision@5, recall@10, mean reciprocal rank. Compares 4 different retrieval strategies in a table.
Tests: 47 unit tests, 8 integration tests, GitHub Actions runs them on every push.
Hardening: rate-limited public API, OpenAI fallback to local model on quota exhaustion, structured logging to a free Logtail instance.
Writeup: 2,400-word blog post titled "Why semantic similarity wasn't enough for clinical trial matching." Discusses the failure of pure vector retrieval on rare disease names, the BM25 lift, the reranker's behavior on negation, and an unsuccessful attempt to use LLM-as-judge that the candidate had to abandon and explain why.
Score: 9/10. Loses 1 point only because there is no on-call/observability layer (out of scope for a portfolio, but a 10/10 would add basic Prometheus metrics + a Grafana dashboard screenshot).
Why it works: Real domain problem with a real user (themselves, or a healthcare friend they tested with). Honest writeup that includes a failure. Quantified evaluation. Hybrid retrieval with a defensible reason. Real deployment that anyone can try. This candidate gets a phone screen at every company they apply to, including frontier labs. The portfolio has done 80% of the recruiter's job for them.
A personal portfolio website is not optional for any candidate targeting Big Tech or frontier labs, and is highly recommended for everyone else.
Four sections, in this order:
Hero (above the fold). Your name. A one-line tagline that says exactly what you do. A "Get in touch" button with your email. Optionally, a small photo.
Featured project (still above the fold on desktop). Your one anchor project. A screenshot or GIF. A 2-sentence description. Buttons: "Live demo" "GitHub" "Writeup."
Selected work. 3-5 other projects with one-paragraph descriptions and links. Not 20. Five maximum.
About + links. A short bio, your stack, a CV download link, links to LinkedIn / GitHub / Twitter / Hugging Face.
Build it with whatever you know. The simplest path is a Next.js + Tailwind site deployed to Vercel free tier. The framework does not matter; what matters is that loading the home page tells a hiring manager in 10 seconds: this person ships, this is what they ship, here is how to reach them.
Bad portfolio sites that get filtered out: broken images, "coming soon" pages, lorem ipsum placeholder text, more than three blog posts that say "I am learning AI," a "skills" section listing 47 technologies the candidate has touched once. None of these signal anything good.
# [Project Name]
[](https://github.com/USER/REPO/actions)
[](LICENSE)
[](https://your-demo.fly.dev)
> One-sentence tagline. What this does, for whom, and why it's interesting.

## What It Does
A short paragraph (3-5 sentences) explaining the use case, the user, and
the core value. Avoid jargon. A non-ML hiring manager should understand
this paragraph.
## Architecture
- **Frontend:** React + Tailwind, deployed to Vercel
- **API:** FastAPI on Fly.io, autoscale 1-3 instances
- **Retrieval:** Hybrid BM25 (rank_bm25) + sentence-transformers + bge-reranker-base
- **LLM:** GPT-4o-mini with Anthropic Claude 3.5 fallback
- **Storage:** Postgres for documents, Pinecone for embeddings
[Architecture diagram (link)](docs/architecture.png)
## Evaluation
I built a 200-pair eval set for [this domain]. Here is how the system
performs vs three baselines:
| Method | Precision@5 | Recall@10 | MRR |
| ------------------------- | ----------- | --------- | ------ |
| BM25 only | 0.42 | 0.61 | 0.51 |
| Vector only (cosine) | 0.55 | 0.72 | 0.63 |
| Hybrid (BM25 + vector) | 0.68 | 0.83 | 0.74 |
| **Hybrid + reranker** | **0.81** | **0.89** | **0.83** |
The reranker is doing most of the work on rare-entity queries. See the
writeup for failure-mode analysis.
## Quick Start
```bash
git clone https://github.com/USER/REPO
cd REPO
cp .env.example .env # add your API keys
docker compose up
```
Open [http://localhost:8000](http://localhost:8000).
## Tests
```bash
pytest # 47 unit, 8 integration, CI green
```
## Writeup
Long-form blog post: [Why semantic similarity wasn't enough for X](https://your-blog.example.com/post-link)
## License
MIT.
## Contact
If you'd like to discuss this project or hire me, reach me at
[your.email@example.com](mailto:your.email@example.com) or on
[LinkedIn](https://linkedin.com/in/yourhandle).
That template scores 8+ on the rubric on its own, before the underlying code is even good. Most candidates do not write READMEs this clear. Doing so puts you in roughly the top 5% of submissions a hiring manager will see this week.
LinkedIn is not optional. About 60% of recruiters source candidates through LinkedIn search; if your profile is weak or absent, you are invisible to that 60%.
The five LinkedIn elements that matter
Headline. Not "Student" or "Open to opportunities." Write what you actually do in concrete terms. Example: "ML Engineer building RAG systems for legal/clinical text | RugvAI Labs alum | Ex-Backend at [Company]." Specific > vague, every time.
Photo. Recent, clear, professional-but-human. Headshot, not wedding photo, not blurry, not avatar.
About. Three paragraphs maximum: what you do, what you have shipped, what you are looking for. Include your portfolio URL in plain text in the last paragraph.
Featured section. Pin 3 things: your portfolio website, your headline GitHub project, your best blog post. Add custom thumbnails (LinkedIn lets you upload one — most people don't).
Experience. STAR-format bullet points with numbers. Not "Worked on machine learning systems." Yes "Built a hybrid retrieval system serving 8k QPS at p99=180ms, reducing inference cost by 37% vs the prior baseline."
LinkedIn posts as portfolio. Posting about your work, building in public, does two things: (a) it forces you to articulate what you are doing, which deepens your understanding; (b) it builds a public track record that recruiters search. Post 1-2 times per week. Examples:
"Just shipped v0.2 of my legal-doc-rag system. Switched from cosine similarity to hybrid BM25+vector after my eval set showed a 30% precision drop on case names. Here is what I learned: [3 bullet points]. Link in comments."
"Three weeks into rebuilding the Karpathy makemore series in JAX. Surprising lesson: pmap is far more ergonomic than I expected. Here is the diff that converted my training loop. [code snippet]."
Posts with specific technical content get 5-10x the engagement of "I'm learning AI" posts. Recruiters DM the candidates who post specifics.
A single 1,500-word technical blog post that hits the front page of Hacker News or gets 5,000 reads on Substack/Medium can be worth more than three GitHub projects. It demonstrates communication ability — a skill ML interviewers explicitly probe for and most candidates fail on.
The 4-part technical-post structure that works
The problem and why it was harder than expected. (200-400 words.) Specific, concrete. "I tried to do X. I expected Y. The actual difficulty was Z."
The approach. (400-700 words.) Walk through what you tried, with code snippets. Include the failures, not just the successes.
The evaluation. (200-400 words.) A table, a chart, or specific numbers. "Approach A gave me 0.42 precision. Approach B gave me 0.68. Here is why."
Lessons learned and what you would do differently. (200-400 words.) Specific, honest. "Next time I would skip step X. The biggest mistake was Y."
Publish on your own domain (best for SEO), on Substack (best for distribution), or on a personal Hashnode/Hashnode-Medium (also fine). Cross-post to LinkedIn. Cross-post to Hacker News once if relevant. Cross-post to /r/MachineLearning if the topic fits their guidelines.
A 60-90 second screen recording of your project running, embedded in your README and on your portfolio site, is the single highest-ROI artifact you can make.
Why it works: A recruiter who is scanning 200 portfolios in 3 hours is going to spend 30 seconds on yours. A static screenshot earns them 30 seconds of attention. A short video earns 60. A 60-second video that actually demonstrates the product earns the full 90 seconds plus a click into your repo.
How to make one:Loom, Tella, QuickTime + iMovie, or the macOS built-in screen recorder (Cmd-Shift-5). Edit out the dead time, add captions to the key moments, export as an MP4 < 25MB, embed in README via GitHub's video upload, and link to it on your portfolio site.
Most candidates do not do this. The marginal effort is one hour. The marginal upside is significant.
#The Seven Portfolio Mistakes That Nuke Strong Candidates
Pinning tutorial reproductions. Pinning your MNIST tutorial signals you did the tutorial. The hiring manager already assumes you did the tutorial. Unpin it.
README of "WIP" or "Coming soon" on the headline project. Either ship it or take it down. "WIP" reads as "abandoned, never finished" to a hiring manager.
Hardcoded API keys committed in commit history. This is an instant rejection at security-conscious companies (and a real-world security incident). Use git filter-branch or BFG Repo-Cleaner to scrub. And rotate the keys.
No commit history. A single commit titled "initial commit" with 5,000 lines of code reads as a copy-paste, not a build. Commit incrementally. Show the work.
pickle.load() of someone else's model. If your "ML project" downloads a pretrained model and runs inference, that is a wrapper, not an ML project. Either train something yourself (even a small one) or be explicit that the project is "applied AI using a pretrained Claude model", which is fine as long as the wrapper does something interesting.
One single 8,000-line file. A senior engineer's instant signal of "this person has never reviewed real production code." Break into modules. Use a package layout. Add __init__.py.
Linking to a private repo on a public resume. The hiring manager clicks, gets a 404, moves on. If a repo cannot be public, do not link it — describe the work in text instead.
Designing Machine Learning Systems
Chip Huyen (2022)
The book most often referenced in hiring-loop calibration meetings. Chapters on evaluation and deployment overlap with the portfolio rubric used in this lesson. Reading it sharpens both your interview answers and your portfolio's substance.
The rule of 1: one great project beats five mediocre projects, every single time. Optimize your best, not your count.
The 10-point rubric scores any portfolio quickly and tells you exactly what to fix. Aim for 8+. Be honest in scoring.
The first 30 seconds on the README decide whether the hiring manager keeps reading. Tagline + screenshot + demo URL + quickstart, above the fold.
A 4-section portfolio website + an optimized LinkedIn + 3-6 blog posts is the maximum-leverage public footprint. Mirror the story across all three.
A 60-90 second demo video is the highest-ROI artifact most candidates skip. One hour of work; large hiring upside.
Seven mistakes nuke strong candidates: tutorial pins, "WIP" headlines, leaked API keys, no commit history, model-wrapping disguised as ML, one giant file, dead links. Audit and remove all seven.
Your portfolio is a forcing function for honesty. Build the one project that holds up to a 30-second scan AND a 30-minute deep dive, and the interviews will come.