Git Basics: History, Branches and Keeping Secrets Out
After this lesson, you will be able to:
- Describe the working tree, the staging area and the commit history, and say which git command moves a change between them
- Create a repository, then use git status, git add, git commit, git diff and git log to save and inspect your work
- Undo an edit with git restore and a saved commit with git revert, and say what git reset --hard destroys
- Merge a branch, and resolve a merge conflict by hand
- Write a .gitignore that keeps .env, .venv, __pycache__ and data files out, and explain what to do when a secret was committed
Before You Start
#The Problem With script_final_v3_REAL.py
train.py for a week. Every time you were afraid of breaking it, you copied it first. Here is the folder.$ ls
train_final_v3_REAL_use_this_one.py
train_final_v3_REAL.py
train_final.py
train_v2.py
train.py
Which one produced the accuracy you wrote in your notes on Tuesday? Which one has the learning rate fix? If you open two of them, you have to compare them by eye. If a teammate sends you a sixth copy, you now have two lines of history to merge, and nothing records who changed what.
The commands below were run in a scratch folder with git 2.54.0. The paths are shortened, and the commit ids (the seven-character codes) will differ on your machine because they are computed from your files, your name and the time. Every session in this lesson is a transcript you read, since the browser cannot run git. You will run the same commands in the exercise at the end.
#Three Places Your Work Lives
The idea that makes git make sense is that a change passes through three places on its way to being saved.
How a change moves between the three places
git status tells you.#Your First Repository
Git needs to know who is making commits. Do this once on your machine (use your own name and email).
# @read-only
git config --global user.name "Your Name"
git config --global user.email "you@example.com"
git init inside your project folder. The -b main part names the first branch main; it needs git 2.28 or newer, and on an older git you run git init and then git branch -M main. Here train.py already exists, with four lines of settings.$ git init -b main
Initialized empty Git repository in /home/you/churn-model/.git/
$ git status
On branch main
No commits yet
Untracked files:
(use "git add <file>..." to include in what will be committed)
train.py
nothing added to commit but untracked files present (use "git add" to track)
git status is the command you will run most. It says which branch you are on, and it sorts every file into a group. train.py is untracked: it is in the folder, but git has not been told to watch it.You run git init in a folder that already contains train.py, and then git status straight away. What does it show for train.py?
git add, check, and then save the commit.$ git add train.py
$ git status
On branch main
No commits yet
Changes to be committed:
(use "git rm --cached <file>..." to unstage)
new file: train.py
$ git commit -m "Add training script with default settings"
[main (root-commit) dfa7fd5] Add training script with default settings
1 file changed, 4 insertions(+)
create mode 100644 train.py
$ git status
On branch main
nothing to commit, working tree clean
$ git log --oneline
dfa7fd5 Add training script with default settings
git add, the file moved from untracked to Changes to be committed: it is in the staging area. After git commit, the working tree is clean, which means the folder matches the last commit exactly. git log --oneline lists the history, newest first, one line per commit.Writing a good commit message
git log in three months, and that is usually you. Write one short line in the imperative that says what the commit does: "Train for 20 epochs", "Fix off-by-one in the batch loop". Say why in the line if the reason is not obvious. "update", "stuff" and "final" tell nobody anything, and they bring you straight back to train_final_v3_REAL.py.git commit with no -m, git opens a text editor for the message. Type it, save and close the editor; in vim that is :wq.#Seeing and Staging Changes
train.py. git status now reports a modified file, and git diff shows what changed, line by line.$ git status
On branch main
Changes not staged for commit:
(use "git add <file>..." to update what will be committed)
(use "git restore <file>..." to discard changes in working directory)
modified: train.py
no changes added to commit (use "git add" and/or "git commit -a")
$ git diff
diff --git a/train.py b/train.py
index 73a81a4..8ba0d47 100644
--- a/train.py
+++ b/train.py
@@ -1,4 +1,4 @@
LEARNING_RATE = 0.1
-EPOCHS = 10
+EPOCHS = 20
BATCH_SIZE = 32
print(f"lr={LEARNING_RATE} epochs={EPOCHS} batch={BATCH_SIZE}")
- was removed and one that starts with + was added. The unchanged lines around it are context. Now stage the file and run both kinds of diff.$ git add train.py
$ git diff
$ git diff --staged
diff --git a/train.py b/train.py
index 73a81a4..8ba0d47 100644
--- a/train.py
+++ b/train.py
@@ -1,4 +1,4 @@
LEARNING_RATE = 0.1
-EPOCHS = 10
+EPOCHS = 20
BATCH_SIZE = 32
print(f"lr={LEARNING_RATE} epochs={EPOCHS} batch={BATCH_SIZE}")
$ git commit -m "Train for 20 epochs"
[main ffccf07] Train for 20 epochs
1 file changed, 1 insertion(+), 1 deletion(-)
git diff printed nothing after the git add. It compares the working tree with the staging area, and they now agree. git diff --staged compares the staging area with the last commit, which is exactly what the next commit will contain. Run it before you commit, as a last look.You edit train.py, run git add train.py, and then run git diff with no options. What does it print?
git commit -am "message" does it. It does not pick up new files, and it skips the look at what you are about to save, so use it when the change is small and you know what it is.#Undoing a Mistake Safely
EPOCHS = 2000 and you have not committed it. git restore puts the file back to what the staging area holds, which here is the last commit.$ git diff
diff --git a/train.py b/train.py
index 8ba0d47..26fd560 100644
--- a/train.py
+++ b/train.py
@@ -1,4 +1,4 @@
LEARNING_RATE = 0.1
-EPOCHS = 20
+EPOCHS = 2000
BATCH_SIZE = 32
print(f"lr={LEARNING_RATE} epochs={EPOCHS} batch={BATCH_SIZE}")
$ git restore train.py
$ git diff
$ cat train.py
LEARNING_RATE = 0.1
EPOCHS = 20
BATCH_SIZE = 32
print(f"lr={LEARNING_RATE} epochs={EPOCHS} batch={BATCH_SIZE}")
git restore cannot bring it back. To take a file out of the staging area while keeping your edits, use git restore --staged train.py.git revert.$ git commit -m "Shrink batch size"
[main 7daeb11] Shrink batch size
1 file changed, 1 insertion(+), 1 deletion(-)
$ git revert HEAD --no-edit
[main 4fdf46a] Revert "Shrink batch size"
Date: Mon Oct 5 01:42:03 2026 +0530
1 file changed, 1 insertion(+), 1 deletion(-)
$ git log --oneline
4fdf46a Revert "Shrink batch size"
7daeb11 Shrink batch size
ffccf07 Train for 20 epochs
dfa7fd5 Add training script with default settings
HEAD means the commit you are standing on, here the latest. The bad commit is still in the log, and the new one cancels it. That is why git revert is the safe way to undo something other people may already have.git reset --hard destroys uncommitted work
git reset --hard in many answers online. It makes the working tree and the staging area match a commit exactly, and it throws away everything that differs, with no confirmation.$ git status --short
M train.py
$ git reset --hard
HEAD is now at 4fdf46a Revert "Shrink batch size"
$ git status --short
$ cat train.py
LEARNING_RATE = 0.1
EPOCHS = 20
BATCH_SIZE = 32
print(f"lr={LEARNING_RATE} epochs={EPOCHS} batch={BATCH_SIZE}")
train.py was never saved anywhere, so it is gone for good. Commits are more forgiving: if you reset past a commit, git reflog still lists it for a while, and you can often get it back. Uncommitted edits have no such safety net. Until you are sure what a command will discard, prefer git restore on one named file, or git revert.#Branches and Merging
git switch -c add-evaluate creates a branch and moves you onto it.$ git switch -c add-evaluate
Switched to a new branch 'add-evaluate'
$ git add evaluate.py
$ git commit -m "Add evaluate script"
[add-evaluate 40c2131] Add evaluate script
1 file changed, 1 insertion(+)
create mode 100644 evaluate.py
$ git switch main
Switched to branch 'main'
$ ls
train.py
$ git merge add-evaluate
Updating 4fdf46a..40c2131
Fast-forward
evaluate.py | 1 +
1 file changed, 1 insertion(+)
create mode 100644 evaluate.py
$ ls
evaluate.py
train.py
$ git branch -d add-evaluate
Deleted branch add-evaluate (was 40c2131).
main, evaluate.py disappeared from the folder, because it exists only on the other branch. Merging brought it in. Nothing had happened on main in the meantime, so git simply moved main forward, which it calls a fast-forward. git branch -d removes the branch name once its work is merged; git refuses with a warning if the branch has commits that are not merged anywhere.#A merge conflict, step by step
main sets it to 0.05.$ git merge lower-lr
Auto-merging train.py
CONFLICT (content): Merge conflict in train.py
Automatic merge failed; fix conflicts and then commit the result.
$ git status
On branch main
You have unmerged paths.
(fix conflicts and run "git commit")
(use "git merge --abort" to abort the merge)
Unmerged paths:
(use "git add <file>..." to mark resolution)
both modified: train.py
no changes added to commit (use "git add" and/or "git commit -a")
$ cat train.py
<<<<<<< HEAD
LEARNING_RATE = 0.05
=======
LEARNING_RATE = 0.01
>>>>>>> lower-lr
EPOCHS = 20
BATCH_SIZE = 32
print(f"lr={LEARNING_RATE} epochs={EPOCHS} batch={BATCH_SIZE}")
<<<<<<< HEAD and ======= is your side, the current branch. The part between ======= and >>>>>>> lower-lr is the other branch's side. Nothing is lost, and nothing is saved either: the merge is paused.You resolve it by editing the file. Decide what the line should be, delete the three marker lines and everything you do not want, and save. Here the decision is to keep 0.01. Then tell git the file is resolved by staging it, and finish the merge.
$ cat train.py
LEARNING_RATE = 0.01
EPOCHS = 20
BATCH_SIZE = 32
print(f"lr={LEARNING_RATE} epochs={EPOCHS} batch={BATCH_SIZE}")
$ git add train.py
$ git commit --no-edit
[main 8337437] Merge branch 'lower-lr'
$ git log --oneline --graph
* 8337437 Merge branch 'lower-lr'
|\
| * 619dcc0 Lower learning rate to 0.01
* | 720b36d Lower learning rate to 0.05
|/
* 40c2131 Add evaluate script
* 4fdf46a Revert "Shrink batch size"
* 7daeb11 Shrink batch size
* ffccf07 Train for 20 epochs
* dfa7fd5 Add training script with default settings
git commit would open the editor with the message "Merge branch 'lower-lr'" already filled in; --no-edit accepts it. The --graph view shows the two lines of work splitting and joining. If you panic in the middle of a conflict, git merge --abort returns you to where you were before the merge.<<<<<<<, git commits it, and your code is broken. Open the file and search for <<<<<<< before you run git add.#Keeping Files Out With .gitignore
pip can rebuild. __pycache__ holds compiled files Python regenerates. Data files can be large and change often. And .env holds secrets. Here is git status in a folder that has all four, before any rule exists.$ git status
On branch main
Untracked files:
(use "git add <file>..." to include in what will be committed)
.env
.venv/
__pycache__/
data/
nothing added to commit but untracked files present (use "git add" to track)
git add . ("add everything here") would stage all of it, secrets included. A file called .gitignore tells git to leave named things alone. Each line is a pattern: a name matches at any depth, and a trailing slash means a folder.$ cat .gitignore
.env
.venv/
__pycache__/
data/
$ git status
On branch main
Untracked files:
(use "git add <file>..." to include in what will be committed)
.gitignore
nothing added to commit but untracked files present (use "git add" to track)
$ git check-ignore -v .env data/train.csv
.gitignore:1:.env .env
.gitignore:4:data/ data/train.csv
.gitignore is new. You commit .gitignore itself, so everyone who clones the project gets the same rules. git check-ignore -v is the tool for asking "which rule hid this file?", and it names the file and line number.*.csv matches every CSV file anywhere. A line starting with ! makes an exception. A line starting with # is a comment. Keep a .env.example in the repository with the variable names and fake values, so a new teammate knows what to fill in, and ignore only the real .env.This script mimics a simplified .gitignore and prints the files that would be committed. The rules are wrong: .env, .venv, __pycache__ and data are all still going in, because each rule is spelled slightly differently from the real name. Fix the rules so that only .env.example and train.py are printed. (A shortcut like .env* also hides .env.example, which you want to keep.)
.env.example train.py
A rule does not untrack a file git already tracks
.gitignore only affects files that are untracked. If you committed .env first and wrote the rule afterwards, git keeps tracking it and reports every edit.$ git status --short
M .env
?? .gitignore
$ git ls-files
.env
a.py
b.py
git ls-files lists the files git tracks, and .env is still on it. To stop tracking a file without deleting it from your folder, run git rm --cached .env, then commit. The next section shows why that is not enough for a secret.#When a Secret Gets Committed
git add ., commit, notice .env was in there, remove it, add the rule, and commit again.$ git add .
$ git status --short
A .env
A app.py
$ git commit -m "First version"
[main (root-commit) 6e40fe7] First version
2 files changed, 2 insertions(+)
create mode 100644 .env
create mode 100644 app.py
$ git rm --cached .env
rm '.env'
$ git add .gitignore
$ git commit -m "Stop tracking .env"
[main c7b7191] Stop tracking .env
2 files changed, 1 insertion(+), 1 deletion(-)
delete mode 100644 .env
create mode 100644 .gitignore
$ ls -a
.
..
.env
.git
.gitignore
app.py
$ git show HEAD~1:.env
API_KEY=sk-FAKE-not-a-real-key-123456
HEAD~1 means "one commit before the latest", and git show HEAD~1:.env prints the file as it was in that commit. The key is right there. The file is untracked now and covered by the ignore rule, yet the earlier commit still contains it, and so does every copy of the repository that anyone made.So when a secret has been committed, do these things in this order.
- Rotate it. Go to the service that issued the key, revoke it and create a new one. This is the only step that actually makes the old key useless.
- Stop tracking the file with
git rm --cached, add the rule to.gitignore, and commit, so the new key is never committed. - Assume the old key was seen if the repository was ever pushed anywhere or shared. Deleting a file, or a commit, does not remove it from history, and tools that rewrite history do not remove copies other people already hold.
.gitignore before the first git add. The next lesson shows how your code reads the key from the environment instead.#Sharing: clone, remote and push
https://github.com/you/churn-model.git, and the commands are identical.$ git remote add origin ../origin-demo.git
$ git remote -v
origin ../origin-demo.git (fetch)
origin ../origin-demo.git (push)
$ git push -u origin main
To ../origin-demo.git
* [new branch] main -> main
branch 'main' set up to track 'origin/main'.
$ git branch -vv
lower-lr 619dcc0 Lower learning rate to 0.01
* main 096fe94 [origin/main] Add .gitignore
origin is the conventional name of the main remote. git push sends your commits to it. The -u option records that your main follows origin/main, so later a plain git push and git pull know where to go. The brackets in git branch -vv show that link.git clone, which copies the whole history and sets up origin for them.$ git clone origin-demo.git teammate-copy
Cloning into 'teammate-copy'...
done.
$ git log --oneline
096fe94 Add .gitignore
8337437 Merge branch 'lower-lr'
720b36d Lower learning rate to 0.05
619dcc0 Lower learning rate to 0.01
40c2131 Add evaluate script
4fdf46a Revert "Shrink batch size"
7daeb11 Shrink batch size
ffccf07 Train for 20 epochs
dfa7fd5 Add training script with default settings
.env, .venv or your data, because those were never committed. That is the point of the ignore rules.Now the teammate pushes a commit, and you make one of your own and push.
$ git push
To ../origin-demo.git
! [rejected] main -> main (fetch first)
error: failed to push some refs to '../origin-demo.git'
hint: Updates were rejected because the remote contains work that you do not
hint: have locally. This is usually caused by another repository pushing to
hint: the same ref. If you want to integrate the remote changes, use
hint: 'git pull' before pushing again.
hint: See the 'Note about fast-forwards' in 'git push --help' for details.
git pull, which fetches the remote commits and merges them into your branch, exactly as in the branch section. If the same lines changed you get a conflict, and you resolve it the same way. When both sides have new commits, a plain git pull on this git version stops with fatal: Need to specify how to reconcile divergent branches until you choose a method. git pull --no-rebase chooses the merge used in this lesson. Git may open an editor for the merge message, and you save and close it.$ git pull --no-rebase
Merge made by the 'ort' strategy.
NOTES.md | 1 +
1 file changed, 1 insertion(+)
create mode 100644 NOTES.md
git push goes through. git fetch downloads remote commits without merging, which is useful when you want to look first.#Common Mistakes
#Try It: Your First Safe Repository
git init prints your own folder path.-
Make a project folder and a repository.
bash# @read-only mkdir ml-notes cd ml-notes git init -b mainExpected:Initialized empty Git repository in .../ml-notes/.git/ -
In your editor, create
score.pycontainingprint("score: 0.82")and.envcontainingTOKEN=not-a-real-token. Then rungit status --short.text$ git status --short ?? .env ?? score.pyThe secret is sitting there waiting for a carelessgit add .. -
Create
.gitignorecontaining the single line.env, and rungit status --shortagain.text$ git status --short ?? .gitignore ?? score.py -
Stage the two files by name, check, and commit.
text$ git add .gitignore score.py $ git status --short A .gitignore A score.py $ git commit -m "Add score script and gitignore" [main (root-commit) 758822b] Add score script and gitignore 2 files changed, 2 insertions(+) create mode 100644 .gitignore create mode 100644 score.py -
Change the file to
print("score: 0.91"), then look at the change and commit it with-am.text$ git diff diff --git a/score.py b/score.py index 6916804..985776a 100644 --- a/score.py +++ b/score.py @@ -1 +1 @@ -print("score: 0.82") +print("score: 0.91") $ git commit -am "Update score" [main d87d1c3] Update score 1 file changed, 1 insertion(+), 1 deletion(-) -
Check the history and the state.
text$ git log --oneline d87d1c3 Update score 758822b Add score script and gitignore $ git status On branch main nothing to commit, working tree clean
.env never appeared in git status after step 3, and it will not in any later commit. Going further: create a branch with git switch -c, change the same line on main and on the branch, and resolve the conflict you cause.#Recap
git add, and into the history with git commit. git status tells you where every file is, git diff shows unstaged edits and git diff --staged shows what the next commit holds, and git log --oneline lists the history.git restore, and a saved commit with git revert, which adds a new commit that cancels it. git reset --hard discards uncommitted work for good. A branch lets you try something apart from main, and git merge brings it back, stopping for you to edit and stage the file when two branches changed the same lines..gitignore written before the first commit keeps .env, .venv, __pycache__ and data out. A committed secret stays in history even after you delete the file, so rotate it. git clone, git remote and git push move commits between copies, and you never force-push history other people share.You edit a.py, run git add a.py, and then edit a.py again before committing. What does git status --short print for it?