Skip to main content
RugvAI Labs
Show chat and tutor
Reinforcement Learning
Rugv
AI
Labs
Learn
Practice
Reference
Blog
Search...
⌘K
తె
Sign In
RugvAI Labs
All playgrounds →
Q-Table Visualization
Learning rate (alpha)
0.3
Discount (gamma)
0.95
Exploration (epsilon)
0.2
Show policy arrows
Step 0/345
0.5x
1x
2x
4x
0.00
0.00
-0.01
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
WALL
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
WALL
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
GOAL
0.00
0.00
0.00
0.00
Episode:
1
Step:
1
TD Error:
-0.0400
Reward:
-0.04
Total Steps:
346
Bellman Backup — Current Update
Q(s,a)
0.000
+
α
0.30
×
[
r
-0.040
+
γ·max Q(s')
0.000
−
Q(s,a)
0.000
]
=
new Q(s,a)
-0.012
State: (0,0) → (1,0)
Action: ↓ Down
TD Error: -0.0400 (Q was overestimated — decrease)