hands-on-modern-rl β tool profile and history
π An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
- Language
- Python
- Created
- Last activity
- Topics
- agent
- agentic
- agentic-ai
- agentic-rl
- dpo
- grpo
- llm
- llm-alignment
- pytorch
- reinforcemen
- rlhf
- tutorial
Current measurement
3β―948GitHub starsβ
estimated momentumMomentum uses available readings from the last 7 days: measured with at least two comparable readings, estimated otherwise.
Reading history
| Date | Value | Metric |
|---|---|---|
| 3β―948 | GitHub stars |
Classifications
- Domains
- Type
- Agents
- Use cases
- No declared use case
Neighbouring tools
These entries declare the same topics. No similarity is inferred: only the shared topics are stated.