cp: fix: fix mcore train_iters in grpo (1383) into r0.4.0 - #1385
Conversation
Signed-off-by: Yuki Huang <yukih@nvidia.com> Signed-off-by: NeMo Bot <nemo-bot@nvidia.com>
ð WalkthroughWalkthroughModifies total_train_iters calculation in GRPO algorithm to incorporate epoch count: changes from Changes
Estimated code review effortðŊ 2 (Simple) | âąïļ ~8 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Pre-merge checks and finishing touchesâ Failed checks (1 warning)
â Passed checks (3 passed)
âĻ Finishing touches
ð§Š Generate unit tests (beta)
ð Recent review detailsConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro ð Files selected for processing (1)
ð§° Additional context usedð Path-based instructions (2)**/*.pyð CodeRabbit inference engine (CODING_GUIDELINES.md)
Files:
nemo_rl/**/*.pyð CodeRabbit inference engine (CODING_GUIDELINES.md)
Files:
ð§Ž Code graph analysis (1)nemo_rl/algorithms/grpo.py (1)
â° Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
ð Additional comments (1)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
beep boop [ðĪ]: Hi @yuki-97 ð,
Summary by CodeRabbit