Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing R…
Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
davila7
cli
free
Others in the same category, ranked by how often they are opened.