Implement GRPO (Group Relative Policy Optimization) fine-tuning for vision-language models on small datasets. Use when …
Implement GRPO (Group Relative Policy Optimization) fine-tuning for vision-language models on small datasets. Use when SFT underperforms or training data is limited (<1000 examples).
aws-solutions-library-samples
cli
free
Others in the same category, ranked by how often they are opened.