Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with sel…
Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system.
davila7
cli
free
Others in the same category, ranked by how often they are opened.