Provides guidance for training and analyzing Sparse Autoencoders (SAEs) using SAELens to decompose neural network activ…
Provides guidance for training and analyzing Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features. Use when discovering interpretable features, analyzing superposition, or studying monosemantic representations in language models.
davila7
cli
free
Others in the same category, ranked by how often they are opened.