Sparse autoencoder (SAE)
AISparse autoencoder (SAE): An interpretability tool that decomposes a model's dense activations into many sparse features, each hopefully corresponding to a single human-understandable concept.
Sparse autoencoder (SAE): An interpretability tool that decomposes a model's dense activations into many sparse features, each hopefully corresponding to a single human-understandable concept.