Mechanistic interpretability
AIMechanistic interpretability: A branch of interpretability that reverse-engineers the internal weights and activations of a network into human-understandable algorithms and circuits.
Mechanistic interpretability: A branch of interpretability that reverse-engineers the internal weights and activations of a network into human-understandable algorithms and circuits.