YFarmX

Mechanistic interpretability

AI

Mechanistic interpretability: A branch of interpretability that reverse-engineers the internal weights and activations of a network into human-understandable algorithms and circuits.

Related terms

Browse the full glossary →