Can Human Perceptual Similarity Alignment Improve Interpretability in Visual Representations?
Authors: Colin, J. , Oliver, N. , Serre, T
Publication: Conference on Cognitive Computational Neuroscience (CCN), 2026
PDF: Click here for the PDF paper
What makes the features learned by vision models interpretable to humans? We introduce a psychophysics protocol that directly measures feature interpretability—the degree to which a person, given a feature visualization and importance maps, can predict where that feature will activate in a new image. Applying this protocol to sparse autoencoder features from a range of vision transformers, we find that self-supervised and foundation models are consistently less interpretable than their supervised counterparts. While fine-grained alignment with human similarity judgments is uniformly high across models and does not predict interpretability, coarse-grained alignment—capturing semantic rather than perceptual structure—is a stronger correlate, suggesting that further improvements in interpretability depend less on perceptual fidelity than on whether a model's representations reflect the semantic organization of human perception.