Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

Title:Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

Abstract:Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential requires that the learned features be both interpretable and steerable. To that end, we introduce two new computationally inexpensive interpretability and steerability metrics and conduct a systematic analysis on LVLMs. Our analysis uncovers two observations; (i) a majority of SAE neurons exhibit either low interpretability or low steerability or both, rendering them ineffective for downstream use; and (ii) due to…

Title:Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

View PDF HTML (experimental)

Abstract:Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential requires that the learned features be both interpretable and steerable. To that end, we introduce two new computationally inexpensive interpretability and steerability metrics and conduct a systematic analysis on LVLMs. Our analysis uncovers two observations; (i) a majority of SAE neurons exhibit either low interpretability or low steerability or both, rendering them ineffective for downstream use; and (ii) due to the unsupervised nature of SAEs, user-desired concepts are often absent in the learned dictionary, thus limiting their practical utility. To address these limitations, we propose Concept Bottleneck Sparse Autoencoders (CB-SAE) - a novel post-hoc framework that prunes low-utility neurons and augments the latent space with a lightweight concept bottleneck aligned to a user-defined concept set. The resulting CB-SAE improves interpretability by +32.1% and steerability by +14.5% across LVLMs and image generation tasks. We will make our code and model weights available.


Subjects:	Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2512.10805 [cs.LG]
	(or arXiv:2512.10805v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2512.10805 arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Akshay Kulkarni [view email] [v1] Thu, 11 Dec 2025 16:48:07 UTC (5,113 KB)

Title:Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

Title:Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

Submission history

Similar Posts