Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
Di cosa parla
Propongono di rendere le caratteristiche interne dei modelli linguistici quasi perpendicolari tra loro, in modo che cambiare una non alteri le altre. Formalizzano come l'intreccio tra queste caratteristiche produca effetti indesiderati e introducono un vincolo che riduce tale interferenza: nei test questo permette di intervenire in modo più isolato su concetti di ragionamento matematico mantenendo le prestazioni.
Cosa permette di osservare
Consente di esplorare se è possibile progettare la struttura interna dei modelli per rendere più semplici e meno intrecciati gli interventi sui concetti che usano, e come ciò influisce sulla loro capacità di risolvere compiti di ragionamento matematico.
Dalla fonte
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. In practice, however, feature entanglement leads to interference such that localized interventions can have unintended downstream effects. Motivated by the _Independent Causal Mechanisms_ principle, we propose to constrain internal features to be almost orthogonal. We argue that this promotes modular representations amenable to causal intervention. We formalize this problem by characterizing the gap between an idealized isolated intervention and its realized effect on model outputs in terms of feature interference. We upper-bound the propagation of feature interference in terms of the self-coherence of the…