Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models
Di cosa parla
Indagano perché i modelli linguistici tendono a produrre risposte simili e poco varie. Tramite esperimenti controllati mostrano che questa omogeneità è spesso appresa già nella fase iniziale di addestramento e che la successiva fase di adattamento alle istruzioni la mette in luce o la amplifica, ma non la crea. Mostrano inoltre che basta porre certe richieste al modello non adattato per far riemergere la stessa tendenza.
Cosa permette di osservare
Permette di esplorare se la somiglianza delle risposte è una proprietà intrinseca dell'addestramento iniziale e quanto gli interventi successivi possano davvero cambiarla, o se basta cambiare le richieste per farla emergere.
Dalla fonte
The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that output homogeneity is likely learned during the pretraining phase, and only \emph{revealed} or magnified during the alignment process. Specifically, we find that semantic convergence is observed from the first alignment stage--the instruction-tuning phase (SFT)--suggesting that homogeneity might already exist in the pre-alignment model. To investigate this, we conduct controlled SFT experiments examining how training data influences output convergence on specific input/output pairs. We find that convergence can be revealed and amplified, but not introduced by the SFT data, supporting its role as a catalyst rather than a cause. To further test whether homogeneity originates before alignment, we measure convergence in b…