Fine-Tuning LLMs Shift Representations, Not Causal Heads
Researchers analyze how fine‑tuning reshapes LLM internals and find that EAP‑identified components stay confined to specific layers, unrelated to the layers with the biggest representational shifts.

## Overview A recent arXiv paper (arXiv:2609.21113v1) investigates the internal impact of fine‑tuning on large language models (LLMs). While fine‑tuning is a common method for adapting LLMs to downstream tasks, its effect on attention patterns, layer‑wise activations, and other internal mechanisms has been unclear.
## Key Findings The authors examine whether changes in internal representations align with task‑relevant components highlighted by the Explainable AI Prompt (EAP) framework, such as specific attention heads and logit‑level activations. They discover that EAP‑identified components cluster within particular layers, suggesting a functional localisation of task‑specific behavior.
## Representation vs. Causal Importance Crucially, the distribution of these EAP components shows little correlation with the layers that experience the most pronounced representational changes during fine‑tuning. This decoupling indicates that the layers most altered by fine‑tuning are not necessarily the ones driving task performance.
## Implications The study highlights a nuanced view of model adaptation: fine‑tuning reshapes internal dynamics without directly shifting the causal elements that underpin task success. Understanding this separation could guide more efficient fine‑tuning strategies and better interpretability tools.
*Source: arXiv cs.AI (2026‑09‑21)*
Read also

RBS-Attention Cuts Prefill Cost for Long-Context LLMs
RBS-Attention offers a training‑free, dual‑branch sparse prefill method that mitigates mean dilution while preserving block‑sparse FlashAttention performance.

New Platform Empowers Clinicians to Evaluate AI Psychiatric Intake Tools
A novel evaluation platform called InterviewPlayground enables clinicians to benchmark AI‑assisted psychiatric intake systems against real‑world clinical standards.

DeepMind Unveils Institute to Expand AGI Dialogue
DeepMind has created an institute to foster a broader conversation on AGI, highlighting potential disagreements and evolving views within the field.