FeatCal: Feature Calibration for Post-Merging Models

ABSTRACT

Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often still underperforms task experts. We study this performance gap through *feature drift*, the difference between features produced by the merged model and by the expert on the same input. Our theory decomposes this drift into upstream propagation and local mismatch, tracks how it propagates and combines through later layers in forward order, and links final feature drift to output drift. This view motivates FeatCal, which uses a small calibration set to calibrate the merged model weights layer by layer in forward order, reducing feature drift while staying close to merged weights and preserving the benefits of model merging. FeatCal uses an efficient closed-form solution to update model weights, with no gradient descent, iterative optimization, or extra modules. On the main CLIP and GLUE benchmarks, FeatCal beats Surgery and ProbSurgery, the closest post-merging calibration baselines: 85.5% vs. 77.0%/78.8% on CLIP-ViT-B/32 Task Arithmetic (TA) and 85.2% vs. 83.7%/82.2% on FLAN-T5-base GLUE. On CLIP-ViT-B/32, 8 examples per task reach 82.9%, and 256 examples per task take 53 seconds, about 4x faster than both baselines, showing better sample efficiency and lower calibration cost.

Model MergingPost-Merging CalibrationFeature Drift
May 13, 2026
Yanggan Gu, Shuo Cai, Zihao Wang, Wenjun Wang, Yuanyi Wang, Pengkai Wang, Sirui Huang, Su Lu, Jianmin Wu, Hongxia Yang

Abstract

Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often still underperforms task experts. We study this performance gap through feature drift, the difference between features produced by the merged model and by the expert on the same input. Our theory decomposes this drift into upstream propagation and local mismatch, tracks how it propagates and combines through later layers in forward order, and links final feature drift to output drift. This view motivates FeatCal, which uses a small calibration set to calibrate the merged model weights layer by layer in forward order, reducing feature drift while staying close to merged weights and preserving the benefits of model merging. FeatCal uses an efficient closed-form solution to update model weights, with no gradient descent, iterative optimization, or extra modules. On the main CLIP and GLUE benchmarks, FeatCal beats Surgery and ProbSurgery, the closest post-merging calibration baselines: 85.5% vs. 77.0%/78.8% on CLIP-ViT-B/32 Task Arithmetic (TA) and 85.2% vs. 83.7%/82.2% on FLAN-T5-base GLUE. On CLIP-ViT-B/32, 8 examples per task reach 82.9%, and 256 examples per task take 53 seconds, about 4x faster than both baselines, showing better sample efficiency and lower calibration cost.

Figure 1. Feature drift after Task Arithmetic (TA) merging and FeatCal calibration in CLIP-ViT-B/32. Panels (a,b) use Stanford Cars: FeatCal moves features toward expert features, raises their mean cosine similarity from 0.60 to 0.84, and reduces mean L2 feature drift, with 46% less final-layer drift. Panel (c) reports per-task accuracy in the 8-task setting. Appendix B gives full 8-task feature views.

1. Introduction

Model merging composes task experts into one model, avoiding joint training, retraining, or deploying a separate model for each task. However, the merged model often still underperforms the experts it is intended to combine. This leaves a clear performance gap after merging in practice.

We analyze this gap through feature drift, the difference between features of the merged model and the task expert on the same input sample. Our layer by layer analysis decomposes this drift into upstream propagation and local mismatch at expert input features, then shows how local mismatches propagate through later layers in forward order and combine into final feature drift. We further analyze how final feature drift reaches the model output and becomes output drift, explaining how feature drift can affect output scores. Figure 1 illustrates this behavior on CLIP-ViT-B/32 merged with Task Arithmetic (TA), where final features move away from the Stanford Cars expert features and drift appears across multiple layers of the network after TA merging.

Surgery and ProbSurgery are related post-merging calibration methods: they identify feature drift at the final layer and train extra modules for calibration. Other intervention methods point to the same lesson: expert signals can help, but current methods often rely on task specific intervention parameters, extra modules at inference, or iterative optimization. These choices make calibration less direct for an already merged model and can leave the final model with an auxiliary inference path. This motivates efficient, direct calibration that preserves inference speed.

FeatCal uses our drift analysis as a design cue for calibrating the merged model. It uses a forward-order schedule suggested by the propagation view and calibrates the model layer by layer. We introduce a regularization term that keeps the calibrated model close to the merged model, which helps preserve the benefits of model merging and reduce overfitting to the small calibration set. The resulting objective has an efficient closed form solution, so calibration needs no gradient descent, iterative optimization, or extra modules at inference. We further introduce feature interpolation and anchor regularization to balance expert signals with the merged model and improve performance.

Empirically, Figure 1 shows the practical effect: FeatCal moves merged features toward expert features, reduces feature drift, and raises per-task accuracy after TA merging. On CLIP-ViT-B/32 8-task TA, it reaches 85.5% versus 77.0%/78.8% for Surgery/ProbSurgery, and on FLAN-T5-base GLUE it reaches 85.2% versus 83.7%/82.2%. The same trend holds on CLIP-ViT-L/14 WUDI, FLAN-T5-large, and MergeBench Llama-family LLM merging, where FeatCal improves TA by +2.0+2.0/+2.3+2.3 average points on 3B/8B models. On CLIP-ViT-B/32 TA, 8 examples per task reach 82.9%, and calibration with 256 examples per task takes 53 seconds, about 4x faster than both baselines under the same calibration protocol.

We summarize the main contributions of this work as follows: ① We develop a theory of feature drift after merging, with an exact decomposition into local mismatch and upstream propagation, forward order propagation, and a link to output drift. ② We introduce FeatCal, which efficiently calibrates merged model weights in forward order with closed form updates, without gradient descent, architecture changes, or extra modules at inference. ③ We validate FeatCal on CLIP, FLAN-T5, and MergeBench LLM benchmarks, where it outperforms related post-merging calibration baselines while using fewer samples and lower calibration cost without adding inference-time modules.

Model Merging. Most model merging methods build a fused model by merging task experts directly in parameter space. Methods based on feature statistics or feature drift, such as RegMean, RegMean++, and LOT Merging, are closer in mechanism: they derive layer updates during merging from feature statistics, regression objectives, or an explicit feature drift objective. These methods define how to build the merged model. In contrast, FeatCal treats that model as the starting point, uses a small calibration set and task experts for post-merging calibration, and controls calibration strength through regularization. This stage separation matters because calibration must work with the feature drift left by a chosen merger instead of changing the merge rule itself. It helps preserve the benefits of model merging while reducing the risk of overfitting to calibration data.

Post-Merging Feature Calibration. Representation Surgery and follow-up methods operate on merged-model features through task-specific plugins, deeper interventions, probabilistic feature-drift modeling, or parameter-efficient modules. These closest post-merging alternatives establish that expert-guided feature calibration is useful. FeatCal differs in parameterization and deployment: instead of learning or deploying auxiliary intervention modules, it folds the calibration into the original linear module weights through closed-form regularized updates, leaving a single architecture-preserving model at inference time rather than an auxiliary intervention path.

3. Post-Merging Feature Drift: Problem Formulation and Properties

Before introducing FeatCal, we formalize post-merging feature drift and show how local mismatch is defined at expert input features, then propagated and combined in forward order across depth.

3.1 Layer-Wise Feature Drift

We consider NN task experts {Miexpert}i=1N\{M_i^{\expert}\}_{i=1}^N and a merged model MmergeM^{\merge}. As in standard weight-space merging, experts are fine-tuned from a common pretrained base. All models share the same architecture and contain LL layers. Let Di\mathcal D_i be the data distribution of task ii. For layer ℓ\ell, fi,ℓexpertf_{i,\ell}^{\expert} and fℓmergef_{\ell}^{\merge} denote the corresponding layer functions of the task expert and the merged model, respectively.

For an input sample xx from task ii, define the expert and merged layer output features recursively as

hi,0expert(x)=x,hi,ℓexpert(x)=fi,ℓexpert ⁣(hi,ℓ−1expert(x)),hi,0merge(x)=x,hi,ℓmerge(x)=fℓmerge ⁣(hi,ℓ−1merge(x)),ℓ=1,…,L. \begin{aligned} h_{i,0}^{\expert}(x)&=x, & h_{i,\ell}^{\expert}(x)&= f_{i,\ell}^{\expert}\!(h_{i,\ell-1}^{\expert}(x)), \\ h_{i,0}^{\merge}(x)&=x, & h_{i,\ell}^{\merge}(x)&= f_{\ell}^{\merge}\!(h_{i,\ell-1}^{\merge}(x)), \end{aligned} \qquad \ell=1,\dots,L.

Definition 1. [Layer-wise feature drift]

For task ii, sample xx, and layer ℓ\ell, the layer-wise feature drift is the difference between the merged-model feature and the corresponding task-expert feature at that layer:

ei,ℓ(x)=hi,ℓmerge(x)−hi,ℓexpert(x). e_{i,\ell}(x) = h_{i,\ell}^{\merge}(x)-h_{i,\ell}^{\expert}(x).

This pointwise drift signal is the object propagated in forward order in the analysis below.

3.2 Local Mismatch and Drift Propagation

Proposition 1. [Exact layer-wise drift decomposition]

For every task ii, sample xx, and intermediate layer ℓ\ell, the drift decomposes as follows. By Eq. 1, pi,1(x)=0p_{i,1}(x)=0.

ei,ℓ(x)=fℓmerge ⁣(hi,ℓ−1expert(x))−fi,ℓexpert ⁣(hi,ℓ−1expert(x))⏟local mismatch mi,ℓ(x)+fℓmerge ⁣(hi,ℓ−1merge(x))−fℓmerge ⁣(hi,ℓ−1expert(x))⏟upstream-drift propagation pi,ℓ(x). e_{i,\ell}(x) = \underbrace{ f_{\ell}^{\merge}\!\left(h_{i,\ell-1}^{\expert}(x)\right) - f_{i,\ell}^{\expert}\!\left(h_{i,\ell-1}^{\expert}(x)\right) }_{\text{local mismatch}\ \scriptstyle m_{i,\ell}(x)} + \underbrace{ f_{\ell}^{\merge}\!\left(h_{i,\ell-1}^{\merge}(x)\right) - f_{\ell}^{\merge}\!\left(h_{i,\ell-1}^{\expert}(x)\right) }_{\text{upstream-drift propagation}\ \scriptstyle p_{i,\ell}(x)}.

Proof. The algebraic identity and its regularity details are deferred to Appendix C.

Interpretation. Proposition 1 decomposes layer-wise feature drift into two terms: local mismatch and upstream-drift propagation. The local mismatch mi,ℓ(x)m_{i,\ell}(x) measures the feature mismatch caused by the difference between the merged and expert layer maps at the same expert input feature. The propagation term pi,ℓ(x)p_{i,\ell}(x) measures how feature drift inherited from earlier layers changes the input feature of layer ℓ\ell and is then carried into the output feature of this layer.

Proposition 2. [Layer-wise propagation of local mismatch]

Fix a task ii, sample xx, and layers 1,…,L1,\dots,L. Suppose that, for each layer ℓ\ell, fℓmergef_\ell^{\merge} is continuously differentiable on an open neighborhood containing the segment between hi,ℓ−1expert(x)h_{i,\ell-1}^{\expert}(x) and hi,ℓ−1merge(x)h_{i,\ell-1}^{\merge}(x). Let Ai,ℓ(x)A_{i,\ell}(x) denote the corresponding path-averaged local sensitivity operator, defined explicitly in Eq. 18. Then the drift obeys the layer-by-layer recursion

pi,ℓ(x)=Ai,ℓ(x)ei,ℓ−1(x),ei,ℓ(x)=Ai,ℓ(x)ei,ℓ−1(x)+mi,ℓ(x).\begin{aligned} p_{i,\ell}(x) &= A_{i,\ell}(x)e_{i,\ell-1}(x), \qquad e_{i,\ell}(x) = A_{i,\ell}(x)e_{i,\ell-1}(x)+m_{i,\ell}(x). \end{aligned}

Since ei,0(x)=0e_{i,0}(x)=0, the final feature drift is

ei,L(x)=∑ℓ=1LPi,ℓ→L(x)mi,ℓ(x),Pi,ℓ→L(x)=Ai,L(x)Ai,L−1(x)⋯Ai,ℓ+1(x),Pi,L→L(x)=I.\begin{aligned} e_{i,L}(x) &= \sum_{\ell=1}^{L} P_{i,\ell\rightarrow L}(x)m_{i,\ell}(x), \notag\\ P_{i,\ell\rightarrow L}(x) &= A_{i,L}(x)A_{i,L-1}(x)\cdots A_{i,\ell+1}(x), \qquad P_{i,L\rightarrow L}(x)=I . \end{aligned}

The product composes compatible local maps. See Eq. 21 for the ss-to-tt form.

Proof. The local sensitivity derivation and unrolled expansion are deferred to Appendix C.

Interpretation. Proposition 2 gives a simple forward-order view: local mismatches arise at individual layers, and their induced feature drift propagates through downstream layers to the final layer. Thus, final feature drift is a downstream combination of local mismatches from different layers. For residual networks, we further show that residual paths carry upstream feature drift through the skip connection. The drift can also grow under specific conditions. See Appendix D.

3.3 From Feature Drift to Output Drift

Definition 2. [Output drift for task scores]

For the merged model MmergeM^{\merge}, let ziexpert(x)z_i^{\expert}(x) and zimerge(x)z_i^{\merge}(x) be the expert and merged task scores, each represented as a score vector. Let ψiexpert\psi_i^{\expert} and ψimerge\psi_i^{\merge} map final features to task scores, with ziexpert(x)=ψiexpert(hi,Lexpert(x))z_i^{\expert}(x)=\psi_i^{\expert}(h_{i,L}^{\expert}(x)) and zimerge(x)=ψimerge(hi,Lmerge(x))z_i^{\merge}(x)=\psi_i^{\merge}(h_{i,L}^{\merge}(x)). The post-merging output drift is

Δzimerge(x)=zimerge(x)−ziexpert(x). \Delta z_i^{\merge}(x)=z_i^{\merge}(x)-z_i^{\expert}(x).

In the logit or similarity settings used below, ziz_i can contain class logits, CLIP candidate scores (scaled similarity scores over a fixed candidate set), or fixed-prefix decoder vocabulary logits.

Proposition 3. [Feature-to-output perturbation]

Suppose ψimerge\psi_i^{\merge} is locally Bimerge(x)B_i^{\merge}(x)-Lipschitz on the segment between hi,Lexpert(x)h_{i,L}^{\expert}(x) and hi,Lmerge(x)h_{i,L}^{\merge}(x). Here Bimerge(x)B_i^{\merge}(x) locally bounds how much the merged task score map can change when its final-feature input moves along this segment. Then

∥Δzimerge(x)∥2≤Bimerge(x)∥ei,L(x)∥2+δiψ,merge(x), \left\Vert \Delta z_i^{\merge}(x)\right\Vert _2 \le B_i^{\merge}(x)\left\Vert e_{i,L}(x)\right\Vert _2 + \delta_i^{\psi,\merge}(x),

where

δiψ,merge(x)=∥ψimerge ⁣(hi,Lexpert(x))−ψiexpert ⁣(hi,Lexpert(x))∥2\delta_i^{\psi,\merge}(x) = \left\Vert \psi_i^{\merge}\!\left(h_{i,L}^{\expert}(x)\right) - \psi_i^{\expert}\!\left(h_{i,L}^{\expert}(x)\right) \right\Vert _2

is the score map mismatch. If the task score map is shared, this term is 00.

Proof. This is the perturbation bound proved in Proposition 6.

Interpretation. Proposition 3 shows that final feature drift can propagate to output drift, thus further affect model outputs and task loss. For example, in a language model under a fixed prefix, token probabilities are obtained by applying softmax to output logits. Holding other logits fixed, a larger token logit gives a larger token probability, so output-logit drift can potentially induce probability drift and change the next-token choice. Cross-entropy loss is also tied to the probability assigned to the target token, so probability drift can induce loss drift. Detailed analysis is deferred to Appendix E.

4. FeatCal: Feature Calibration for Post-Merging Models

The preceding drift analysis guides the design of a direct calibration procedure for an already merged model. FeatCal visits layers in forward order, takes a current feature snapshot after earlier layer updates, and solves expert-guided calibration objectives for the modules in that layer with explicit regularization.

4.1 Basic Calibration Objective

Forward-order calibration. The drift analysis in Section 3.2 shows that feature drift arises from local mismatch terms that propagate layer by layer in forward order. FeatCal follows the same order by calibrating the merged model layer by layer: after earlier layers are calibrated, it recollects current calibrated-model features for the next layer before fitting that layer's module objectives. This schedule also reduces mismatch between features used during calibration and features exposed by the deployed calibrated model; Appendix F formalizes this source/deployed feature mismatch.

Why calibrate linear modules? Under this schedule, each layer provides a shared feature snapshot. Given the cached input features and expert target features for a module in that snapshot, the local surrogate can be defined for any linear module. In practice, FeatCal applies it to the modules configured for calibration. Linear modules are natural feature-mixing and projection points in the architectures we study, including attention and MLP linear modules. Once their calibrated-model input features are fixed, each linear module gives a tractable regularized regression problem with a closed-form update. The update replaces the existing merged weight, preserving the architecture without gradient descent, adapters, or inference-time modules. For other affine components, including bias parameters and LayerNorm affine parameters, we also design calibration updates, as detailed in Appendix G. In practice, these extra updates give limited gains: on CLIP, their average accuracy improvement is less than 0.5 percentage points. The main gains come from calibrating linear modules.

Linear module feature drift. For a fixed linear module inside the current layer, we define a module-local version of feature drift using the variables available at that module. Let Wi,Wmerge,Wbase∈Rm×dW_i,W^{\merge},W^{\base}\in\mathbb{R}^{m\times d} be the task-ii expert, merged, and base weights for this module. When processing the layer, we cache fixed input feature matrices XicaliX_i^{\cali} and XiexpertX_i^{\expert} for this module on the same task-ii calibration samples. Here XicaliX_i^{\cali} is produced by the prefix-calibrated model and is therefore the deployed input source, while XiexpertX_i^{\expert} is the corresponding task-expert feature matrix.

Following the layer-wise feature drift definition in Eq. 2, we define the linear module feature drift of a candidate calibrated weight WW by

eilinear(W)=WXicali−WiXiexpert. e_i^{\mathrm{linear}}(W) = W X_i^{\cali} - W_i X_i^{\expert}.

As in the layer-wise analysis, this module-level drift can be interpreted as a combination of module-local mismatch and upstream-drift propagation.

Basic calibration objective. With input features fixed, the per-module calibration objective minimizes the overall linear module feature-drift error with a merged-weight penalty:

W⋆=arg⁡min⁡W∈Rm×d∑i=1N∥WXicali−WiXiexpert∥F2  +  λmerge∥W−Wmerge∥F2. W^\star = \arg\min_{W \in \mathbb{R}^{m \times d}} \sum_{i=1}^N \Vert W X_i^{\cali} - W_i X_i^{\expert} \Vert _F^2 \;+\; \lambda_{\merge} \Vert W - W^{\merge} \Vert _F^2.

The quadratic penalty controls how far a single module update can move from the merged weights. This objective is a tractable module-local surrogate for reducing linear module feature drift, rather than an exact objective for end-to-end task risk or all-layer feature drift.

4.2 Feature Interpolation for Calibration Targets

Under the forward-order schedule, each layer's calibration should focus on the local mismatch at its current linear module rather than compensate for feature drift caused by upstream layers. Directly using expert input features for calibration can violate this goal. The gap between XicaliX_i^{\cali} and XiexpertX_i^{\expert} already contains upstream feature drift, so fitting the direct target WiXiexpertW_iX_i^{\expert} may force the current linear module to fit drift left by earlier layers. We therefore introduce an interpolated target feature:

Xitgt=αXiexpert+(1−α)Xicali,α∈[0,1]. X_i^{\tgt} = \alpha X_i^{\expert} + (1-\alpha) X_i^{\cali}, \qquad \alpha \in [0,1].

In Eq. 9, we replace the target term WiXiexpertW_iX_i^{\expert} with WiXitgtW_iX_i^{\tgt}. This target keeps the expert signal while making the local regression less aggressive under upstream drift.

4.3 Anchor Regularization

During calibration, the objective should use targets formed from expert features while keeping the update tied to the merged weights. The base model provides another useful reference. It is pretrained on large data before task specialization and can contain knowledge that a small calibration set does not cover. To use this reference, we include the base weight in the regularization term. The coefficients below let us control how calibration uses the merged and base references. For a linear module, we define the anchor weight as

Wanchor=ρWmerge+(1−ρ)Wbase,ρ∈R. W^{\anchor} = \rho W^{\merge} + (1-\rho) W^{\base}, \qquad \rho\in\mathbb{R}.

The coefficient ρ\rho controls the anchor used by the quadratic penalty.

Combining the target interpolation in Eq. 10 and the anchor in Eq. 11 gives the practical objective:

W⋆=arg⁡min⁡W∈Rm×d∑i=1N∥WXicali−WiXitgt∥F2  +  λ∥W−Wanchor∥F2. W^\star = \arg\min_{W \in \mathbb{R}^{m \times d}} \sum_{i=1}^N \Vert W X_i^{\cali} - W_i X_i^{\tgt} \Vert _F^2 \;+\; \lambda \Vert W - W^{\anchor} \Vert _F^2.

The first term uses the interpolated target from Eq. 10 to fit the expert feature signal without using the raw expert input feature as a hard target. The second term uses the anchor from Eq. 11 to keep the update close to the chosen reference, with λ>0\lambda>0 controlling the regularization strength.

4.4 Task-Wise Scale Normalization and Closed-Form Solution

This subsection describes the update applied after feature collection. At each forward-order layer, FeatCal first caches current calibrated-model features and expert features for the modules calibrated in that layer. The cached features are then used to compute each linear module update separately, and the layer parameters are loaded after the layer's updates are formed. Thus the statistics below are module-local, even though feature collection is organized by layer.

For a fixed linear module in the current layer, let nn denote the calibration sample count for each task at this module. The 1/n1/n factors form per-task empirical moments and prevent each task contribution from scaling directly with the sample count. We summarize task ii by the empirical feature statistics

Gi=1nXicaliXicali⊤∈Rd×d,Ci=1nXitgtXicali⊤∈Rd×d. G_i = \frac{1}{n} X_i^{\cali} {X_i^{\cali}}^\top \in \mathbb{R}^{d \times d}, \qquad C_i = \frac{1}{n} X_i^{\tgt} {X_i^{\cali}}^\top \in \mathbb{R}^{d \times d}.

Here, GiG_i is the input second moment and CiC_i is the target-input cross moment.

The objective in Eq. 12 is a matrix-valued ridge regression problem. To reduce scale sensitivity, we use stabilized inverse task weights with ϵ>0\epsilon>0,

νi=max⁡{∥Gi∥F,ϵ},ωi=νi−1, \nu_i = \max\{\Vert G_i\Vert _F,\epsilon\}, \qquad \omega_i = \nu_i^{-1},

With these stabilized task weights fixed, the module-wise objective becomes

W⋆=arg⁡min⁡W∈Rm×d∑i=1Nωin∥WXicali−WiXitgt∥F2  +  λ∥W−Wanchor∥F2. W^\star = \arg\min_{W \in \mathbb{R}^{m \times d}} \sum_{i=1}^N \frac{\omega_i}{n} \Vert W X_i^{\cali} - W_i X_i^{\tgt} \Vert _F^2 \;+\; \lambda \Vert W - W^{\anchor} \Vert _F^2.

The corresponding stationary condition for this quadratic objective is

W(∑i=1NωiGi+λId)=∑i=1NωiWiCi+λWanchor, W \Bigl( \sum_{i=1}^N \omega_i G_i + \lambda I_d \Bigr) = \sum_{i=1}^N \omega_i W_i C_i + \lambda W^{\anchor},

Solving this linear system gives the closed-form update for this module

W⋆=(∑i=1NωiWiCi+λWanchor)(∑i=1NωiGi+λId)−1. W^\star = \Bigl( \sum_{i=1}^N \omega_i W_i C_i + \lambda W^{\anchor} \Bigr) \Bigl( \sum_{i=1}^N \omega_i G_i + \lambda I_d \Bigr)^{-1}.

Because Gi⪰0G_i\succeq 0, ωi>0\omega_i>0, and λ>0\lambda>0, the inverse is well defined. In implementation, we also add ϵId\epsilon I_d to the solve matrix as a numerical stabilizer.

The complete forward-order calibration procedure is given in Appendix H.

5. Experiments

5.1 Setup

Benchmarks. We use two public FusionBench settings and one MergeBench LLM setting : ① CLIP image classification. We use the FusionBench CLIP model merging benchmark with CLIP-ViT-B/32 and CLIP-ViT-L/14. The primary setting has 8 image tasks: SUN397, Stanford Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, and DTD. The 14-task suite adds Flowers102, PCAM, FER2013, Oxford-IIIT Pet, STL10, and CIFAR100. The 20-task suite further adds CIFAR10, Food101, Fashion-MNIST, EMNIST Letters, KMNIST, and Rendered SST2. We follow prior merging protocols, report top-1 accuracy and task averages, and defer full per-task extended results to Appendix I. ② FLAN-T5 text generation. We evaluate FusionBench FLAN-T5-base and FLAN-T5-large merging on 8 prompted GLUE tasks: CoLA, MNLI, MRPC, QNLI, QQP, RTE, SST-2, and STS-B. The base experts are full fine-tuned models, while the large experts use LoRA fine-tuning. We merge task experts and evaluate generated text outputs. We report exact match accuracy except for STS-B, where we report Spearman's ρ\rho, and average the 8 task scores. ③ MergeBench LLM merging. We evaluate Llama-3.2-3B-Instruct and Llama-3.1-8B-Instruct in the MergeBench domain-expert setting. The task suite covers mathematics, coding, instruction following, and general knowledge through MATH-500, GSM8K, HumanEval+, MBPP+, IFEval, and ARC-Challenge. HumanEval+ and MBPP+ report pass@1, and each table average is the mean over all 6 reported MergeBench tasks in this LLM setting.

Compared methods. We compare against pre-trained, single-task, and multi-task references when available. For CLIP, upstream mergers include Simple Averaging, Task Arithmetic, AdaMerging, and WUDI-Merging; for FLAN-T5 and MergeBench, we use Task Arithmetic. We apply FeatCal on top of different upstream mergers and compare with Surgery and ProbSurgery where available. Unless otherwise stated, upstream and baseline hyperparameters follow the FusionBench recipes.

Calibration setup. By default, FeatCal uses 256 calibration samples per task, calibrates layers in forward order, and applies the linear-weight update in Eq. 17. When enabled, it also applies the bias and LayerNorm affine updates in Appendix G; the main CLIP runs enable both, while the FLAN-T5 runs calibrate linear-module bias parameters but not LayerNorm affine parameters. For the main CLIP accuracy tables, we fix λ=0.05\lambda=0.05, ρ=2.0\rho=2.0, and α=0.3\alpha=0.3. For FLAN-T5, we use λ=10−5\lambda=10^{-5}, ρ=2.0\rho=2.0, and α=0.6\alpha=0.6. For MergeBench, we use (λ,ρ,α)=(10−5,0.5,0.15)(\lambda,\rho,\alpha)=(10^{-5},0.5,0.15) for Llama-3.2-3B-Instruct and (3.0,0.5,0.15)(3.0,0.5,0.15) for Llama-3.1-8B-Instruct. Test sets are used only for final reporting. Section 5.5 includes a compact sensitivity diagnostic for the 8-task TA setting.

5.2 Results on CLIP Models

Table 1. Multi-task performance of CLIP-ViT-B/32 models on 8 image-classification tasks. All numbers are top-1 accuracy (%). Gray arrows show changes over upstream mergers.

MethodSUN397CarsRESISC45EuroSATSVHNGTSRBMNISTDTDAvg.
Pre-trained63.259.860.746.031.632.548.243.948.2
Fine-tuned (STL)75.078.395.299.097.398.999.679.790.3
Traditional MTL72.376.692.297.995.597.799.377.788.6
Simple Averaging65.4↑0.062.4↑0.070.6↑0.075.7↑0.064.5↑0.055.0↑0.086.3↑0.050.6↑0.066.3↑0.0
w/ Surgery67.4↑2.063.5↑1.180.5↑9.994.7↑19.70.8↑6.379.7↑24.96.8↑10.64.7↑14.77.3↑11.
w/ ProbSurgery69.2↑3.865.8↑3.485.8↑15.93.4↑17.70.4↑5.987.0↑32.96.5↑10.69.1↑18.79.7↑13.
w/ FeatCal69.7↑4.370.4↑8.085.0↑14.95.4↑19.92.6↑28.87.1↑32.97.9↑11.67.4↑16.83.2↑16.
Task Arithmetic57.0↑0.055.7↑0.064.7↑0.073.3↑0.077.9↑0.068.5↑0.096.1↑0.047.1↑0.067.5↑0.0
w/ Surgery59.9↑2.959.9↑4.276.1↑11.92.4↑19.83.6↑5.784.9↑16.98.0↑1.961.1↑14.77.0↑9.5
w/ ProbSurgery60.6↑3.661.1↑5.481.0↑16.93.8↑20.85.3↑7.487.1↑18.97.9↑1.863.8↑16.78.8↑11.
w/ FeatCal70.1↑13.72.5↑16.88.1↑23.96.3↑23.95.0↑17.93.0↑24.98.8↑2.769.8↑22.85.5↑18.
AdaMerging67.9↑0.071.2↑0.084.0↑0.092.3↑0.087.6↑0.093.1↑0.098.2↑0.066.9↑0.082.7↑0.0
w/ Surgery69.8↑1.972.1↑0.988.7↑4.795.3↑3.090.5↑2.995.7↑2.698.7↑0.573.3↑6.485.5↑2.8
w/ ProbSurgery70.6↑2.772.9↑1.790.4↑6.495.8↑3.590.2↑2.695.0↑1.998.8↑0.673.8↑6.985.9↑3.2
w/ FeatCal73.0↑5.177.4↑6.291.3↑7.396.8↑4.594.2↑6.696.7↑3.699.0↑0.876.4↑9.588.1↑5.4
WUDI-Merging68.0↑0.072.5↑0.085.0↑0.094.6↑0.094.8↑0.095.0↑0.099.3↑0.066.7↑0.084.5↑0.0
w/ Surgery69.0↑1.071.7↓0.889.1↑4.197.2↑2.695.6↑0.896.7↑1.799.3↑0.072.7↑6.086.4↑1.9
w/ ProbSurgery69.5↑1.572.7↑0.290.6↑5.697.3↑2.795.6↑0.897.1↑2.199.3↑0.073.3↑6.686.9↑2.4
w/ FeatCal72.2↑4.276.0↑3.593.0↑8.098.1↑3.596.7↑1.997.9↑2.999.4↑0.177.1↑10.88.8↑4.3

Table 2. CLIP Extended Average Accuracy.

2*Method8-Task14-Task14-Task20-Task20-Task
L/14B/32L/14B/32L/14
Pre-trained64.658.869.155.665.6
Fine-tuned (STL)94.390.092.890.393.1
Task Arithmetic80.5↑0.066.2↑0.077.3↑0.060.6↑0.070.3↑0.0
w/ Surgery86.0↑5.576.8↑10.83.8↑6.575.2↑14.82.4↑12.
w/ ProbSurgery87.6↑7.178.3↑12.86.1↑8.877.7↑17.83.3↑13.
w/ FeatCal91.6↑11.81.5↑15.90.7↑13.79.4↑18.84.7↑14.
WUDI-Merging92.2↑0.078.7↑0.088.8↑0.067.1↑0.075.8↑0.0
w/ Surgery92.8↑0.682.5↑3.890.3↑1.576.5↑9.484.9↑9.1
w/ ProbSurgery93.0↑0.882.7↑4.090.7↑1.977.6↑10.86.5↑10.
w/ FeatCal93.5↑1.386.2↑7.591.5↑2.783.5↑16.89.8↑14.

On B/32 8-task CLIP, FeatCal raises the 4 upstream averages from 66.3/67.5/82.7/84.5 to 83.2/85.5/88.1/88.8, beating Surgery and ProbSurgery in each block. The gains are largest for the weaker upstream mergers, but FeatCal also improves AdaMerging and WUDI-Merging, where the merged models are already close to the multi-task reference. This pattern aligns with the feature drift motivation: the same calibration step can recover large lost accuracy and still refine strong merged models without changing the merger itself. The best average, 88.8, is close to the 90.3 task expert average and above the 88.6 multi-task reference. For the TA and WUDI rows in Table 2, FeatCal remains above Surgery and ProbSurgery across all extended averages. On 20-task TA, FeatCal reaches 79.4 on B/32 and 84.7 on L/14, giving +18.8+18.8 and +14.4+14.4 points over TA. Full per-task results are in Appendix I.

5.3 Results on FLAN-T5 Models

Table 3. FLAN-T5 GLUE generation results. Scores are percentages: exact-match accuracy except STS-B, which reports Spearman's ρ\rho; arrows show changes over Task Arithmetic, and bold marks best non-reference entries.

2*Model2*MethodGLUE Tasks2*Avg.
CoLAMNLIMRPCQNLIQQPRTESST-2STS-B
Pre-trained69.156.576.288.482.180.191.262.275.7
Fine-tuned (STL)75.083.487.591.585.485.993.688.786.4
Task Arithmetic70.5↑0.057.8↑0.078.4↑0.090.2↑0.083.6↑0.080.5↑0.092.3↑0.077.8↑0.078.9↑0.0
w/ Surgery70.8↑0.382.4↑24.82.4↑4.089.8↓0.484.2↑0.683.0↑2.592.1↓0.285.2↑7.483.7↑4.8
w/ ProbSurgery82.4↑11.69.1↑11.78.3↓0.180.6↓9.689.8↑6.283.4↑2.981.2↓11.92.5↑14.82.2↑3.3
w/ FeatCal72.2↑1.782.6↑24.85.1↑6.791.1↑0.984.7↑1.185.2↑4.793.0↑0.787.7↑9.985.2↑6.3
Pre-trained73.756.682.491.185.585.694.387.582.1
Fine-tuned (STL)80.288.589.294.487.291.795.290.989.6
Task Arithmetic76.8↑0.085.4↑0.085.3↑0.094.0↑0.085.8↑0.088.1↑0.095.2↑0.087.7↑0.087.3↑0.0
w/ Surgery76.0↓0.887.9↑2.586.0↑0.793.9↓0.186.2↑0.489.9↑1.895.2↑0.089.0↑1.388.0↑0.7
w/ ProbSurgery77.2↑0.487.6↑2.285.0↓0.393.8↓0.286.0↑0.288.4↑0.395.2↑0.087.8↑0.187.6↑0.3
w/ FeatCal79.2↑2.487.8↑2.489.0↑3.793.9↓0.186.9↑1.188.4↑0.395.3↑0.190.8↑3.188.9↑1.6

Table 3 shows that the GLUE gains hold for both FLAN-T5-base and FLAN-T5-large. For base, FeatCal improves Task Arithmetic by +6.3+6.3 average points, exceeds Surgery and ProbSurgery, and improves all 8 tasks, with the largest gains on MNLI and STS-B. For large, where LoRA-based Task Arithmetic is already strong, FeatCal gives smaller gains but still reaches the best post-TA average. This contrast suggests that feature calibration helps most when the merged generator has clear headroom, while still helping in the stronger setting.

5.4 Results on MergeBench LLMs

Table 4. MergeBench results on Llama-family models. All Task Arithmetic rows use scale 0.30.3; arrows show changes over Task Arithmetic within each model block. Bold marks the best non-reference entry for each metric within the corresponding model block.

ModelMethodMATH-500GSM8KHumanEval+MBPP+IFEvalARC-CAvg.
Base47.872.649.456.667.373.161.1
Fine-tuned (STL)49.080.654.957.169.773.164.1
Task Arithmetic44.0↑0.074.8↑0.052.4↑0.053.4↑0.063.2↑0.072.6↑0.060.1↑0.0
w/ Surgery45.8↑1.874.5↓0.353.0↑0.654.5↑1.164.3↑1.172.9↑0.360.8↑0.7
w/ ProbSurgery46.0↑2.074.8↑0.054.3↑1.954.5↑1.164.1↑0.973.0↑0.461.1↑1.0
w/ FeatCal47.0↑3.078.0↑3.253.0↑0.655.3↑1.966.5↑3.372.5↓0.162.1↑2.0
Base48.683.562.263.072.880.968.5
Fine-tuned (STL)52.685.367.163.872.880.970.4
Task Arithmetic49.0↑0.085.6↑0.062.8↑0.061.6↑0.045.3↑0.076.9↑0.063.5↑0.0
w/ Surgery47.2↓1.885.2↓0.464.6↑1.862.2↑0.647.1↑1.877.6↑0.764.0↑0.5
w/ ProbSurgery49.6↑0.685.3↓0.362.8↑0.061.9↑0.348.8↑3.577.8↑0.964.4↑0.9
w/ FeatCal47.6↓1.484.5↓1.159.8↓3.062.7↑1.160.8↑15.79.5↑2.665.8↑2.3

Table 4 shows that the gains extend to LLM domain expert merging. FeatCal improves the 6-task average over Task Arithmetic by +2.0+2.0 points on the 3B model and +2.3+2.3 points on the 8B model, outperforming Surgery and ProbSurgery in both model blocks. The largest gain is on IFEval for the 8B setting, where FeatCal improves Task Arithmetic by +15.+15. points.

5.5 Analysis

We collect 4 diagnostics that probe the mechanism, practical cost, robustness, and stability of post-merging feature calibration.

Figure 2. Feature-calibration diagnostics for TA on CLIP-ViT-B/32. (a) Task-wise final-feature cosine to experts; (b) per-sample cosine on Cars; (c) expert-cosine gain versus accuracy gain over TA w/ Surgery; (d) backbone task-vector cosine to same-task experts.

Feature-calibration diagnostics. Figure 2 tests whether TA w/ FeatCal gains coincide with final features closer to task experts. Panels (a) and (b) show larger expert cosine than Surgery in this setting: the macro average rises from 0.785 to 0.850 over TA w/ Surgery, and Stanford Cars has a mean per-sample gain of 0.084. This matches the intended mechanism because FeatCal calibrates the feature distribution reached by the deployed merged model. Panel (c) shows that expert-feature cosine is a diagnostic rather than a complete explanation: larger cosine gains often track accuracy gains, with an average +8.64+8.64-point gain over TA w/ Surgery, but EuroSAT and MNIST improve with small cosine changes, while GTSRB improves despite a slight decrease. Panel (d) highlights a feature-space and parameter-space mismatch: although FeatCal moves final features closer to experts, the calibrated backbone has lower same-task expert task-vector cosine than the TA baseline on every task. Thus, FeatCal calibrates the deployed feature distribution without bringing the backbone parameters into closer task-vector alignment, and the ProbSurgery comparison in Appendix J shows the same pattern with a stronger adapter baseline.

Table 5. Post-TA runtime at n=256n=256, excluding final evaluation. Speedup is relative to Surgery.

[2][4.35em][c]#1×\times #2

CPU RSS
(Time)(Wh)(GiB)
Surgery1.0x(217s)18.169.3
ProbSurgery1.0x(224s)18.387.0
FeatCal4.1x**(053s)**01.922.8

Figure 5. Sample Efficiency.

Sample efficiency and calibration cost. Figure 5 reports post-TA sample efficiency on CLIP-ViT-B/32 TA with matched per-task budgets nn. All methods start from the same Task Arithmetic model and differ only in the post-merging calibration stream. With 88 examples per task, FeatCal already reaches 82.9%82.9\% average accuracy and then saturates near 85.5%85.5\%, while Surgery and ProbSurgery rise more slowly. Table 5 fixes n=256n=256: excluding final evaluation, the closed-form update takes 53s, 4.1×4.1\times faster than Surgery and 4.2×4.2\times faster than ProbSurgery, with lower GPU energy and much less peak CPU RSS. Together, FeatCal reaches most of its TA gain with few examples and keeps calibration cost low at n=256n=256. Full per-budget accuracy, resource averages, hardware, and logging details are in Appendix K.

Robustness to corrupted calibration data.

Table 6. 8-task TA accuracy under corrupted calibration. Avg. includes clean, Gaussian, blur, and fog.

MethodCleanGauss.BlurFogAvg.
Task Arithmetic67.5--------
w/ Surgery76.872.467.569.471.5
w/ ProbSurgery79.173.265.868.271.6
w/ FeatCal85.576.674.275.678.0

Figure 6. Calibration examples under corruptions.

We corrupt only the images used for post-merging calibration and evaluate on clean test sets from 8-task TA, isolating calibration data quality rather than test-time corruption robustness. Protocol details are in Appendix L.

FeatCal remains strongest in every reported setting, with a 78.0%78.0\% average over clean, Gaussian noise, motion blur, and fog, compared with 71.5%71.5\% for Surgery and 71.6%71.6\% for ProbSurgery. Corrupted runs remain below the clean reference, but the stable ordering in Figure 6 suggests that FeatCal extracts a reliable expert-feature calibration signal from imperfect data.

Ablation study.

Figure 3. CLIP-ViT-B/32 TA coefficient sweeps.

Figure 3 reports single-factor sensitivity in the 8-task TA setting. FeatCal works best in a conservative calibration regime. Overly aggressive targets or regularization can degrade performance. The ridge sweep is flat over small positive values, so the closed-form solve is not tied to a narrow regularization choice, while very large ridge strength suppresses the update. The interpolation sweep has a sharper target-side tradeoff: small and medium α\alpha values help, whereas the hard-expert endpoint can force local updates to absorb upstream drift that should be calibrated gradually across layers. The merged-base blend curve is most asymmetric: moderate extrapolation improves both upstream mergers, but large ρ\rho values push calibrated weights too far from the merged-base reference and especially hurt the WUDI variant. Overall, the ablation is a stability diagnostic with a broad conservative region rather than a single sharp optimum.

6. Conclusion

Feature drift frames post-merging calibration by showing how local mismatches can propagate through a merged model and affect outputs. FeatCal uses this signal by capturing features layer by layer and applying closed-form updates to linear modules, improving CLIP and FLAN-T5 mergers over Surgery and ProbSurgery in the main settings without adding modules at inference.

Acknowledgments

This work was supported by the Hong Kong RGC (TRS: T41-517/25-N; GRF: 15228325), ITC RAISe+ (RAI/24/1/086A), and the Research Institute for Generative Artificial Intelligence at PolyU.

_2026

BibTeX

@misc{gu-2026-featcal,
      title={FeatCal: Feature Calibration for Post-Merging Models},
      author={Yanggan Gu and Shuo Cai and Zihao Wang and Wenjun Wang and Yuanyi Wang and Pengkai Wang and Sirui Huang and Su Lu and Jianmin Wu and Hongxia Yang},
      year={2026},
      eprint={2605.13030},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2605.13030},
}