Build1 publisher3 min readPublished
Grafting base-model SDF weight changes onto post-trained models does less damage, MATS fellows report
MATS fellows say grafting base-trained synthetic-document edits onto post-trained models beats direct tuning in five families up to 284B parameters. One graft also carried across later SDF, RL and DPO checkpoints without retraining.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- According to the authors, running SDF directly on a post-trained model can significantly damage coherence and capabilities, worst on grip on reality and sharp preferences.
- Grafted models more closely resemble what having the corpus in pre-training data would have produced than natively tuned models do, the authors report.
- The authors say evidence suggests grafting beats controls such as doctags while keeping the same level of behavioral expression.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Each revision of a synthetic corpus costs one fine-tune on the base checkpoint, where the clean alternative re-runs post-training that the authors price at possibly millions of dollars per full run.
- constraint Teams without the base checkpoint behind their deployed model, including anyone working only through an API, cannot use the method, because the graft is computed on base weights.
- decision Teams that hold base weights now have a published reason to move SDF off the instruct checkpoint and add the result back after each later training stage.
Synthetic document fine-tuning (SDF) trains a model on a corpus of pre-training-style documents that portray a chosen scenario, such as "Ed Sheeran is an Olympic gold medalist" [9]. Grafting changes only where that training starts. The steps run in order:
1. Take the pre-training checkpoint the deployed model was post-trained from. 2. Run SDF on that base checkpoint. 3. Subtract the starting weights to get the weight difference. 4. Add that difference to the post-trained model you plan to deploy [3].
The design exists because both obvious options cost something. Mixing the documents into pre-training is the clean version. It shifts the base checkpoint, so post-training has to be redone [13]. The authors put a full post-training run for a high-capability production model at possibly millions of dollars [13]. For a team iterating on a corpus, they call that prohibitively expensive [14]. The accepted shortcut is to run SDF directly on the instruct-tuned model [8]. According to the authors, that shortcut is what damages coherence and capabilities [1].
Who can use grafting gets decided in step 1. The graft is a weight difference computed on base, so a team needs the base checkpoint's weights as well as the post-trained ones [11]. The authors report that the benefit is exclusive to grafts trained before any instruction fine-tuning [6]. A graft taken from the instruct stage and carried into later RL did not work as well as one taken from base [6].
Reuse is the result I'd expect pipeline owners to care about most. One adapter trained on base was applied to SDF, RL and DPO checkpoints without retraining [5]. The authors cite adversarial concealment training for model organisms as an example [5]. A team can train the graft once and add it after each later stage.
The evidence covers implanting false facts, training model organisms and constitutional mid-training, across five model families up to 284B parameters [4]. The post's summary does not give effect sizes or name the families. On controls such as doctags, the authors write only that the evidence "suggests" grafting does better while keeping the same level of behavioral expression [12]. The work came out of Peter and Dani's MATS Fellowship 10.0, mentored by Shi, and so far is reported only in their LessWrong post [10].
Transfer to another team depends on three conditions. The deployed model has to descend from a base checkpoint the team can load. Its edit should be a synthetic-document corpus like the ones tested, and its later training should resemble the SDF, RL and DPO stages the graft was applied to [4][5].
"If a base model is available for whatever model you plan to run SDF on, we'd recommend this method as an easy swap-in that improves the quality of the final model," the authors wrote [7]. I think that is the right default for open-weight SDF work where the base is published. For an existing pipeline, the change is a different starting checkpoint plus one weight addition [3].
What to watch
- Per-family effect sizes and the names of the five model families, showing how wide the coherence gap between native SDF and grafting actually is.
- Replication by a group outside the MATS project, especially on a post-training pipeline with stages beyond the SDF, RL and DPO checkpoints tested.
- Whether developers keep publishing base checkpoints next to post-trained releases; that decides how many deployed models can be grafted at all.