Skip to content

Build1 publisher3 min readPublished

Patterned PEFT LoRA adapters run at the wrong scale in vLLM 0.30.0

vLLM 0.30.0 ignores the per-module rank and alpha patterns in PEFT LoRA adapters, and one test put the error against PEFT at 70 times the unpatterned baseline. The adapter loads without complaint, so only a config check or a side-by-side against PEFT will catch it.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Patterned PEFT LoRA adapters run at the wrong scale in vLLM 0.30.0
Generated illustration

What happened

  • With no pattern set, vLLM matched PEFT's prompt log-probabilities to a mean gap of 0.0003, within fp16 rounding.
  • An adapter with only alpha_pattern set was off by a mean of 0.0168, so a differing rank is not needed to trigger the error.
  • Multiplying the affected modules' lora_B by the missing factor and clearing the pattern brought the gap back to 0.0003.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Mixture-of-experts fine-tunes built to PEFT's documented recipe, with a smaller rank per expert, are the adapters most likely to be served off-scale on 0.30.0.
  • cost If the OLMoE result holds, about 40% of what a fine-tune gained is lost at serving time, and the team that paid for training absorbs the loss with no error to point at.
  • decision Teams serving patterned adapters on 0.30.0 now choose between rewriting weights so the patterns are empty and keeping those adapters off vLLM until a fix such as #59801 lands.

The cause is in vllm/lora/peft_helper.py. In 0.30.0 that file declares the fields vLLM reads from adapter_config.json: r, lora_alpha, target_modules, bias, modules_to_save, use_rslora, use_dora and vllm_lora_scaling_factor [14]. From them it computes one number. With use_rslora set, the scale is lora_alpha / sqrt(r). Otherwise it is lora_alpha / r [15]. Neither pattern has a field, and the dev.to author's code search of the repository found no other reference to them [16].

Rank survives. vLLM takes each module's rank from the tensor shapes in the safetensors file, so only the scale is shared [17]. PEFT scales each module by its own alpha_m / r_m [1]. Any module whose pattern moves that ratio off lora_alpha / r is wrong by exactly the ratio [17]. In the test adapter, q_proj should run at 32/4 = 8, and vLLM ran it at 2, a factor of 4 [17].

vLLM does refuse some adapters. Its _validate_features check turns away DoRA and modules_to_save, so an adapter that loads without complaint looks supported [9]. No line in the test run's log mentions rank_pattern or alpha_pattern [8]. The output is still the adapter's output, only stronger or weaker in the patterned modules than it was trained to be [10]. According to the author, it takes a PEFT comparison or a measurement of the trained-for quality to see it [10].

The test design is careful. PEFT's own prompt log-probabilities on four 48-token prompts are the reference, and vLLM serves the same adapter on the same base [5]. The setup is an 8 GB RTX 2070 under WSL2 with vLLM 0.30.0, torch 2.13.0+cu130 and PEFT 0.21.2 [4]. The model is a randomly initialised two-layer Qwen3 with hidden size 256 [4]. The card has no bf16, so the runs used fp16 where the upstream report used bf16 [4]. The author also scaled lora_B up so the adapter's effect would stand clear of fp16 noise [5]. That choice makes the gaps visible. I would not carry the absolute figures, including the patterned run's largest single gap of 0.1106 [22], over to a trained adapter.

For the finding to transfer, the drift has to show up on a real model with a trained adapter. The upstream reporter's OLMoE-1B-7B run is that case. On its perplexity figures, vLLM gave up about 40% of the gain the adapter had made over the base model [1]. The dev.to author notes those numbers come from hardware the lab doesn't have, and that their direction and size match the small model [13].

PEFT steers users toward the patterns in one common case. According to the report, PEFT's LoRA documentation recommends rank_pattern for mixture-of-experts models, giving each expert r // num_experts, and points to vLLM for serving [11]. One patterned case escapes. When save_as_lora is called with a dynamic rank, PEFT sets both r and lora_alpha to 1 and gives each module an identical value in rank_pattern and alpha_pattern, so every scale is 1 and vLLM gets it right by coincidence [18].

The author's toolkit includes check-lora-patterns.py, a script that takes an adapter's config and reports each module vLLM will scale wrongly, along with the size of the error [19]. The upstream issue is vllm-project/vllm#59799, and a fix is proposed in #59801 [20].

What to watch

  • Whether #59801 merges and which release carries it; once vLLM honours the patterns, every patterned adapter already in production changes its outputs.
  • Whether vLLM adds a load-time refusal or warning for rank_pattern and alpha_pattern, as _validate_features already does for DoRA and modules_to_save.
  • An independent bf16 reproduction of the OLMoE-1B-7B perplexity gap on production hardware.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories