NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1284 most downloaded on PyPI
Parameter-Efficient Fine-Tuning (PEFT)
Last release 3 days ago
01 Oct 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 36 of 38 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
39 releases · first in 2023
One column per quarter.
Fix: the evaluation_strategy is deprecated by @yuanwu2017 in #2487
In #2468, @AaronZLT added the LoRA-FA optimizer to PEFT. This optimizer is based on AdamW and it increases memory efficiency of LoRA training. This means that you can train LoRA with less memory, or, with the same memory budget, use higher LoRA ranks, potentially getting better results.
Thanks to @PaulAlbert31, a new PEFT method called RandLoRA was added to PEFT (#2464). Similarly to VeRA, it uses non-learnable random low rank matrices that are combined through learnable matrices. This way, RandLoRA can approximate full rank updates of the weights. Training models quantized with bitsandbytes is supported.
@Phoveran added Circular Convolution Adaptation, C3A, in #2577. This new PEFT method can overcome the limit of low rank adaptations as seen e.g. in LoRA while still promising to be fast and memory efficient.
Thanks to @gslama12 and @SP1029, LoRA now supports Conv2d layers with groups != 1. This requires the rank r being divisible by groups. See #2403 and #2567 for context.
@dsocek added support for Intel Neural Compressor (INC) quantization to LoRA in #2499.
DoRA now supports Conv1d layers thanks to @EskildAndersen (#2531).
Passing init_lora_weights="orthogonal" now enables orthogonal weight initialization for LoRA (#2498).
@gapsong brought us Quantization-Aware LoRA training in #2571. This can make QLoRA training more efficient, please check the included example. Right now, only GPTQ is supported.
There has been a big refactor of Orthogonal Finetuning, OFT, thanks to @zqiu24 (#2575). This makes the PEFT method run more quickly and require less memory. It is, however, incompatible with old OFT checkpoints. If you have old OFT checkpoints, either pin the PEFT version to <0.16.0 or retrain it with the new PEFT version.
Thanks to @keepdying, LoRA hotswapping with compiled models no longer leads to CUDA graph re-records (#2611).
required_grads_ of modules_to_save is now set to True when used directly with inject_adapter. This is relevant for PEFT integrations, e.g. Transformers or Diffusers.vlm.language_model, it will no longer work, please apply it to vlm directly (see #2554 for context). Morever, the refactor results in different checkpoints. We managed to ensure backwards compatability in PEFT, i.e. old checkpoints can be loaded successfully. There is, however, no forward compatibility, i.e. loading checkpoints trained after the refactor is not possible with package versions from before the refactor. In this case, you need to upgrade PEFT and transformers. More context in #2574.<0.16.0 and <4.52.0, respectively).modules_to_save by @githubnemo in #2481add_weighted_adapter by @Beinsezii in #2512rank_pattern, rank_alpha for add_weighted_adapter by @Beinsezii in #2550prepare_model_for_gradient_checkpointing protected to public by @qgallouedec in #2569Note truncated.
This patch fixes a bug that resulted in prompt learning methods like P-tuning not to work ( #2477 ).
This patch fixes a bug that resulted in prompt learning methods like P-tuning not to work (#2477).
This patch includes a fix for #2450. In this bug modules_to_save was not handled correctly when used in conjunction with DeepSpeed ZeRO stage 3 which
This patch includes a fix for #2450. In this bug modules_to_save was not handled correctly when used in conjunction with DeepSpeed ZeRO stage 3 which resulted in those modules being placeholder values in the saved checkpoints.
Full Changelog: https://github.com/huggingface/peft/compare/v0.15.0...v0.15.1
PEFT_TYPE_TO_MODEL_MAPPING is now deprecated and should not be relied upon. Use PEFT_TYPE_TO_TUNER_MAPPING instead.
@iboing and @5eqn contributed CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning . This task-driven initialization method has two modes, knowledge-preservation and instruction-preservation, both using external data to select ranks intelligently. The former can be used to select those ranks that correspond to weights not affiliated with knowledge from, say, a QA dataset. The latter can be used to select those ranks that correspond most to the task at hand (e.g., a classification task). (#2231)
The new Trainable Tokens tuner allows for selective training of tokens without re-training the full embedding matrix, e.g. when adding support for reasoning / thinking tokens. This is a lot more memory efficient and the saved checkpoint is much smaller. It can be used standalone or in conjunction with LoRA adapters by passing trainable_token_indices to LoraConfig. (#2376)
LoRA now supports targeting multihead attention modules (but for now only those with _qkv_same_embed_dim=True). These modules were tricky as they may expose linear submodules but won't use their forward methods, therefore needing explicit support. (#1324)
Hotswapping now allows different alpha scalings and ranks without recompilation of the model when the model is prepared using a call to prepare_model_for_compiled_hotswap() before compiling the model. (#2177)
GPTQModel support was added in #2247 as a replacement for AutoGPTQ which is not maintained anymore.
all-linear as target_modules for custom (non-transformers) models (#2267). With this change comes a bugfix where it was possible that non-linear layers were selected when they shared the same name with a linear layer (e.g., bar.foo and baz.foo).register_peft_method() call. (#2282)PEFT_TYPE_TO_MODEL_MAPPING is now deprecated and should not be relied upon. Use PEFT_TYPE_TO_TUNER_MAPPING instead. (#2282)modules_to_save keys wrongly matched parts of the state dict if the key was a substring of another key (e.g., classifier and classifier2). (#2334)disable_input_dtype_casting=True. (#2353)rank_pattern and alpha_pattern used by many adapters now supports matching full paths as well by specifying the pattern with a caret in front, for example: ^foo to target model.foo but not model.bar.foo. (#2419)adapter_name conflict with tuner by @pzdkn in https://github.com/huggingface/peft/pull/2254"all-linear" to target custom models by @BenjaminBossan in https://github.com/huggingface/peft/pull/2267__all__ by @bluenote10 in https://github.com/huggingface/peft/pull/2280config.py by @innerlee in https://github.com/huggingface/peft/pull/2297prepare_model_for_kbit_training docstring by @NilBiescas in https://github.com/huggingface/peft/pull/2305resize_token_embeddings to docs by @bingwork in https://github.com/huggingface/peft/pull/2290get_peft_model() for in-place base model modification by @d-kleine in https://github.com/huggingface/peft/pull/2313low_cpu_mem_usage=True with 8bit bitsandbytes by @BenjaminBossan in https://github.com/huggingface/peft/pull/2325PEFT_TYPE_TO_MODEL_MAPPING variable with deprecation by @BenjaminBossan in https://github.com/huggingface/peft/pull/2328modules_to_save loading if substring by @BenjaminBossan in https://github.com/huggingface/peft/pull/2334modules_to_save by @BenjaminBossan in https://github.com/huggingface/peft/pull/2220torch.compile tests and docs by @BenjaminBossan in https://github.com/huggingface/peft/pull/2332nn.Conv1d by @CCLDArjun in https://github.com/huggingface/peft/pull/2333prepare_model_for_compiled_hotswap raises when no adapter was found by @BenjaminBossan in https://github.com/huggingface/peft/pull/2375hf_hub_download arguments are used when loading locally by @henryzhengr in https://github.com/huggingface/peft/pull/2373all-linear target modules by @BenjaminBossan in https://github.com/huggingface/peft/pull/2391PeftConfig.from_pretrained by @BenjaminBossan in https://github.com/huggingface/peft/pull/2397.eval() for inference by @faaany in https://github.com/huggingface/peft/pull/2408Full Changelog: https://github.com/huggingface/peft/compare/v0.14.0...v0.15.0
@tsachiblau added a new soft prompt method called Context-aware Prompt Tuning (CPT) which is a combination of In-Context Learning and Prompt Tuning in
@tsachiblau added a new soft prompt method called Context-aware Prompt Tuning (CPT) which is a combination of In-Context Learning and Prompt Tuning in the sense that, for each training sample, it builds a learnable context from training examples in addition to the single training sample. Allows for sample- and parameter-efficient few-shot classification and addresses recency-bias.
@sirluk contributed a new LoRA initialization method called Explained Variance Adaptation (EVA). Instead of randomly initializing LoRA weights, this method uses SVD on minibatches of finetuning data to initialize the LoRA weights and is also able to re-allocate the ranks of the adapter based on the explained variance ratio (derived from SVD). Thus, this initialization method can yield better initial values and better rank distribution.
@JL-er added an implementation for Block Affine (Bone) Adaptation which utilizes presumed sparsity in the base layer weights to divide them into multiple sub-spaces that share a single low-rank matrix for updates. Compared to LoRA, Bone has the potential to significantly reduce memory usage and achieve faster computation.
PEFT now supports LoRAs for int8 torchao quantized models (check this and this notebook) . In addition, VeRA can now be used with 4 and 8 bit bitsandbytes quantization thanks to @ZiadHelal.
Hot-swapping of LoRA adapters is now possible using the hotswap_adapter function. Now you are able to load one LoRA and replace its weights in-place with the LoRA weights of another adapter which, in general, should be faster than deleting one adapter and loading the other adapter in its place. The feature is built so that no re-compilation of the model is necessary if torch.compile was called on the model (right now, this requires ranks and alphas to be the same for the adapters).
LoRA and IA³ now support Conv3d layers thanks to @jsilter, and @JINO-ROHIT added a notebook showcasing PEFT model evaluation using lm-eval-harness toolkit.
With the target_modules argument, you can specify which layers to target with the adapter (e.g. LoRA). Now you can also specify which modules not to target by using the exclude_modules parameter (thanks @JINO-ROHIT).
DynamicCache caching infrastructure of transformers (see #2096). If you are using this PEFT version and a recent version of transformers with an old prefix tuning checkpoint, you should double check that it still works correctly and retrain it if it doesn't.lora_bias parameter to LoRA layers to enable bias on LoRA B matrix. This is useful when extracting LoRA weights from fully fine-tuned parameters with bias vectors so that these can be taken into account.from_pretrained now warns the user if PEFT keys are missing.modules_to_save is now properly and transparently handled.SFTConfig instead of SFTTrainer keyword args by @qgallouedec in https://github.com/huggingface/peft/pull/2150eval and no dropout by @ariG23498 in https://github.com/huggingface/peft/pull/2122rank_pattern and alpha_pattern together in LoraConfig by @sirluk in https://github.com/huggingface/peft/pull/2195meta device check bug + add multi-gpu functionality by @sirluk in https://github.com/huggingface/peft/pull/2218None check for loftq_config attribute in LoraConfig by @sirluk in https://github.com/huggingface/peft/pull/2215task_type in PEFT Configurations by @d-kleine in https://github.com/huggingface/peft/pull/2210Full Changelog: https://github.com/huggingface/peft/compare/v0.13.2...v0.14.0
This patch release contains a small bug fix for an issue that prevented some LoRA checkpoints to be loaded correctly (mostly concerning stable diffusi
This patch release contains a small bug fix for an issue that prevented some LoRA checkpoints to be loaded correctly (mostly concerning stable diffusion checkpoints not trained with PEFT when loaded in diffusers, #2144).
Full Changelog: https://github.com/huggingface/peft/compare/v0.13.1...v0.13.2
This patch release contains a small bug fix for the low_cpu_mem_usage=True option (#2113).
This patch release contains a small bug fix for the low_cpu_mem_usage=True option (#2113).
Full Changelog: https://github.com/huggingface/peft/compare/v0.13.0...v0.13.1
Fix usage of deprecated parameters/functions in X-LoRA by @EricLBuehler in https://github.com/huggingface/peft/pull/2010
@kallewoof added LoRA+ to PEFT (#1915). This is a function that allows to initialize an optimizer with settings that are better suited for training a LoRA adapter.
@leo-yangli added a new method to PEFT called VB-LoRA (#2039). The idea is to have LoRA layers be composed from a single vector bank (hence "VB") that is shared among all layers. This makes VB-LoRA extremely parameter efficient and the checkpoints especially small (comparable to the VeRA method), while still promising good fine-tuning performance. Check the VB-LoRA docs and example.
New Hugging Face team member @ariG23498 added the helper function rescale_adapter_scale to PEFT (#1951). Use this context manager to temporarily increase or decrease the scaling of the LoRA adapter of a model. It also works for PEFT adapters loaded directly into a transformers or diffusers model.
@ariG23498 also added DoRA support for embedding layers (#2006). So if you're using the use_dora=True option in the LoraConfig, you can now also target embedding layers.
For some time now, we support inference with batches that are using different adapters for different samples, so e.g. sample 1-5 use "adapter1" and samples 6-10 use "adapter2". However, this only worked for LoRA layers so far. @saeid93 extended this to also work with layers targeted by modules_to_save (#1990).
When loading a PEFT adapter, you now have the option to pass low_cpu_mem_usage=True (#1961). This will initialize the adapter with empty weights ("meta" device) before loading the weights instead of initializing on CPU or GPU. This can speed up loading PEFT adapters. So use this option especially if you have a lot of adapters to load at the same time or if these adapters are very big. Please let us know if you encounter issues with this option, as we may make this the default in the future.
Unless indicated otherwise, PEFT adapters are saved and loaded using the secure safetensors format. However, we also support the PyTorch format for checkpoints, which relies on the inherently insecure pickle protocol from Python. In the future, PyTorch will be more strict when loading these files to improve security by making the option weights_only=True the default. This is generally recommended and should not cause any trouble with PEFT checkpoints, which is why with this release, PEFT will enable this by default. Please open an issue if this causes trouble.
merge_and_unload by @snarayan21 in https://github.com/huggingface/peft/pull/1978helper.rescale_adapter_scale by @ariG23498 in https://github.com/huggingface/peft/pull/1989test_vera_dtypes on XPU by @faaany in https://github.com/huggingface/peft/pull/2017TestModelAndLayerStatus device-agnostic by @faaany in https://github.com/huggingface/peft/pull/2026test_mixed_adapter_batches_lora_opt_timing on XPU by @faaany in https://github.com/huggingface/peft/pull/2021test_common_gpu.py to work on XPU by @faaany in https://github.com/huggingface/peft/pull/2031test_gpu_examples.py on XPU by @faaany in https://github.com/huggingface/peft/pull/2036tie_word_embeddings by @ltoniazzi in https://github.com/huggingface/peft/pull/2025evaluation_strategy by @muellerzr in https://github.com/huggingface/peft/pull/1664Full Changelog: https://github.com/huggingface/peft/compare/v0.12.0...v0.13.0
Calling save_pretrained with the convert_pissa_to_lora argument is deprecated, the argument was renamed to path_initial_model_for_weight_conversion (#…
@tokenizer-decode added support for a new LoRA initialization strategy called OLoRA (#1828). With this initialization option, the LoRA weights are initialized to be orthonormal, which promises to improve training convergence. Similar to PiSSA, this can also be applied to models quantized with bitsandbytes. Check out the accompanying OLoRA examples.
@EricLBuehler added the X-LoRA method to PEFT (#1491). This is a mixture of experts approach that combines the strength of multiple pre-trained LoRA adapters. Documentation has yet to be added but check out the X-LoRA tests for how to use it.
@Phoveran, @zqgao22, @Chaos96, and @DSAILatHKUST added discrete Fourier transform fine-tuning to PEFT (#1838). This method promises to match LoRA in terms of performance while reducing the number of parameters even further. Check out the included FourierFT notebook.
@DaShenZi721 added support for Householder Reflection Adaptation (#1864). This method bridges the gap between low rank adapters like LoRA on the one hand and orthogonal fine-tuning techniques such as OFT and BOFT on the other. As such, it is interesting for both LLMs and image generation models. Check out the HRA example on how to perform DreamBooth fine-tuning.
add_weighted_adapter method thanks to @alexrs (#1701).peft_model.get_layer_status() and peft_model.get_model_status() to get an overview of the layer/model status of the PEFT model. This can be especially helpful when dealing with multiple adapters or for debugging purposes. More information can be found in the docs (#1743).Important: If the base model is loaded in float16 (fp16) or bfloat16 (bf16), PEFT now autocasts adapter weights to float32 (fp32) instead of using the dtype of the base model (#1706). This requires more memory than previously but stabilizes training, so it's the more sensible default. To prevent this, pass autocast_adapter_dtype=False when calling get_peft_model, PeftModel.from_pretrained, or PeftModel.load_adapter.
The logic of device placement when loading multiple adapters on the same model has been changed (#1742). Previously, PEFT would move all adapters to the device of the base model. Now, only the newly loaded/created adapter is moved to the base model's device. This allows users to have more fine-grained control over the adapter devices, e.g. allowing them to offload unused adapters to CPU more easily.
save_pretrained with the convert_pissa_to_lora argument is deprecated, the argument was renamed to path_initial_model_for_weight_conversion (#1828). Also, calling this no longer deletes the original adapter (#1933).path_initial_model_for_weight_conversion) while also using use_rslora=True and rank_pattern or alpha_pattern now raises an error (#1930). This used not to raise but inference would return incorrect outputs. We also warn about this setting during initialization.We are now making sure to tag appropriate issues with the contributions welcome label. If you are looking for a way to contribute to PEFT, check out these issues.
config.json when the base model_id is local. by @elementary-particle in https://github.com/huggingface/peft/pull/1668merge_and_unload docs by @younesbelkada in https://github.com/huggingface/peft/pull/1805merge_and_unload by @snarayan21 in https://github.com/huggingface/peft/pull/1944Full Changelog: https://github.com/huggingface/peft/compare/v0.11.1...v0.12.0
Fix a bug that could lead to C++ compilation errors after importing PEFT (#1738 #1739).
Fix a bug that could lead to C++ compilation errors after importing PEFT (#1738 #1739).
Full Changelog: https://github.com/huggingface/peft/compare/v0.11.0...v0.11.1
Don't use deprecated Repository anymore by @Wauplin in https://github.com/huggingface/peft/pull/1641
Thanks to @yfeng95, @Zeju1997, and @YuliangXiu, PEFT was extended with BOFT: Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization (#1326, BOFT paper link). In PEFT v0.7.0, we already added OFT, but BOFT is even more parameter efficient. Check out the included BOFT controlnet and BOFT dreambooth examples.
If the parameter reduction of LoRA is not enough for your use case, you should take a close look at VeRA: Vector-based Random Matrix Adaptation (#1564, VeRA paper link). This method resembles LoRA but adds two learnable scaling vectors to the two LoRA weight matrices. However, the LoRA weights themselves are shared across all layers, considerably reducing the number of trainable parameters.
The bulk of this PR was implemented by contributor @vvvm23 with the help of @dkopi.
PiSSA, Principal Singular values and Singular vectors Adaptation, is a new initialization method for LoRA, which was added by @fxmeng (#1626, PiSSA paper link). The improved initialization promises to speed up convergence and improve the final performance of LoRA models. When using models quantized with bitsandbytes, PiSSA initialization should reduce the quantization error, similar to LoftQ.
Thanks to @fahadh4ilyas, PEFT LoRA linear layers now support Half-Quadratic Quantization, HQQ (#1618, HQQ repo). HQQ is fast and efficient (down to 2 bits), while not requiring calibration data.
Another new quantization method supported in PEFT is Easy & Efficient Quantization for Transformers, EETQ (#1675, EETQ repo). This 8 bit quantization method works for LoRA linear layers and should be faster than bitsandbytes.
We added a feature to show adapter layer and model status of PEFT models in #1663. With the newly added methods, you can easily check what adapters exist on your model, whether gradients are active, whether they are enabled, which ones are active or merged. You will also be informed if irregularities have been detected.
To use this new feature, call model.get_layer_status() for layer-level information, and model.get_model_status() for model-level information. For more details, check out our docs on layer and model status.
modules_to_saveWe had the issue that when we were using classes such as PeftModelForSequenceClassification, we implicitly added the classifier layers to model.modules_to_save. However, this would only add a new ModulesToSaveWrapper instance for the first adapter being initialized. When initializing a 2nd adapter via model.add_adapter, this information was ignored. Now, peft_config.modules_to_save is updated explicitly to add the classifier layers (#1615). This is a departure from how this worked previously, but it reflects the intended behavior better.
Furthermore, when merging together multiple LoRA adapters using model.add_weighted_adapter, if these adapters had modules_to_save, the original parameters of these modules would be used. This is unexpected and will most likely result in bad outputs. As there is no clear way to merge these modules, we decided to raise an error in this case (#1615).
lru_cache to import_utils calls that did not previously have it by @tisles in https://github.com/huggingface/peft/pull/1584Repository anymore by @Wauplin in https://github.com/huggingface/peft/pull/1641% to be sensible by @stas00 in https://github.com/huggingface/peft/pull/1648dreambooth Git link by @charliermarsh in https://github.com/huggingface/peft/pull/1660Full Changelog: https://github.com/huggingface/peft/compare/v0.10.0...v0.11.0
The function prepare_model_for_int8_training was deprecated for quite some time and is now removed completely. Use prepare_model_for_kbit_training ins…
We added a couple of changes to allow QLoRA to work with DeepSpeed ZeRO3 and Fully Sharded Data Parallel (FSDP). For instance, this allows you to fine-tune a 70B Llama model on two GPUs with 24GB memory each. Besides the latest version of PEFT, this requires bitsandbytes>=0.43.0, accelerate>=0.28.0, transformers>4.38.2, trl>0.7.11. Check out our docs on DeepSpeed and FSDP with PEFT, as well as this blogpost from answer.ai, for more details.
First time contributor @siddartha-RE added support for layer replication with LoRA. This allows you to duplicate layers of a model and apply LoRA adapters to them. Since the base weights are shared, this costs only very little extra memory, but can lead to a nice improvement of model performance. Find out more in our docs.
Last release, we added the option to enable DoRA in PEFT by simply adding use_dora=True to your LoraConfig. However, this only worked for non-quantized linear layers. With this PEFT release, we now also support Conv2d layers, as well as linear layers quantized with bitsandbytes.
If you have a PEFT model with multiple LoRA adapters attached to it, it's now possible to apply different adapters (or, in fact, no adapter) on different samples in the same batch. To do this, pass a list of adapter names as an additional argument. For example, if you have a batch of three samples:
output = model(**inputs, adapter_names=["adapter1", "adapter2", "__base__"])`
Here, "adapter1" and "adapter2" should be the same name as your corresponding LoRA adapters and "__base__" is a special name that refers to the base model without any adapter. Find more details in our docs.
Without this feature, if you wanted to run inference with different LoRA adapters, you'd have to use single samples or try to group batches with the same adapter, then switch between adapters using set_adapter -- this is inefficient and inconvenient. Therefore, it is recommended to use this new, faster method from now on when encountering this scenario.
We added an alternative way to initialize LoRA weights for a quantized model using the LoftQ method, which can be more convenient than the existing method. Right now, using LoftQ requires you to go through multiple steps as shown here. Furthermore, it's necessary to keep a separate copy of the quantized weights, as those are not identical to the quantized weights from the default model.
Using the new replace_lora_weights_loftq function, it's now possible to apply LoftQ initialization in a single step and without the need for extra copies of the weights. Check out the docs and this example notebook to see how it works. Right now, this method only supports 4bit quantization with bitsandbytes, and the model has to be stored in the safetensors format.
The function prepare_model_for_int8_training was deprecated for quite some time and is now removed completely. Use prepare_model_for_kbit_training instead.
Besides these highlights, we added many small improvements and fixed a couple of bugs. All these changes are listed below. As always, we thank all the awesome contributors who helped us improve PEFT.
CI / Docker] Follow up from #1481 by @younesbelkada in https://github.com/huggingface/peft/pull/1487Docs/ bnb / DeepSpeed] Add clarification on bnb + PEFT + DS compatibilities by @younesbelkada in https://github.com/huggingface/peft/pull/1529num_parameters() and get_nb_trainable_parameters() in PEFT by @kmehant in https://github.com/huggingface/peft/pull/1531prompt_tuning_init==TEXT by @kmehant in https://github.com/huggingface/peft/pull/1519levenshtein_distance algorithm in peft_lora_seq2seq_accelera… by @SUNGOD3 in https://github.com/huggingface/peft/pull/1527prompt_based_methods.md by @insist93 in https://github.com/huggingface/peft/pull/1548BitsAndBytesConfig as load_in_* is deprecated by @BenjaminBossan in https://github.com/huggingface/peft/pull/1552CI] Fix test docker CI by @younesbelkada in https://github.com/huggingface/peft/pull/1535Full Changelog: https://github.com/huggingface/peft/compare/v0.9.0...v0.10.0
[core/TPLinear] Fix breaking change by @younesbelkada in https://github.com/huggingface/peft/pull/1439
With PR #1364, we added new methods for merging LoRA weights together. This is not about merging LoRA weights into the base model. Instead, this is about merging the weights from different LoRA adapters into a single adapter by calling add_weighted_adapter. This allows you to combine the strength from multiple LoRA adapters into a single adapter, while being faster than activating each of these adapters individually.
Although this feature has already existed in PEFT for some time, we have added new merging methods that promise much better results. The first is based on TIES, the second on DARE and a new one inspired by both called Magnitude Prune. If you haven't tried these new methods, or haven't touched the LoRA weight merging feature at all, you can find more information here:
Via #1394, we now support AutoAWQ in PEFT. This is a new method for 4bit quantization of model weights.
<img width="1197" alt="Screenshot 2024-02-28 at 09 41 40" src="https://github.com/huggingface/peft/assets/49240599/431d485b-c2b9-4e49-b407-89977875e6ef">
Similarly, we now support AQLM via #1476. This method allows to quantize weights to as low as 2 bits. Both methods support quantizing nn.Linear layers. To find out more about all the quantization options that work with PEFT, check out our docs here.
<img width="1197" alt="Screenshot 2024-02-28 at 09 42 22" src="https://github.com/huggingface/peft/assets/49240599/6f1e250b-8981-4e2a-9fa2-028d76150912">
Note these integrations do not support merge_and_unload() yet, meaning for inference you need to always attach the adapter weights into the base model
We now support Weight-Decomposed Low-Rank Adaptation aka DoRA via #1474. This new method is builds on top of LoRA and has shown very promising results. Especially at lower ranks (e.g. r=8), it should perform much better than LoRA. Right now, only non-quantized nn.Linear layers are supported. If you'd like to give it a try, just pass use_dora=True to your LoraConfig and you're good to go.
Thanks to @stevhliu and many other contributors, there have been big improvements to the documentation. You should find it more organized and more up-to-date. Our DeepSpeed and FSDP guides have also been much improved.
Check out our improved docs if you haven't already!
If you're implementing custom adapter layers, for instance a custom LoraLayer, note that all subclasses should now implement update_layer -- unless they want to use the default method by the parent class. In particular, this means you should no longer use different method names for the subclass, like update_layer_embedding. Also, we generally don't permit ranks (r) of 0 anymore. For more, see this PR.
Developers should have an easier time now since we fully embrace ruff. If you're the type of person who forgets to call make style before pushing to a PR, consider adding a pre-commit hook. Tests are now a bit less verbose by using plain asserts and generally embracing pytest features more fully. All of this comes thanks to @akx.
On top of these changes, we have added a lot of small changes since the last release, check out the full changes below. As always, we had a lot of support by many contributors, you're awesome!
MatMul8bitLtBackward view issue by @younesbelkada in https://github.com/huggingface/peft/pull/1425core/TPLinear] Fix breaking change by @younesbelkada in https://github.com/huggingface/peft/pull/1439set_adapters() after add_weighted_adapter by @sayakpaul in https://github.com/huggingface/peft/pull/1444modules_to_save config option when using DeepSpeed ZeRO-3 with ZeRO init enabled. by @pacman100 in https://github.com/huggingface/peft/pull/1450core / get_peft_state_dict] Ignore all exceptions to avoid unexpected errors by @younesbelkada in https://github.com/huggingface/peft/pull/1458Adaptation Prompt] Fix llama rotary embedding issue with transformers main by @younesbelkada in https://github.com/huggingface/peft/pull/1459CI] Add CI tests on transformers main to catch early bugs by @younesbelkada in https://github.com/huggingface/peft/pull/1461magnitude_prune merging method by @pacman100 in https://github.com/huggingface/peft/pull/1466CI] Fix adaptation prompt CI on transformers main by @younesbelkada in https://github.com/huggingface/peft/pull/1465CI] Run tests only when relevant files are modified by @younesbelkada in https://github.com/huggingface/peft/pull/1482CI / bnb] Fix failing bnb workflow by @younesbelkada in https://github.com/huggingface/peft/pull/1480PromptTuning] Simple fix for transformers >= 4.38 by @younesbelkada in https://github.com/huggingface/peft/pull/1484CI / Docker]: Create a workflow to temporarly build docker images in case dockerfiles are modified by @younesbelkada in https://github.com/huggingface/peft/pull/1481CI / Adaptation Prompt] Fix CI on transformers main by @younesbelkada in https://github.com/huggingface/peft/pull/1493Docker] Notify us when docker build pass or fail by @younesbelkada in https://github.com/huggingface/peft/pull/1503Full Changelog: https://github.com/huggingface/peft/compare/v0.8.2...v0.9.0
Release v0.8.2.dev0 by @pacman100 in https://github.com/huggingface/peft/pull/1416
core] fix critical bug in diffusers by @younesbelkada in https://github.com/huggingface/peft/pull/1427Full Changelog: https://github.com/huggingface/peft/compare/v0.8.1...v0.8.2
Fix breaking change related to support for saving resized embedding layers and Diffusers models. Contributed by @younesbelkada in https://github.com/h…
This is a small patch release of PEFT that should:
Full Changelog: https://github.com/huggingface/peft/compare/v0.8.0...v0.8.1
Parameter-efficient fine-tuning (PEFT) for cross-task generalization consists of pre-training adapters on a multi-task training set before few-shot ad
Parameter-efficient fine-tuning (PEFT) for cross-task generalization consists of pre-training adapters on a multi-task training set before few-shot adaptation to test tasks. Polytropon [Ponti et al., 2023] (𝙿𝚘𝚕𝚢) jointly learns an inventory of adapters and a routing function that selects a (variable-size) subset of adapters for each task during both pre-training and few-shot adaptation. To put simply, you can think of it as Mixture of Expert Adapters. 𝙼𝙷𝚁 (Multi-Head Routing) combines subsets of adapter parameters and outperforms 𝙿𝚘𝚕𝚢 under a comparable parameter budget; by only fine-tuning the routing function and not the adapters (𝙼𝙷𝚁-z) they achieve competitive performance with extreme parameter efficiency.
Now, you can specify all-linear to target_modules param of LoraConfig to target all the linear layers which has shown to perform better in QLoRA paper than only targeting query and valuer attention layers
Embedding layers of base models are now automatically saved when the embedding layers are resized when fine-tuning with PEFT approaches like LoRA. This enables extending the vocabulary of tokenizer to include special tokens. This is a common use-case when doing the following:
New option use_rslora in LoraConfig. Use it for ranks greater than 32 and see the increase in fine-tuning performance (same or better performance for ranks lower than 32 as well).
all-linear flag by @SumanthRH in https://github.com/huggingface/peft/pull/1357Tests] Add bitsandbytes installed from source on new docker images by @younesbelkada in https://github.com/huggingface/peft/pull/1275bnb] Add bnb nightly workflow by @younesbelkada in https://github.com/huggingface/peft/pull/1282bnb-nightly] Address final comments by @younesbelkada in https://github.com/huggingface/peft/pull/1287prepare_inputs_for_generation logic for Prompt Learning methods by @pacman100 in https://github.com/huggingface/peft/pull/1352all-linear flag by @SumanthRH in https://github.com/huggingface/peft/pull/1357Full Changelog: https://github.com/huggingface/peft/compare/v0.7.1...v0.8.0
This is a small patch release of PEFT that should handle:
This is a small patch release of PEFT that should handle:
Full Changelog: https://github.com/huggingface/peft/compare/v0.7.0...v0.7.1
Orthogonal Fine-Tuning (OFT): A new adapter that is similar to LoRA and shows a lot of promise for Stable Diffusion, especially with regard to control
merge (#1132)"gaussian" (#1189)adapter_model.bin, calling save_pretrained now creates adapter_model.safetensors. Safetensors have numerous advantages over pickle files (which is the PyTorch default format) and well supported on Hugging Face Hub.add_weighted_adapter with the option combination_type="linear", the scaling of the adapter weights is now performed differently, leading to improved results.peft.lora.Linear is no longer a subclass of nn.Linear, so isinstance checks may need updating). Also, to retrieve the original weight of an adapted layer, now use self.get_base_layer().weight, not self.weight (same for bias).As always, a bunch of small improvements, bug fixes and doc improvements were added. We thank all the external contributors, both new and recurring. Below is the list of all changes since the last release.
Docker] Update Dockerfile to force-use transformers main by @younesbelkada in https://github.com/huggingface/peft/pull/1085core] Fix safetensors serialization for shared tensors by @younesbelkada in https://github.com/huggingface/peft/pull/1101id_tensor_storage by @younesbelkada in https://github.com/huggingface/peft/pull/1116ModulesToSaveWrapper when using Low-level API by @younesbelkada in https://github.com/huggingface/peft/pull/1112adapter_names when calling merge by @younesbelkada in https://github.com/huggingface/peft/pull/1132Tests] Fix daily CI by @younesbelkada in https://github.com/huggingface/peft/pull/1136core / LoRA] Add adapter_names in bnb layers by @younesbelkada in https://github.com/huggingface/peft/pull/1139Tests] Do not stop tests if a job failed by @younesbelkada in https://github.com/huggingface/peft/pull/1141huggingface_hub.file_exists instead of custom helper by @Wauplin in https://github.com/huggingface/peft/pull/1145add_weighted_adapter method by @pacman100 in https://github.com/huggingface/peft/pull/1169Tests] Migrate to AWS runners by @younesbelkada in https://github.com/huggingface/peft/pull/1185modules_to_save is specified and multiple adapters are being unloaded by @pacman100 in https://github.com/huggingface/peft/pull/1137Full Changelog: https://github.com/huggingface/peft/compare/v0.6.2...v0.7.0
The following contributors have made significant changes to the library over the last release:
@alexrs
@callanwu
@elyxlz
@lukaskuhn-lku
@okotaku
@yxli2123
@zhangsheng377
This patch release refactors the adapter deletion API and fixes to ModulesToSaveWrapper when using Low-level API.
This patch release refactors the adapter deletion API and fixes to ModulesToSaveWrapper when using Low-level API.
ModulesToSaveWrapper when using Low-level APIModulesToSaveWrapper when using Low-level API by @younesbelkada in https://github.com/huggingface/peft/pull/1112id_tensor_storage by @younesbelkada in https://github.com/huggingface/peft/pull/1116ModulesToSaveWrapper when using Low-level API by @younesbelkada in https://github.com/huggingface/peft/pull/1112Full Changelog: https://github.com/huggingface/peft/compare/v0.6.1...v0.6.2
This patch release fixes the compatbility issues with Adaptation Prompt that users faced with transformers 4.35.0. Moreover, it fixes an issue with to
This patch release fixes the compatbility issues with Adaptation Prompt that users faced with transformers 4.35.0. Moreover, it fixes an issue with token classification PEFT models when saving them using safetensors
core] Fix safetensors serialization for shared tensors by @younesbelkada in https://github.com/huggingface/peft/pull/1101Docker] Update Dockerfile to force-use transformers main by @younesbelkada in https://github.com/huggingface/peft/pull/1085Full Changelog: https://github.com/huggingface/peft/compare/v0.6.0...v0.6.1
🧨 Diffusers now leverage PEFT as a backend for LoRA inference for Stable Diffusion models (#873, #993, #961). Relevant PRs on 🧨 Diffusers are https://
<img src="https://github.com/huggingface/peft/assets/13534540/57eea400-ce5f-4bf8-a499-711fd83f590b" width="600" height="600">
🧨 Diffusers now leverage PEFT as a backend for LoRA inference for Stable Diffusion models (#873, #993, #961). Relevant PRs on 🧨 Diffusers are https://github.com/huggingface/diffusers/pull/5058, https://github.com/huggingface/diffusers/pull/5147, https://github.com/huggingface/diffusers/pull/5151 and https://github.com/huggingface/diffusers/pull/5359. This helps in unlocking a vast number of practically demanding use cases around adapter-based inference 🚀. Now you can do the following with easy-to-use APIs and it supports different checkpoint formats (Diffusers format, Kohya format ...):
For details, refer to the documentation at Inference with PEFT.
r=0). This used to be possible, in which case the adapter was ignored.As always, a bunch of small improvements, bug fixes and doc improvements were added. We thank all the external contributors, both new and recurring. Below is the list of all changes since the last release.
CI] Pin diffusers by @younesbelkada in https://github.com/huggingface/peft/pull/936LoRA] Add scale_layer / unscale_layer by @younesbelkada in https://github.com/huggingface/peft/pull/935tests] add transformers & diffusers integration tests by @younesbelkada in https://github.com/huggingface/peft/pull/962safe_merge option in merge by @younesbelkada in https://github.com/huggingface/peft/pull/1001core / LoRA] Add safe_merge to bnb layers by @younesbelkada in https://github.com/huggingface/peft/pull/1009LoRA] Revert original behavior for scale / unscale by @younesbelkada in https://github.com/huggingface/peft/pull/1029LoRA] Raise error when adapter name not found in set_scale by @younesbelkada in https://github.com/huggingface/peft/pull/1034core] Fix use_reentrant issues by @younesbelkada in https://github.com/huggingface/peft/pull/1036tests] Update Dockerfile to use cuda 12.2 by @younesbelkada in https://github.com/huggingface/peft/pull/1050Full Changelog: https://github.com/huggingface/peft/compare/v0.5.0...v0.6.0
Now, you can finetune GPTQ quantized models using PEFT. Here are some examples of how to use PEFT with a GPTQ model: colab notebook and finetuning scr
Now, you can finetune GPTQ quantized models using PEFT. Here are some examples of how to use PEFT with a GPTQ model: colab notebook and finetuning script.
Enables users and developers to use PEFT as a utility library, at least for injectable adapters (LoRA, IA3, AdaLoRA). It exposes an API to modify the model in place to inject the new layers into the model.
core] PEFT refactor + introducing inject_adapter_in_model public method by @younesbelkada https://github.com/huggingface/peft/pull/749Low-level-API] Add docs about LLAPI by @younesbelkada in https://github.com/huggingface/peft/pull/836Leverage the support for more devices for loading and fine-tuning PEFT adapters.
Stable support and new ways of merging multiple LoRAs. There are currently 3 ways of merging loras supported: linear, svd and cat.
Llama2] Add disabling TP behavior by @younesbelkada in https://github.com/huggingface/peft/pull/728Patch] patch trainable params for 4bit layers by @younesbelkada in https://github.com/huggingface/peft/pull/733AdaLora] Fix adalora inference issue by @younesbelkada in https://github.com/huggingface/peft/pull/745ModulesToSave] add correct hook management for modules to save by @younesbelkada in https://github.com/huggingface/peft/pull/755core] PEFT refactor + introducing inject_adapter_in_model public method by @younesbelkada in https://github.com/huggingface/peft/pull/749Docker] Fix gptq dockerfile by @younesbelkada in https://github.com/huggingface/peft/pull/835Tests] Add 4bit slow training tests by @younesbelkada in https://github.com/huggingface/peft/pull/834Low-level-API] Add docs about LLAPI by @younesbelkada in https://github.com/huggingface/peft/pull/836Full Changelog: https://github.com/huggingface/peft/compare/v0.4.0...v0.5.0
QLoRA uses 4-bit quantization to compress a pretrained language model. The LM parameters are then frozen and a relatively small number of trainable pa
QLoRA uses 4-bit quantization to compress a pretrained language model. The LM parameters are then frozen and a relatively small number of trainable parameters are added to the model in the form of Low-Rank Adapters. During finetuning, QLoRA backpropagates gradients through the frozen 4-bit quantized pretrained language model into the Low-Rank Adapters. The LoRA layers are the only parameters being updated during training. For more details read the blog Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA
core] Protect 4bit import by @younesbelkada in https://github.com/huggingface/peft/pull/480core] Raise warning on using prepare_model_for_int8_training by @younesbelkada in https://github.com/huggingface/peft/pull/483To make fine-tuning more efficient, IA3 (Infused Adapter by Inhibiting and Amplifying Inner Activations) rescales inner activations with learned vectors. These learned vectors are injected into the attention and feedforward modules in a typical transformer-based architecture. These learned vectors are the only trainable parameters during fine-tuning, and thus the original weights remain frozen. Dealing with learned vectors (as opposed to learned low-rank updates to a weight matrix like LoRA) keeps the number of trainable parameters much smaller. For more details, read the paper Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
Addition of PeftModelForQuestionAnswering and PeftModelForFeatureExtraction classes to support QA and Feature Extraction tasks, respectively. This enables exciting new use-cases with PEFT, e.g., LoRA for semantic similarity tasks.
Introduces a new paradigm, AutoPeftModelForxxx intended for users that want to rapidly load and run peft models.
from peft import AutoPeftModelForCausalLM
peft_model = AutoPeftModelForCausalLM.from_pretrained("ybelkada/opt-350m-lora")
AutoPeftModelForxxx by @younesbelkada in https://github.com/huggingface/peft/pull/694Not a transformer model, no problem, we have got you covered. PEFT now enables the usage of LoRA with custom models.
Improvements to add_weighted_adapter method to support SVD for combining multiple LoRAs when creating new LoRA.
New utils such as unload and delete_adapter providing users much better control about how they deal with the adapters.
PEFT is very extensible and easy to use for performing DreamBooth of Stable Diffusion. Community has added conversion scripts to be able to use PEFT models with Civitai/webui format and vice-versa.
CI] Fix CI - pin urlib by @younesbelkada in https://github.com/huggingface/peft/pull/402Tests] Add soundfile to docker images by @younesbelkada in https://github.com/huggingface/peft/pull/401core] Protect 4bit import by @younesbelkada in https://github.com/huggingface/peft/pull/480core] Raise warning on using prepare_model_for_int8_training by @younesbelkada in https://github.com/huggingface/peft/pull/483core] Add gradient checkpointing check by @younesbelkada in https://github.com/huggingface/peft/pull/404LoRA] Allow applying LoRA at different stages by @younesbelkada in https://github.com/huggingface/peft/pull/429Llama-Adapter] fix half precision inference + add tests by @younesbelkada in https://github.com/huggingface/peft/pull/456core] Add safetensors integration by @younesbelkada in https://github.com/huggingface/peft/pull/553core] Fix config kwargs by @younesbelkada in https://github.com/huggingface/peft/pull/561openai/whisper-large-v2 by @alvarobartt in https://github.com/huggingface/peft/pull/563get_peft_model by @samsja in https://github.com/huggingface/peft/pull/566core] Correctly passing the kwargs all over the place by @younesbelkada in https://github.com/huggingface/peft/pull/575test] Adds more CI tests by @younesbelkada in https://github.com/huggingface/peft/pull/586tests] Fix dockerfile by @younesbelkada in https://github.com/huggingface/peft/pull/608core] Add adapter_name in get_peft_model by @younesbelkada in https://github.com/huggingface/peft/pull/610core] Stronger import of bnb by @younesbelkada in https://github.com/huggingface/peft/pull/605Adalora] Add adalora 4bit by @younesbelkada in https://github.com/huggingface/peft/pull/598AdaptionPrompt] Add 8bit + 4bit support for adaption prompt by @younesbelkada in https://github.com/huggingface/peft/pull/604PeftModel.disable_adapter by @ain-soph in https://github.com/huggingface/peft/pull/644AutoPeftModelForxxx by @younesbelkada in https://github.com/huggingface/peft/pull/694Feature] Save only selected adapters for LoRA by @younesbelkada in https://github.com/huggingface/peft/pull/705Auto] Support AutoPeftModel for custom HF models by @younesbelkada in https://github.com/huggingface/peft/pull/707core] Better hub kwargs management by @younesbelkada in https://github.com/huggingface/peft/pull/712Full Changelog: https://github.com/huggingface/peft/compare/v0.3.0...v0.4.0
The following contributors have made significant changes to the library over the last release:
@TimDettmers
@SumanthRH
@kovalexal
@sywangyi
@aarnphm
@martin-liu
@thomas-schillaci
With task guides, conceptual guides, integration guides, and code references all available at your fingertips, 🤗 PEFT's docs (found at https://hugging
With task guides, conceptual guides, integration guides, and code references all available at your fingertips, 🤗 PEFT's docs (found at https://huggingface.co/docs/peft) provide an insightful and easy-to-follow resource for anyone looking to how to use 🤗 PEFT. Whether you're a seasoned pro or just starting out, PEFT's documentation will help you to get the most out of it.
Comprised of both unit and integration tests, it rigorously tests core features, examples, and various models on different setups, including single and multiple GPUs. This commitment to testing helps ensure that PEFT maintains the highest levels of correctness, usability, and performance, while continuously improving in all areas.
CI] Add ci tests by @younesbelkada in https://github.com/huggingface/peft/pull/203CI] Add more ci tests by @younesbelkada in https://github.com/huggingface/peft/pull/223tests] Adds more tests + fix failing tests by @younesbelkada in https://github.com/huggingface/peft/pull/238tests] Adds GPU tests by @younesbelkada in https://github.com/huggingface/peft/pull/256tests] add slow tests to GH workflow by @younesbelkada in https://github.com/huggingface/peft/pull/304core] Better log messages by @younesbelkada in https://github.com/huggingface/peft/pull/366PEFT just got even more versatile with its new Multi Adapter Support! Now you can train and infer with multiple adapters, or even combine multiple LoRA adapters in a weighted combination. This is especially handy for RLHF training, where you can save memory by using a single base model with multiple adapters for actor, critic, reward, and reference. And the icing on the cake? Check out the LoRA Dreambooth inference example notebook to see this feature in action.
PEFT just got even better, thanks to the contributions of the community! The AdaLoRA method is one of the exciting new additions. It takes the highly regarded LoRA method and improves it by allocating trainable parameters across the model to maximize performance within a given parameter budget. Another standout is the Adaption Prompt method, which enhances the already popular Prefix Tuning by introducing zero init attention.
Good news for LoRA users! PEFT now allows you to merge LoRA parameters into the base model's parameters, giving you the freedom to remove the PEFT wrapper and apply downstream optimizations related to inference and deployment. Plus, you can use all the features that are compatible with the base model without any issues.
utils] add merge_lora utility function by @younesbelkada in https://github.com/huggingface/peft/pull/227core] Fix peft multi-gpu issue by @younesbelkada in https://github.com/huggingface/peft/pull/145CI] Add ci tests by @younesbelkada in https://github.com/huggingface/peft/pull/203main by @younesbelkada in https://github.com/huggingface/peft/pull/224CI] Add more ci tests by @younesbelkada in https://github.com/huggingface/peft/pull/223utils] add merge_lora utility function by @younesbelkada in https://github.com/huggingface/peft/pull/227core] Fix offload issue by @younesbelkada in https://github.com/huggingface/peft/pull/248Automation] Add stale bot by @younesbelkada in https://github.com/huggingface/peft/pull/247Automation] Update stale.py by @younesbelkada in https://github.com/huggingface/peft/pull/254tests] Adds more tests + fix failing tests by @younesbelkada in https://github.com/huggingface/peft/pull/238tests] Adds GPU tests by @younesbelkada in https://github.com/huggingface/peft/pull/256test] Add Dockerfile by @younesbelkada in https://github.com/huggingface/peft/pull/278tests] add CI training tests by @younesbelkada in https://github.com/huggingface/peft/pull/311merge_and_unload when having additional trainable modules by @pacman100 in https://github.com/huggingface/peft/pull/322pip caching to CI by @SauravMaheshkar in https://github.com/huggingface/peft/pull/314tests] add slow tests to GH workflow by @younesbelkada in https://github.com/huggingface/peft/pull/304core] Better log messages by @younesbelkada in https://github.com/huggingface/peft/pull/366try and finally in disable_adapter() to catch exceptions by @mukobi in https://github.com/huggingface/peft/pull/368CI] Fix nightly CI issues by @younesbelkada in https://github.com/huggingface/peft/pull/375The following contributors have made significant changes to the library over the last release:
@QingruZhang
@yeoedward
@Splo2t
We tested PEFT on @OpenAI's Whisper Large model and got: i) 5x larger batch sizes ii) Less than 8GB GPU VRAM iii) Best part? Almost no degredation to
We tested PEFT on @OpenAI's Whisper Large model and got: i) 5x larger batch sizes ii) Less than 8GB GPU VRAM iii) Best part? Almost no degredation to WER 🤯
Without PEFT:
With PEFT:
prepare_for_int8_training utilityThis utility enables preprocessing the base model to be ready for INT8 training.
core] add prepare_model_for_training by @younesbelkada in https://github.com/huggingface/peft/pull/85core] Some changes with prepare_model_for_training & few fixes by @younesbelkada in https://github.com/huggingface/peft/pull/105disable_adapter() context managerEnables to disable adapter layers to get the outputs from the frozen base models. An exciting application of this feature allows only a single model copy to be used for policy model and reference model generations in RLHF.
core] add prepare_model_for_training by @younesbelkada in https://github.com/huggingface/peft/pull/85bnb] add flan-t5 example by @younesbelkada in https://github.com/huggingface/peft/pull/86prepare_model_for_training flexible by @pacman100 in https://github.com/huggingface/peft/pull/90bnb optional by @pacman100 in https://github.com/huggingface/peft/pull/97core] Some changes with prepare_model_for_training & few fixes by @younesbelkada in https://github.com/huggingface/peft/pull/105EleutherAI/gpt-neox-20b to support matrix by @pacman100 in https://github.com/huggingface/peft/pull/109core] Fix autocast issue by @younesbelkada in https://github.com/huggingface/peft/pull/121prepare_for_int8_training by @pacman100 in https://github.com/huggingface/peft/pull/127pyproject.toml by @SauravMaheshkar in https://github.com/huggingface/peft/pull/125The following contributors have made significant changes to the library over the last release:
Full Changelog: https://github.com/huggingface/peft/compare/v0.1.0...v0.2.0
Initial release of 🤗 PEFT. Checkout the main README to learn more about it!
Initial release of 🤗 PEFT. Checkout the main README to learn more about it!
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →