# Qwen3.8 微调指南(Unsloth) > 使用 Unsloth 微调 Qwen3.8-27B 的完整教程:SFT、视觉微调、强化学习、导出 # Qwen3.8 微调指南 Qwen3.8-27B 现在可以通过 [Unsloth](https://github.com/unslothai/unsloth) 进行微调和强化学习(RL)训练。它是一个 27B 参数的稠密统一视觉语言模型,原生支持文本、图像和视频输入,具备思维控制功能,上下文窗口为 262K。 * Unsloth 训练 Qwen3.8 的速度比 FA2 设置**快约 1.5 倍**,VRAM 使用量**减少约 50%**(无精度损失) * 通过我们**免费**的 **Kaggle notebooks** 微调 Qwen3.8-27B: | [**对话式微调**](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Qwen3.8_\(27B\)-Conversational.ipynb&accelerator=nvidiaTeslaT4)(可启用视觉) | [**RL GRPO**](https://www.kaggle.com/notebooks/welcome?src=https://github.com/unslothai/notebooks/blob/main/nb/Kaggle-Muse_Glimmer_\(30B\)-GRPO.ipynb&accelerator=nvidiaTeslaT4)(将模型改为 Qwen3.8) | | --- | --- | * **QLoRA 微调只需 24GB VRAM**,LoRA 需要 >36GB * 你也可以在 24GB VRAM 上免费进行 Qwen3.8 的[强化学习](#强化学习-rl)(RL)训练 * **全量微调(FFT)** 也可以,但 VRAM 使用量是 4 倍 * Unsloth 使用 **Flash Linear Attention 内核** 来高效训练 Qwen3.8 * 使用 [`Qwen3.8-27B-unsloth-bnb-4bit`](https://huggingface.co/unsloth/Qwen3.8-27B-unsloth-bnb-4bit) 进行 4-bit QLoRA 训练,然后**导出**为 NVFP4、FP8、GGUF 等格式 * 如果你想保留推理能力,请将推理风格的示例与直接回答混合,并保持至少 75% 的推理数据 请使用最新的 Transformers v5。Qwen3.8 使用 `qwen3_5` 架构。首次运行时 Gated DeltaNet 内核编译可能需要更长时间。 > **提示**:通过 [Unsloth](https://github.com/unslothai/unsloth) 免费微调 Qwen3.8-27B,我们提供了多个 Kaggle notebooks,提供 **30 小时免费 GPU 使用时间,配备 2× Tesla T4 GPU**。Kaggle 是 Google 的产品,类似 Google Colab,提供了一种无需自己拥有 GPU 硬件即可运行微调工作流的便捷方式。 ## Unsloth 指南 Qwen3.8 可以在我们新的开源本地 AI Web UI——[Unsloth Studio](/docs/new/studio) 中运行和微调。 使用 Unsloth Studio,你可以在 **macOS、Windows、Linux** 上本地运行模型,并在 NVIDIA GPU 上进行训练。Intel、MLX 和 AMD 训练支持将于本月推出。 ### 步骤 1:安装 Unsloth 最简单的入门方式是下载 [Unsloth Desktop 应用](/docs/desktop)。支持 [macOS](/docs/get-started/install/mac)、[Windows](/docs/get-started/install/windows-installation) 和 [Linux](/docs/get-started/install/linux)。 [下载 Unsloth](https://unsloth.ai/download) * [下载 macOS 版](https://unsloth.ai/download/mac) * [下载 Windows 版](https://unsloth.ai/download/windows) * [下载 Linux 版](https://unsloth.ai/download/linux) 或者,如果你喜欢手动安装: macOS、Linux、WSL: ```bash curl -fsSL https://unsloth.ai/install.sh | sh ``` Windows PowerShell: ```bash irm https://unsloth.ai/install.ps1 | iex ``` ### 步骤 2:训练 Qwen3.8 进入 Train 选项卡,然后在搜索栏中搜索 Qwen3.8-27B,选择你想要的模型和数据集。接下来,根据需要调整超参数和上下文长度。 ### 步骤 3:监控训练进度 点击开始训练后,你将能够监控和观察模型的训练进度。训练损失应该稳步下降。完成后,模型将自动保存。 ### 步骤 4:导出微调后的模型 完成后,Unsloth Studio 允许你将模型导出为 GGUF、safetensor 等格式。 ### 步骤 5:对比微调模型与原始模型 点击 `Compare Mode` 对比 LoRA adapter 和原始模型。 ## SFT 教程 以下是一个用于纯文本微调的最小化 SFT 教程。你的数据集需要一个已经使用 Qwen chat template 渲染的 `text` 列。 ```python from unsloth import FastModel from datasets import load_dataset from trl import SFTTrainer, SFTConfig max_seq_length = 2048 dataset = load_dataset( "json", data_files = "train.jsonl", split = "train", ) model, tokenizer = FastModel.from_pretrained( model_name = "unsloth/Qwen3.8-27B-unsloth-bnb-4bit", max_seq_length = max_seq_length, load_in_4bit = True, full_finetuning = False, offload_embedding = True, ) model = FastModel.get_peft_model( model, finetune_vision_layers = False, finetune_language_layers = True, finetune_attention_modules = True, finetune_mlp_modules = True, r = 16, lora_alpha = 16, lora_dropout = 0, bias = "none", use_gradient_checkpointing = "unsloth", random_state = 3407, use_rslora = False, loftq_config = None, ) trainer = SFTTrainer( model = model, tokenizer = tokenizer, train_dataset = dataset, args = SFTConfig( dataset_text_field = "text", max_seq_length = max_seq_length, per_device_train_batch_size = 1, gradient_accumulation_steps = 4, warmup_steps = 10, max_steps = 100, learning_rate = 2e-4, logging_steps = 1, optim = "adamw_8bit", output_dir = "outputs_qwen38", seed = 3407, dataset_num_proc = 1, report_to = "none", ), ) trainer.train() ``` `offload_embedding=True` 是可选的,通过将大型未绑定的输入 embedding 保留在 RAM 中来减少常驻 VRAM。Unsloth 在不支持的平台上会自动禁用它。 如果遇到 OOM(内存不足),请减小 `max_seq_length` 并将 batch size 保持为 1,使用 `"unsloth"` 梯度检查点。 ## 视觉微调 Qwen3.8 支持原生图像和视频输入。对于多模态训练,启用视觉层并使用带有 `UnslothVisionDataCollator` 的对话式视觉数据集。 ```python model = FastModel.get_peft_model( model, finetune_vision_layers = True, finetune_language_layers = True, finetune_attention_modules = True, finetune_mlp_modules = True, r = 16, lora_alpha = 16, lora_dropout = 0, bias = "none", use_gradient_checkpointing = "unsloth", random_state = 3407, use_rslora = False, loftq_config = None, ) ``` 当你的数据集仅包含文本时,请设置 `finetune_vision_layers=False`。 ## 强化学习(RL) Qwen3.8 使用相同的 `qwen3_5` 架构,因此使用 Qwen3.5 的 Unsloth RL 路径并禁用快速 vLLM 推理: ```python from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name = "unsloth/Qwen3.8-27B-unsloth-bnb-4bit", max_seq_length = 2048, load_in_4bit = True, fast_inference = False, ) ``` ## 保存/导出微调后的模型 部署时请使用与训练时相同的 chat template 和 EOS token。 ### 保存为 GGUF ```python model.save_pretrained_gguf( "qwen38_gguf", tokenizer, quantization_method = "q4_k_m", ) # model.push_to_hub_gguf( # "hf_username/qwen38_gguf", # tokenizer, # quantization_method = "q4_k_m", # ) ``` ### 保存为 vLLM ```python model.save_pretrained_merged( "qwen38_finetuned", tokenizer, save_method = "merged_16bit", ) ``` ### 仅保存 LoRA adapters ```python model.save_pretrained("qwen38_lora") tokenizer.save_pretrained("qwen38_lora") ``` --- > 原文:https://unsloth.ai/docs/models/qwen3.8/train > 翻译时间:2026-08-26 --- **分类**:教程 **标签**:微调 · Qwen3.8 · unsloth **作者**:子龙 **链接**:https://octohz.com/p/2053