
LlamaFactory
Open SourceUnified efficient fine-tuning of 100+ LLMs and VLMs.
Published 2 October 2026
Scores
Popularity3/5
Around 75K GitHub stars puts it level with Unsloth among fine-tuning projects; fine-tuning itself is work few application developers ever do, so the name rarely comes up outside ML engineering.
Learning Curve3/5
LLaMA Board makes a first run a matter of filling in forms, but choosing between training methods, quantization levels, and hyperparameters still requires a working grasp of fine-tuning concepts.
Flexibility5/5
Over 100 model families, every common post-training method from SFT to PPO and SimPO, a dozen parameter-efficient techniques, and both config-file and web UI workflows.
Performance4/5
FlashAttention-2, Liger Kernel, and the optional Unsloth backend bring single-GPU QLoRA close to Unsloth's own times, though its memory use is higher than Unsloth running alone.
Portability5/5
Apache-2.0, installs with pip or Docker on Linux, Windows, and macOS, runs on NVIDIA, AMD, and Ascend hardware, and exports standard Hugging Face weights that any runtime can load.
About LlamaFactory
LlamaFactory, long written LLaMA-Factory, is an open-source framework for fine-tuning open-weight language and vision-language models. One codebase covers more than 100 model families, including Llama, Qwen, DeepSeek, Mistral, Gemma, GLM, and Phi, plus multimodal models such as Qwen-VL and LLaVA. The project began at Beihang University, was published at ACL 2024, and is one of the most-starred fine-tuning projects on GitHub.
It bundles the whole post-training toolbox in one place: continued pre-training, supervised fine-tuning, reward modeling, PPO, DPO, KTO, ORPO, and SimPO. Each can run as full fine-tuning, freeze-tuning, LoRA, or QLoRA with 2- to 8-bit quantization, alongside newer methods such as DoRA, PiSSA, GaLore, and OFT. Speed-ups come from FlashAttention-2, Liger Kernel, and Unsloth's kernels as an optional backend, and DeepSpeed and Megatron-core handle multi-GPU and multi-node runs.
There are two ways to drive it. The llamafactory-cli runs training, evaluation, merging, and export from YAML config files, which keeps runs reproducible. LLaMA Board, a Gradio web UI, exposes the same options as forms, so a first fine-tune needs no code at all. After training, the CLI can chat with the model or serve it through an OpenAI-compatible API backed by vLLM or SGLang, and export merged weights for other runtimes.
LlamaFactory is Apache-2.0 and free, with no hosted plan of its own. It runs on the user's own GPUs, a cloud GPU instance, or a Docker image, and the hardware needed ranges from a single consumer GPU for QLoRA on a 7B model to a cluster for full training of a 70B one.
Key Features
- Fine-tuning for 100+ language and vision-language model families
- SFT, reward modeling, PPO, DPO, KTO, ORPO, and SimPO training
- Full, freeze, LoRA, and 2- to 8-bit QLoRA tuning
- LLaMA Board web UI for no-code training runs
- YAML-driven CLI for reproducible training, merging, and export
- FlashAttention-2, Liger Kernel, and Unsloth acceleration backends
- DeepSpeed and Megatron-core for multi-GPU training
- OpenAI-compatible inference through vLLM or SGLang after training
Pros
- The web UI gets a first fine-tune running without writing code, which suits newcomers
- Covers far more training methods and model families than most single-purpose tools
- Can use Unsloth's kernels as a backend, so single-GPU runs get most of Unsloth's speed
- Scales from one consumer GPU to multi-node clusters with DeepSpeed or Megatron
Cons
- Uses noticeably more VRAM than Unsloth for the same QLoRA job in community benchmarks
- The sheer number of options in the UI and configs is overwhelming once past the defaults
- Documentation and community discussion lean toward Chinese-language sources
- No hosted service, so users must provision and manage their own GPUs
LlamaFactory Pricing
Open SourceTools Related to LlamaFactory
Works well with LlamaFactory(1)
LlamaFactory is built on Hugging Face Transformers, PEFT, and TRL: it loads base models and datasets from the Hub by repository ID and exports merged weights in the standard Hub format.
Built on (1)
LlamaFactory is a Python toolkit installed with pip; its CLI and web UI run fine-tuning jobs on the Python machine learning stack.
Required by LlamaFactory(1)
LlamaFactory trains models on PyTorch through Hugging Face Transformers and PEFT; there is no other training backend.
Alternatives to LlamaFactory(1)
LlamaFactory covers more training methods and model families, driven from YAML configs or a no-code web UI, and scales to multi-GPU clusters; Unsloth is faster and lighter on memory on a single GPU and also runs models locally. LlamaFactory can use Unsloth's kernels as a backend, but a project picks one as its toolkit.