QLoRA on a consumer GPU: fitting a 13B fine-tune in 24GB
You don't need an H100 to fine-tune a 13B model. With 4-bit base quantization, paged optimizer state, and gradient checkpointing, a 4090 will train one in an evening. Here's the recipe and the trap doors.