How to Resume an Interrupted ai-toolkit LoRA Training on RunPod

Let me know if this sounds familiar. You set up a fresh GPU instance on RunPod, fire up ai-toolkit to train a Flux.1-dev LoRA, and even smartly run an rclone job in tmux to push backups to your Google Drive.

You step away to help the kids with homework or play a quick round of your favorite video game. When you come back, your RunPod credits have run completely dry. The instance stopped, and your training environment is gone.

It reminds me of the old 486 days when a sudden power blink would wipe out hours of work because I forgot to hit save. Or, more accurately, it is like running out of gas right before you reach the mechanic’s garage.

But do not panic. If you check your Google Drive and see your config.yaml, optimizer.pt, and your last .safetensors file (let’s say it stopped at step 1400), you are in a very good spot. You can absolutely resume this without starting from zero.

Here is exactly how to fix it and get your training running again safely.

Step 1: Spin Up a New Instance and Prep the Server

First, you need a new workspace. Deploy a new RunPod instance and get your environment ready. Make sure you allocate enough volume storage—running out of disk space halfway through a resumed training is just as bad as running out of credits.

You will need to clone the ai-toolkit repository again and install all the necessary dependencies, just like you did the first time. Treat this as a completely fresh start for your server setup. Update your package manager and make sure your Python environment is clean and ready to go.

Step 2: Restore Your Dataset (Do Not Skip This)

This is the step people often forget in a panic. The training script does not just pick up the .safetensors file and guess the rest; it still needs your original images to continue learning and calculating gradients.

Upload your dataset to the exact folder path specified in your configuration file. For example, if your config.yaml expects the images at /workspace/ai-toolkit/dataset/my_custom_concept, make sure the images are exactly there.

Check the folder permissions just to be safe. I have lost an hour troubleshooting failed scripts before only to realize Linux was blocking read access to the image folder.

Step 3: Recreate the Output Folder Structure

Next, we need to rebuild the directory where the training script expects to find your historical progress.

In ai-toolkit, the output directory is built automatically by combining two variables from your config file: the training_folder and the project name. For example, if your training_folder is set to output and your project name is my_project_lora, the system expects a folder called /workspace/ai-toolkit/output/my_project_lora.

Since this is a fresh instance, you have to create this directory manually. Just use mkdir -p /workspace/ai-toolkit/output/my_project_lora in your terminal.

Step 4: Transfer Your Backup Files Safely

Now it is time to move your saved files from Google Drive into that newly created output folder. You can use rclone to pull them down, or just download them to your local PC and upload them directly via RunPod’s web interface.

You need to place three specific files into your output folder:

  1. Your configuration file. (If your cloud drive added numbers to it like config (2).yaml, rename it back to exactly config.yaml).
  2. The optimizer.pt file.
  3. Your last weights file, which should look something like my_project_lora_000001400.safetensors.

Having the optimizer.pt file is critical here. It contains the learning state and momentum of your training. If you just load the .safetensors file without the optimizer, the AI loses its recent memory of how it was adjusting weights, which can result in a sudden drop in output quality.

Step 5: Start the Engine

Once all your files are in the right place, do yourself a favor and start a fresh tmux session. If your SSH connection drops or your home internet blinks out, tmux ensures the training keeps running in the background.

Navigate to your ai-toolkit directory and run the training command. You want to point the script directly at the config file sitting inside your output folder:

python run.py output/my_project_lora/config.yaml

Because ai-toolkit is built with practical auto-resuming logic, you do not need to add any special flags or edit the code. The script will automatically read the step count (1400) directly from your .safetensors filename. It will then inject the optimizer.pt state and continue training right where it left off, heading straight for your total step goal.

Quick Recap

Running out of RunPod credits is annoying, but it is not a disaster if your backups are solid. Just rebuild the instance, put the dataset and backup files in their exact expected folders, and run the script.

Always keep an eye on your instance balance, but rest easy knowing you can recover from a hard stop without wasting hours of GPU time. Good luck with your model training, and I hope this practical fix saves you some stress!

Until next time, keep fixing things.

No Comments

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.