Let me know if this sounds familiar. You set up a fresh GPU instance on RunPod, fire up ai-toolkit to train a Flux.1-dev LoRA, and even smartly run an rclone job in tmux to push backups to your Google Drive.
You step away to help the kids with homework or play a quick round of your favorite video game. When you come back, your RunPod credits have run completely dry. The instance stopped, and your training environment is gone.
It reminds me of the old 486 days when a sudden power blink would wipe out hours of work because I forgot to hit save. Or, more accurately, it is like running out of gas right before you reach the mechanic’s garage.
But do not panic. If you check your Google Drive and see your config.yaml, optimizer.pt, and your last .safetensors file (let’s say it stopped at step 1400), you are in a very good spot. You can absolutely resume this without starting from zero.
Here is exactly how to fix it and get your training running again safely.
Step 1: Spin Up a New Instance and Prep the Server
First, you need a new workspace. Deploy a new RunPod instance and get your environment ready. Make sure you allocate enough volume storage—running out of disk space halfway through a resumed training is just as bad as running out of credits.
You will need to clone the ai-toolkit repository again and install all the necessary dependencies, just like you did the first time. Treat this as a completely fresh start for your server setup. Update your package manager and make sure your Python environment is clean and ready to go.
Step 2: Restore Your Dataset (Do Not Skip This)
This is the step people often forget in a panic. The training script does not just pick up the .safetensors file and guess the rest; it still needs your original images to continue learning and calculating gradients.
Upload your dataset to the exact folder path specified in your configuration file. For example, if your config.yaml expects the images at /workspace/ai-toolkit/dataset/my_custom_concept, make sure the images are exactly there.
Check the folder permissions just to be safe. I have lost an hour troubleshooting failed scripts before only to realize Linux was blocking read access to the image folder.
Step 3: Recreate the Output Folder Structure
Next, we need to rebuild the directory where the training script expects to find your historical progress.
In ai-toolkit, the output directory is built automatically by combining two variables from your config file: the training_folder and the project name. For example, if your training_folder is set to output and your project name is my_project_lora, the system expects a folder called /workspace/ai-toolkit/output/my_project_lora.
Since this is a fresh instance, you have to create this directory manually. Just use mkdir -p /workspace/ai-toolkit/output/my_project_lora in your terminal.
Step 4: Transfer Your Backup Files Safely
Now it is time to move your saved files from Google Drive into that newly created output folder. You can use rclone to pull them down, or just download them to your local PC and upload them directly via RunPod’s web interface.
You need to place three specific files into your output folder:
- Your configuration file. (If your cloud drive added numbers to it like
config (2).yaml, rename it back to exactlyconfig.yaml). - The
optimizer.ptfile. - Your last weights file, which should look something like
my_project_lora_000001400.safetensors.
Having the optimizer.pt file is critical here. It contains the learning state and momentum of your training. If you just load the .safetensors file without the optimizer, the AI loses its recent memory of how it was adjusting weights, which can result in a sudden drop in output quality.
Step 5: Start the Engine
Once all your files are in the right place, do yourself a favor and start a fresh tmux session. If your SSH connection drops or your home internet blinks out, tmux ensures the training keeps running in the background.
Navigate to your ai-toolkit directory and run the training command. You want to point the script directly at the config file sitting inside your output folder:
python run.py output/my_project_lora/config.yaml
Because ai-toolkit is built with practical auto-resuming logic, you do not need to add any special flags or edit the code. The script will automatically read the step count (1400) directly from your .safetensors filename. It will then inject the optimizer.pt state and continue training right where it left off, heading straight for your total step goal.
Quick Recap
Running out of RunPod credits is annoying, but it is not a disaster if your backups are solid. Just rebuild the instance, put the dataset and backup files in their exact expected folders, and run the script.
Always keep an eye on your instance balance, but rest easy knowing you can recover from a hard stop without wasting hours of GPU time. Good luck with your model training, and I hope this practical fix saves you some stress!
Until next time, keep fixing things.
No Comments