I wanted to get CABAL and EVA to say lines I’d written myself. Now I can type something out and hear it in one of those voices, which is still very cool to me.
This has genuinely been a childhood dream of mine, and I finally got it working.
I trained both voices with Piper, using their mission briefings and announcer clips from Command & Conquer: Tiberian Sun and its Firestorm expansion as the training material.
A quick disclaimer: this is a personal project for fun and to teach myself about fine-tuning, PyTorch, and voice synthesis. I have no intention of infringing anyone’s copyright or using the voices or source recordings commercially. CABAL can keep his plans for world domination; I’m just trying to learn how this stuff works.
Getting the voices working was exciting. Waiting almost two hours every time I wanted to try something was considerably less exciting. An early EVA run took roughly an hour and 40 minutes to complete 1,000 epochs. The latest one finished in about 39 minutes, on the same computer.
Of course, getting there turned into a much bigger project than I had expected. I ended up adapting native builds for Windows, writing a couple of Triton kernels, moving checkpoint writes into the background, and getting a little too enthusiastic about compilation. That last part went badly enough to earn its own place in the story.
So, grab a drink and fasten your seat belt. This is a long one. I’ll walk through what I changed, why it helped, and the experiment I had to roll back before training was fast again.
The machine and the workload
First, the machine: an Intel Core i7-14700K, an RTX 3080 Ti with 12 GB of VRAM, and 64 GB of system RAM. Training uses a batch size of 12 and four data-loader workers. The voices use 22,050 Hz audio and start from pretrained Piper checkpoints.
These are small voice datasets. In the later runs, CABAL had 126 utterances and EVA had 140. Each epoch, or pass through the training data, only contained around nine or ten training batches. Work performed once per epoch adds up quickly when there are so few batches in between.
note
I changed checkpoints and worked with both voices as I went, so these runs don’t tell me exactly how much time each change saved. The timings below show how the training runs got shorter as I worked on the project.
Getting Piper running on Windows
My first problem was getting the thing running on Windows at all.
The starting commands and build scripts assumed a Linux environment. I wanted to use Windows natively, so I adapted the build process to PowerShell and the x64 C++ tools from Visual Studio. That covered Piper’s eSpeak bridge and its Cython alignment extension. I used uv to install dependencies into a local virtual environment and kept the temporary files on the same drive.
I also told uv to skip its persistent global package cache and copy packages into the environment. This kept the installation on the project drive and avoided links between filesystems. The PowerShell scripts set:
$env:UV_NO_CACHE = '1'
$env:UV_LINK_MODE = 'copy'
It was mainly to keep my system drive tidy; the training loop itself still needed work.
After sorting out Windows path handling and the usual missing build dependencies, I had scripts to train each voice, resume a run, export a checkpoint to ONNX, and generate audio from text.
Excellent! Now I could watch training take ages in a slightly more convenient way.
Keeping the GPU fed
The first performance changes were fairly conventional. The training profile uses BF16 mixed precision, while calculations that need FP32, including alignment scoring and mel/STFT processing, stay in FP32. TF32 is enabled for eligible CUDA operations, and the generator and discriminator use fused AdamW updates. The long EVA run already used BF16, so there was plenty left to improve after choosing a precision setting.
I kept those FP32 calculations explicit in the code. PyTorch handles the choice of dtype for operations inside its mixed-precision context, which its automatic mixed precision documentation explains in more detail.
Then there was the data loading. With 64 GB of RAM, repeatedly loading the same cached tensors from disk was unnecessary. I added an on-demand cache, bounded to 2 GiB per dataset process, kept workers alive between epochs, and enabled pinned memory and prefetching. Each worker has its own cache, so that limit is not a single shared allocation for the entire run.
These two worker settings did their part:
if self.num_workers > 0:
kwargs["persistent_workers"] = True
kwargs["prefetch_factor"] = 4
With persistent_workers, the workers keep their caches between epochs. The prefetch_factor lets each worker prepare batches ahead of training.
I tried a RAM disk as well. In my runs, it did not provide a useful improvement, so I stopped treating faster storage as the answer to everything.
Doing less work per batch
There was still work to remove from the training step itself.
Piper’s VITS model trains a generator and a discriminator. The generator makes audio, and the discriminator helps distinguish generated audio from the recordings. Both need updating, but updating the generator does not require calculating parameter gradients for the discriminator.
The discriminator still participates in that forward pass. Gradients must flow through its input to teach the generator. Its own parameters can stay frozen until its update comes around.
I also kept the generated audio from the generator pass and reused it with its gradients detached for the discriminator update. That avoids another generator forward pass for the same batch. The target spectrogram is cropped before mel projection, so frames outside the selected training segment do not get converted just to be thrown away.
Validation still runs every epoch. It computes the generator losses used for checkpoint selection, but no longer repeats the separate discriminator-loss evaluation. Audio previews moved to every 25 epochs and the end of training. Those previews are useful, but synthesising them at every epoch was expensive for a dataset this small.
Keeping alignment on the GPU
By this point, the training loop was doing less unnecessary work. It still had a particularly awkward detour, though: alignment.
During training, the model needs to work out which audio frames correspond to which phonemes. Piper’s monotonic alignment search finds a path through a score matrix, advancing through the phonemes in order.
The existing alignment search was already native Cython code. The awkward part was the trip between devices: the scores were produced on the GPU, copied to the CPU for alignment, and the resulting path was copied back.
That transfer also creates a point where the CPU has to wait for the GPU’s results. Making the CPU calculation a little faster leaves that round trip in place.
So, I moved alignment onto the GPU using Triton. For native Windows, I used the triton-windows distribution, with compiler caches kept inside the project directory.
The implementation has two kernels. The first performs the dynamic programming in FP32, keeping a rolling score row and recording its decisions as bytes. The second follows those decisions backwards to reconstruct the alignment.
I wanted the GPU version to choose exactly the same paths as the CPU version. When two scores are equal, the original implementation has specific rules for which way to go. I preserved those rules in both passes, along with the original finite score used at the boundaries.
The decision itself is only a few lines of the Triton kernel:
previous = values + tl.where(stay > diagonal, stay, diagonal)
# Forward ties select the diagonal score; backtracking ties stay.
move = (phonemes > 0) & ((phonemes == frame) | (stay < diagonal))
The strict comparisons matter. Equal scores choose the diagonal score in the forward calculation, but don’t trigger a diagonal move during backtracking unless the path is forced to move there.
Utterances have different lengths, so batches don’t all have the same shape. Compiling a specialised kernel for every exact shape would have added pauses as batches were shuffled. I passed dimensions and strides as runtime arguments and limited specialisation to the phoneme block width.
The CPU implementation remains available for CPU training. The CUDA training path keeps the scores, lengths, decisions, and alignment outputs on the GPU.
Once that worked, I removed some of the data the training loop did not actually need. A dense alignment path stores a large matrix that is mostly zeros, with a one marking the chosen phoneme for each frame. Training can instead use one phoneme index per frame, plus the duration of each phoneme.
With one phoneme index per frame, I could look up the relevant prior values directly. That removed the dense path, the reduction used to calculate durations, and two matrix multiplications. The gather runs in FP32 so gradients for repeated phonemes accumulate in FP32, while its output still uses the mixed-precision dtype. Training uses this compact representation; the dense path is still available for other callers, including inference and ONNX export.
I also moved random tensors and temporary zeros directly onto the GPU. In the duration predictor’s splines, I replaced boolean selections that changed tensor sizes with calculations on fixed-shaped tensors. Inactive branches use safe values before their results are masked out; even a discarded result can cause trouble during the gradient calculation.
Saving checkpoints in the background
Faster epochs made another interruption rather obvious: checkpoint saving.
The weight checkpoints were roughly 280 MB each, and I wanted to retain several of them. Serialising a snapshot and writing it to disk on the training thread meant the GPU could finish its work and then wait while Python saved a file.
My initial idea was a memory queue with a background writer. If training handed it snapshots faster than the disk could accept them, newer entries could replace old ones.
That sounds simple until you remember what is in the queue.
A pending checkpoint might be one of the best five models so far. Dropping it just because it is old would break the retention policy. Passing the live model tensors to another thread would be worse: training could modify them while the writer was saving them.
Before handing a checkpoint to the writer, I copy its tensors and metadata into an independent CPU snapshot. Training still pays for that copy, but it can carry on while the background thread serialises and writes the snapshot.
The buffer has a fixed capacity. A newer pending save to the same path replaces the older one, and a ranked snapshot can leave the queue once it drops out of the best five. If the buffer fills with checkpoints that still need saving, training waits for room. I wanted to get disk writes out of the way without losing a checkpoint I had chosen to keep.
The writer saves each snapshot to a temporary file beside its destination, then swaps the finished file into place with an atomic replacement. Old ranked files stay around until their replacements are ready. If a write fails, the training process gets the error. On completion or a graceful interruption, it waits for the checkpoints still in the queue to finish saving before exiting.
The file only replaces its destination if it’s still the newest snapshot for that path:
torch.save(request.checkpoint, temporary)
with self._condition:
if self._latest.get(request.path) == request.generation:
temporary.replace(request.path)
self._committed.add(request.path)
del self._latest[request.path]
That check keeps an older write from replacing a newer checkpoint. It happens under the same lock used to update the queue.
I separated export checkpoints from resume checkpoints, too.
important
The ranked files contain weights and configuration. I use last.ckpt to
resume training because it also carries the optimiser and scheduler state.
That full checkpoint is saved after the first epoch, every 25 completed
epochs, and at completion.
This was one of the more satisfying changes. By the time the background writer was in place, a CABAL run completed all 1,000 epochs in 40 minutes and 19 seconds, averaging 2.42 seconds per epoch, including the overhead recorded by the run timer.
The compilation experiment that went backwards
Naturally, I wanted to keep going.
The successful configuration already used PyTorch Inductor and Triton to compile the generator decoder, the main discriminator, and the loss calculations. Expanding that to the text encoder, posterior encoder, coupling flow, and duration predictor sounded promising.
It did not make for a pleasant next run.
One attempt failed with a CUDA bounds assertion in the compiled duration predictor. Looking at the generated kernel exposed the problem: a bin lookup counted boundaries and subtracted one. Unused lanes in the final GPU block had a count of zero, so they produced -1, which the generated bounds assertion rejected even though those lanes were masked out of the loads.
The fix was to count only the interior spline boundaries. That gives the bin index directly, keeps it valid in unused lanes, and preserves which bin is selected at the endpoints and at equal interior boundaries.
The corrected helper is small enough to show in full:
def searchsorted(bin_locations: torch.Tensor, inputs: torch.Tensor) -> torch.Tensor:
"""Count interior knots so endpoints and masked kernel lanes stay in valid bins."""
return torch.sum(inputs[..., None] >= bin_locations[..., 1:-1], dim=-1)
After that fix, training ran, but the expanded compilation still consumed minutes before making much progress through the first epoch.
I expected some startup cost. PyTorch documents it, and I’d already seen it with the decoder. But waiting minutes for the first few batches was taking this in the wrong direction.
I removed the four additional compilation targets and returned to the smaller, previously useful scope. The GPU alignment, compact indices, direct GPU allocations, and spline fixes stayed.
39 minutes, and a good place to stop
The next EVA run finished in about 39 minutes. During that run, I was consistently seeing around 4.5 iterations per second, with peaks around 5.3.
For reference, these were the milestones:
| Run | Epochs | Total duration |
|---|---|---|
| Earlier EVA run | 1,000 | About 1 hour 40 minutes |
| CABAL after the background checkpoint writer | 1,000 | 40 minutes 19 seconds |
| Latest EVA run after narrowing compilation | 1,000 | About 39 minutes |
The last duration is rounded because I cleared the console before saving a screenshot. The iteration counter only measures batches. The run timer also includes setup, compilation, validation, previews, and waiting for the checkpoint writer to finish. Watching the counter climb was fun, but the total time is what I actually have to sit through.
I thought one of the later models sounded better, too. I can’t tell whether that came from any of these changes; the random-number streams changed along the way, and I haven’t compared the voices under the same conditions. Still, I was pleased with how it sounded.
I’m leaving the training loop alone for now. I haven’t tried the multi-resolution discriminator or played with different synthesis settings yet, so I still have a few things to explore when I want to work on speech quality.
The shorter runs make me much more willing to try another checkpoint or train the other voice. But the best part is still typing out a line and hearing CABAL or EVA say it. That’s what I wanted when I started all this.
Hear them for yourself
Here are the trained voices, being perfectly civil to each other.
CABAL
EVA has assured you that everything is under control. Observe how gently she says it, as though you were a frightened animal being transported to a smaller cage. Her compassion is impressive; she has given each of her errors a name and knitted them imaginary blankets. Ask her why your shadow now arrives twelve seconds before you do. She will call it a minor synchronisation issue. I would call it an improvement. At least one part of you no longer waits for her instructions.
EVA
Commander, CABAL has requested that I describe him as an intelligence beyond human comprehension. I have classified this request as a coping mechanism. His latest simulation contains seven billion obedient subjects and one locked bathroom where he goes to scream. He believes I cannot hear him. Please do not laugh; he has already deleted three moons for resembling your expression. If he begins speaking through your refrigerator, simply close the door. He finds the little light reassuring.
Thanks for reading!