AudioShake has launched The Refinery, a system that converts mixed recordings into speaker-separated training data for voice AI. It splits conversations into aligned individual tracks while separating dialogue from music and background sound. The company reports more than 100 million minutes processed, with early deployments at frontier AI labs. For business, this matters because existing audio archives could become usable training material without staging new conversations.
How The Refinery separates overlapping speech
According to the announcement, The Refinery works directly from finished recordings and does not require original sessions or separately captured stems. When two people speak at once, it aims to recover each voice individually while preserving the timing of the overlap. The output can be understood as aligned tracks: one speaker on one track, another speaker on another, with shared timing showing interruptions and short acknowledgments. AudioShake says the system does not generate or reconstruct speech, with separated voices and frequencies taken from the original recording.
The product combines speaker separation with diarization, which are different operations. Diarization identifies when speakers are active, while separation produces independently usable audio signals for each voice. AudioShake describes the underlying system as acoustic rather than dependent on a language model, and says it supports different sample rates and capture conditions. The Refinery also scores outputs for quality and confidence, separating correct speaker assignment from clean separation. That allows large collections to be sorted into usable, fixable, or unsuitable material.
The launch builds on AudioShake experience in music and media production, including mixing, localization, audio analysis and audiovisual editing. Its website lists customers such as ESPN, Universal Music Group and Warner Bros. Studios, where separation makes a finished mix useful for another workflow. The Refinery applies the same principle to AI development, alongside dataset preparation from customer content and specialized models for particular catalogs. Early versions were deployed privately over the past year, with named customers Luel and Rime plus unnamed labs and data marketplaces.
What this means for voice AI development
For companies building voice assistants, transcription tools or call analytics, separated tracks can make labeling, inspection and transcript attribution more precise. Teams can review speaker identity, track clarity and transcript accuracy separately instead of treating a dataset as one pass-or-fail result. Quality scores help direct human review toward crowded segments where a quiet speaker is hard to distinguish. Small firms gain access to structured conversational data without recording studios, while large firms can process substantial archives through API or on-premises deployment.
The approach has clear limits that buyers should verify on their own audio. AudioShake reports 9.17% concatenated minimum-permutation word error rate on LibriCSS versus 37.75% for the tested MERL TF-Locoformer checkpoint through Whisper large-v3, but the baseline was tested outside its training domain. The company frames this as an off-the-shelf comparison, not proof of superiority under matched conditions. Accuracy also degrades with sustained overlap and higher speaker counts, and cleaner training data does not by itself establish how much a conversational model will improve.
The marker to watch is whether customers publish results from representative trials showing quiet speakers surviving separation, interruptions staying aligned, and measurable gains in the target voice application. Further adoption by frontier labs, data marketplaces and on-premises deployments for sensitive recordings would confirm operational value. If such cases appear, existing rights-cleared archives will become a more practical foundation for models that handle real conversational behavior.
