Skip to content
Data

Voice AI’s Training Data Problem Is Hiding Between the Words

9 min read · 8:51 PM ET

AudioShake’s October 8 launch reveals a voice AI training problem: transcripts can lose the overlap, corrections and timing that make a conversation work.

AudioShake co-founders Luke Miner, left, and Jessica Powell standing together in a company portrait.
AudioShake co-founders Luke Miner and Jessica Powell in an archival company photograph. AudioShake introduced The Refinery on October 8, 2026.Courtesy of AudioShake · Company press kit
Key takeaways
  • Speech recognition, speaker attribution and conversational turn-taking need different information from the same recording.
  • Recent speech research shows why difficult conversations can be valuable when their labels and timing are reliable.
  • A useful audio data offer explains what preparation preserves, what it changes and which uses are permitted.

A customer says yes while an employee is still explaining the question. A second later, the customer interrupts to correct a detail. A polished transcript can turn that exchange into two orderly sentences. For a voice model, the missing second may be the most useful part of the recording.

On October 8, AudioShake introduced The Refinery, a service that separates existing recordings into speaker tracks and other audio components for AI training. The company says it can work from a finished mix, attach quality scores and recover voices that overlap. It reports processing more than 100 million minutes through earlier private versions. That is a company-reported processing total, not a measure of how much any particular model learned.

The announcement points to a question that gets lost when training data is counted in hours: which parts of an interaction survive the preparation? A recording, a transcript and a summary can describe the same call while teaching a model different things. One preserves the sound. Another preserves the words. The third preserves an editor’s account of what mattered.

For businesses that hold recorded conversations, this makes the original recording more than a container for text. Its timing can show how people correct each other, acknowledge an explanation and decide when to speak. Keeping those relationships intact creates different possibilities from simply producing a larger pile of transcripts.

The Training Example Has a Clock

Speech recognition is the task of turning sound into words. A model trained for that job needs examples connecting the sound to an accurate transcription. Speaker labeling adds another question: who produced each part? A conversational system faces a further problem: when should it answer, and when should it keep listening? These jobs overlap, but success at one does not establish success at all three.

Consider a hypothetical recording in which an employee asks, “Would you like Friday morning?” The customer begins saying “Friday” before the question ends, then adds “afternoon, actually.” A transcript that keeps the correction can communicate the final preference. A summary might reduce everything to “Friday afternoon.” Neither necessarily preserves the moment when the employee had enough information to respond.

The timing matters because a system used in a live call cannot read the future. If a training example includes the final preference before the corresponding sound has arrived, it can make a task look easier than it will be in use. A useful example must preserve what was available at each point, including uncertainty that the next words eventually resolve.

That is a different way to think about a dataset. The unit is an unfolding interaction, with words and actions arranged on the same clock. An isolated sentence can still be useful for learning pronunciation or vocabulary. It cannot, by itself, show how one person adjusted to another person’s interruption.

What a Recent Benchmark Reveals

David AI’s September 30 DAI-ASR-I18N report, updated October 2, tests speech systems on conversations in 21 languages. Participants spoke on separate synchronized channels, and the transcripts retained repetitions and false starts. The researchers found that overlapping speech was the hardest condition across both transcription and speaker-labeling tests. This is evidence about the tested recordings and systems, not a diagnosis of every voice model’s training set.

For a buyer, that distinction changes the sample review. A clear recording of one person reading a prepared sentence may be accurately labeled and still leave the buyer’s hardest problem untouched. An interrupted conversation may be more relevant, provided the buyer can establish who said what and when. “Clean data” needs a definition tied to the job.

The benchmark also supplies a useful boundary. Its data-use terms permit evaluation, including of commercial systems, but prohibit training and fine-tuning. These conversations can reveal where a model struggles without becoming its next lesson. An evaluation set and a training set should not be treated as interchangeable inventory.

Tidying Is an Editorial Decision

Preparing speech data involves choices that can look small on a spreadsheet. Should a repeated word remain? Should a laugh receive a label? Should an unfinished phrase be completed for readability? Should two short pauses count as silence or as part of one speaking turn? Each choice changes the answer the model is asked to learn.

A dataset listed on Mozilla Data Collective on October 6 offers a concrete example. Its 8,000-utterance Javanese subset comes from an older OpenSLR corpus, rather than newly recorded October speech. The contributor describes using the audio and transcripts to adapt speech-recognition models for a competition. The package retains both original transcripts and the versions used in training, documenting changes to spelling, diacritics and annotation tags.

That publication is useful because the transformation is visible. A researcher can distinguish what a person said, how it was written in the original collection and how it was represented for the particular task. The same principle applies to a company’s English-language call archive. A final transcript alone may conceal several decisions made between recording and delivery.

Suppose an internal team removes every hesitation to make a customer-service transcript easier to scan. That is a reasonable choice for some business uses. A buyer studying hesitation before an interruption would need a different version. There is no need to declare one universally superior. The mistake is to deliver either version without knowing which task it is supposed to support.

Keeping a transformation history makes more than one use possible. A company can retain its readable transcript while also preserving an approved version with timing and fuller speech detail. The relationship between those versions is useful information in its own right. It tells a buyer which differences came from the speaker and which came from preparation.

Recover the Sound. Keep the Reference.

Audio separation attempts to recover individual sound sources from a mixture. It does not amount to discovering an untouched original microphone channel. AudioShake says The Refinery extracts recorded sound rather than generating replacement speech. That product claim matters, but a buyer still needs to inspect the result against the original recording. A confidence score can help choose what to review; it cannot make every uncertain word certain.

The practical question is whether preparation preserves the distinction a model needs. If one voice leaks into another speaker’s track, a model may receive an incorrect pairing between sound and identity. If processing removes a quiet acknowledgment, the example may become easier to transcribe while losing a conversational cue. These are reasons to compare versions, not reasons to reject separation as a method.

The research history helps explain why a good demonstration is only a starting point. Microsoft’s earlier LibriCSS study evaluated continuous recordings containing both overlapping and non-overlapping speech. Its dataset replayed assembled utterances through a room and recorded them with distant microphones. The authors also cautioned that improvements in audio-signal measures did not closely track transcription accuracy. The study dates to 2020; it is background for evaluating today’s tools, not a new October result.

For a business buying prepared recordings, a useful acceptance test therefore follows the intended task. Can a reviewer recover the important correction? Are speaker changes attached to the correct points in time? Does the preparation improve difficult passages without damaging the straightforward ones? The answers require a representative sample, including examples that the supplier’s process finds difficult.

What a Company Actually Has to Offer

A company with recorded calls might assume its main asset is the number of hours. A more informative description would explain the interactions those hours contain: the languages spoken, how the recording was captured, whether the speakers have separate channels, whether corrections are preserved and whether a useful result can be connected to the conversation.

One Conversation, Several Possible Learning Tasks

Material retainedPossible learning taskWhat preparation must preserve
Sound and verbatim wordsRecognize what was saidQuiet speech, corrections and uncertain passages
Synchronized speaker tracksSeparate and attribute voicesOverlap and the shared timeline
The conversation as it unfoldedStudy turn-takingWhat each participant had heard before responding
A conversation linked to its resultStudy whether an exchange workedThe connection between the request, correction and outcome

The right preparation depends on what the model is supposed to learn.

SnowRock analysis. These are illustrative uses, not claims about a particular buyer’s requirements.

There is a cost to producing that package. Someone must examine recording formats, check labels, review permissions and correct failures. Some material will not justify that work. A low price per raw hour can become expensive if only a small share survives review. Conversely, an archive with reliable channels and well-maintained timing may require less repair, even if its headline volume is smaller.

The commercial comparison should include that difference. Is the buyer paying for access to recordings, for reviewed transcripts, for separated tracks, or for a maintained collection that adds new examples? Those are different deliveries. Pricing them as if they were the same can leave the preparation team doing work that neither side included in the deal.

Permission is another property of the delivery. A company should establish which conversations it can offer for the proposed use before treating every recorded call as available training material. Separating a voice or removing background music does not itself settle that question. This is especially relevant where an archive was originally created for internal quality review rather than outside model development.

The Conversation Is the Record

The most interesting implication of this week’s launch is that the preparation step can determine what a model has the opportunity to learn. The same recorded exchange can become a pronunciation example, a speaker-separation example, a turn-taking example or a short piece of text. None is a complete substitute for the others.

This gives companies a practical reason to understand the recording systems they already use. Before replacing a rich archive with a tidy text export, establish what will disappear. A future buyer may care about details that were irrelevant to the original reporting dashboard. Preserving useful structure does not guarantee a sale, but discarding it can narrow the questions the data is capable of answering.

An AI voice can sound fluent and still enter a conversation at the wrong moment. Teaching it to participate requires attention to the interaction itself: the overlap, the correction, the pause and the point at which an answer became possible. Those details often sit between the words people remember to count.

All Insights

More from SnowRock

Microsoft CEO Satya Nadella in a studio portrait
Technology

Microsoft’s New Copilot Will Test How Your Company Works.

The red facade of No. 3 Circle Square in Manchester as the building neared completion
Technology

The UK’s AI Privacy Review Puts Training Data and Autonomous Agents Under Scrutiny.

Anthropic co-founder Dario Amodei gesturing during a discussion at TechCrunch Disrupt
Strategy

AI Is Getting Cheaper. Finished Work Is the Real Price.