Engineering
Speech-To-Text
Voice AI
Why Better Audio Made Our Speech-to-Text Worse
Arne Vervaet

Our first engineering post: how one technician's garbled voice notes taught us that "high quality" audio and good input for a speech-to-text model are not the same thing.
Voice notes are at the heart of Highsail's voice AI. A technician finishes a job, taps the microphone, and talks. Speech-to-text turns that into a transcript, and the transcript becomes structured data. If the transcript is wrong, everything built on top of it is wrong too.
Most of the time this works well. So when one technician's transcription accuracy suddenly dropped, it stood out immediately.
One technician, and only one
The report was short: this technician's transcripts made no sense. Some were garbled from start to finish. One began correctly and then ended with phrases that had no basis in what the technician said.
The strangest part was how much the results varied. Transcribe the same recording twice and you got two different stories. Some variation between runs is normal, but for clean audio, repeated runs should tell substantially the same story. These didn't come close. Something in the audio was confusing the model.
Other technicians used the same app, the same pipeline and the same model, and none of them had this problem. We had an outlier.
The test that lied
The obvious first step was to listen to the recordings. They sounded fine: clear speech, some background noise, nothing unusual.
Then we tried a replay test: play the file through a laptop's audio setup, capture it again, and transcribe the result. It came back clean. Only the original file failed.
That "successful" test turned out to be the most important clue. The playback and capture path wasn't a neutral pipe. It altered the signal through filtering and resampling. What we transcribed sounded like the same audio, but it was no longer the same signal. The path had removed exactly what the model was tripping over, so our test fixed the bug without us noticing.
That's the lesson at the centre of this post: listening to a file doesn't tell you what a model hears.
What was hiding in the audio signal
Once we stopped trusting our ears and looked at the full frequency spectrum, the picture changed. Well above the range where speech lives, the recordings carried a strong band of noise. It was easy to miss when you were listening to the speech, but it was very much there in the data.
The noise appeared in every failing recording from this technician, across different days and locations. Noise that follows the recordings from place to place points to the recording device, not the environment.
That's less surprising than it sounds. A phone packs a lot of electronics close to its microphone, and some of it can introduce electrical noise into the audio path. Worn or damaged microphones, cases and accessories can add their own artefacts. Two phones that sound identical in a call can produce very different signals at the top of the spectrum. We believe that's what happened here: this technician's device was adding its own noise, well above the speech band, where no one would hear it.
That's why we saw an outlier and not a trend. Field teams use a wide variety of phones, often for years, and most will never cause this. But when you record voice from real devices in the field, you can't assume clean input. The pipeline has to cope with whatever the hardware adds.
Why audio aliasing breaks speech-to-text
Why would noise well above the speech band damage a transcript? Our best explanation is a classic effect from signal processing called aliasing. We haven't confirmed exactly what happens inside the model's pipeline, so treat this as a well-supported hypothesis, not a proven diagnosis.
Digital audio is a series of samples, and the sample rate sets a ceiling, the Nyquist limit, at half that rate. In an ideal sampled system, frequencies below that ceiling can be represented without aliasing. Frequencies above it cannot.
This matters when audio is converted to a lower sample rate, which happens often in real-world pipelines. A well-designed resampler filters out content above the new ceiling first. But if that content survives into the conversion and the filtering isn't adequate, it doesn't disappear. It folds back into the remaining range and shows up as a false signal at a lower frequency. That's the same effect that makes wagon wheels appear to spin backwards on film.
So noise that was harmless far above the speech can land right in the middle of it after a conversion. To the model, it's no longer background. It's interference in the frequencies it relies on to recognise words. That fits what we saw: garbled words, phrases with no clear basis in the speech, and results that shifted from run to run.
What we know for certain is narrower: removing the high-frequency content fixed the problem. Aliasing is our best explanation of why, but other steps in the processing could also play a part.
More isn't always better for speech recognition
Our app had captured voice notes with a general-purpose "high quality" setting, the kind that makes sense for music. More detail sounds like it should always help. But the information that makes speech understandable sits in a fairly narrow band. Much of what sits well above the speech band adds little useful information for recognition, while giving unwanted noise more room to enter the pipeline.
That doesn't make a higher sample rate harmful in itself. Well-captured, high-rate audio can be excellent input. The problem was the combination: noise from this particular phone, and the way the audio was processed after recording.
Before settling on a fix, we compared what each approach did to the same failing recordings:
Approach | Effect on the transcript |
|---|---|
Original "high quality" capture | Garbled, and different on every run |
Replaying through a laptop | Clean, but only by accident: the replay had changed the signal |
Filtering out the high frequencies | Accurate, and the same on every run |
Capturing at a rate suited to speech | Accurate, and the same on every run, with smaller files |
Filtering and a speech-suited capture rate both worked. We chose the second: it prevents the problematic high-frequency content from reaching the rest of the pipeline, and the files are smaller.
We tested the change end to end, and the affected technician's transcripts went from unreliable to consistent. Because it applies to every recording, it also protects other users from the same class of problem.

Top: the average spectrum of three voice notes from the affected phone. Bottom: the same recordings brought down to the new capture setting. Below the cutoff nothing changes. The range where the device noise sat is simply no longer captured.
What we took away
Match the input to the model, not to the word "quality." A signal can be excellent by one measure and harmful by another.
Check what your test setup does to the data. A test that transforms its input can hide the bug it's meant to find.
Outliers are worth chasing. One technician's problem exposed a weakness that could have hit anyone with the wrong combination of hardware and noise.
Our users record voice notes in vans, basements, plant rooms and on rooftops, and each place brings its own noise. The model should hear what they said and nothing else. Sometimes that means capturing less.
