Audio and voice AI jobs: tasks, pay and who qualifies
Overview · 1 week ago
Audio and voice AI jobs cover recording, transcription, speech data auditing, listening QA and audio engineering. From 107 listings on this site: what each type involves, what it pays (median midpoint $28 an hour, voice cloning up to $150) and what you need to qualify.
What audio and voice AI work involves
Speech recognition, text-to-speech and voice assistants all learn from audio that people have recorded, transcribed or graded. On 29 September 2026 this site listed 107 audio and voice AI jobs under Audio & Voice, out of 699 listings in total. It is the most accessible field on the board: every listing is rated entry, junior or mid level, and none senior.
The work splits into five types.
Recording. You read scripts or speak prompts into a decent microphone. The best-paid version is voice cloning: Mercor's voice actor listings pay $50–150 an hour for a session of about four hours, and the recordings build a synthetic copy of your voice. Read our privacy guide before licensing yours.
Transcription. micro1's language-specific transcription listings, such as Norwegian, turn audio into text to strict spelling and formatting rules, and pay roughly $10–35.
Auditing speech data. Mercor's Sonic Audit listings, for example German at $41.50, check other annotators' transcripts against error codes and correct word-level timestamps.
Listening QA. Mercor's audiobook QA family pays $15–20 to mark every skipped word, mispronunciation and glitch in AI-narrated books.
Audio engineering. micro1's Audio Engineer listing labels background noise, clipping, distortion, echo, hum and dropouts in recordings, at $50–90, and asks for professional or academic audio experience.
Inside transcription and auditing work, the same few tasks keep coming up. Speaker diarization means marking who spoke and when, usually as Speaker 1, Speaker 2 and so on. Timestamping pins words, phrases or speaker turns to the right point in the recording. Audio classification tags a whole clip: speech or music, crosstalk, a synthetic voice, too much background noise. Some projects also ask you to mark code-switching, where a speaker moves between languages mid-sentence. And a growing share of the work is correcting a transcript that an AI model has already drafted rather than typing one from scratch, which rewards a sharp ear for small errors more than typing speed.
Read the task description, not just the job title. A listing called transcription can turn out to need annotation software, a long list of labels and judgment calls on unclear audio.
What it pays
104 of the 107 quote an hourly rate, from $6 to $150. The median midpoint is $28 an hour, one of the lowest on the site alongside general language work. Voice cloning and audio engineering sit at the top; transcription, audiobook QA and Invisible Technologies' flat $17 voice evaluation family at the bottom. Two Invisible speaker projects pay $1.30 per unit instead of by the hour.
Projects that pay per unit or per audio minute need a different sum. A minute of clear, single-speaker audio can go quickly, but a minute of noisy audio with several people talking over each other, plus timestamps and labels, can take many times its length once you count replays and checking. Before you accept a per-minute rate, time yourself on a sample clip like the ones the project uses and work out the hourly figure from that.
Language drives much of the spread, because the same template can pay very differently by market: Mercor's Sonic Audit pays $41.50 for German and $18 for Hindi. Our bilingual guide covers that pattern in detail.
Who qualifies
The main requirement is a native ear. 93 of the 107 listings mention native or native-level fluency, and many name a specific variety, such as German as spoken in Germany or Brazilian Portuguese. Equipment is the second filter: recording roles want a quality microphone, a quiet room and sometimes a pop filter. Voice-cloning roles also require you to live in the country of the accent.
Beyond that, careful listening and consistency count for more than credentials. Transcription and annotation experience is a strong signal, and formal qualifications are rarely asked for outside audio engineering. 91 of the listings are open worldwide. micro1 posts 50, Mercor 43 and Invisible Technologies 14.
Most transcription and annotation roles screen you with a sample task, and it tests rule-following more than spelling. Projects usually set their own rules for the fine details, so read the guidelines before you open the audio and keep them open while you work. The points they typically settle: how to mark speech you are unsure of versus speech that is inaudible, whether to keep false starts and filler words, how to tag laughter, coughs and other non-speech sounds, and what to do when a speaker switches language. Follow the project's convention every time, even where it looks odd next to ordinary written text.
The false starts and filler words question is really a choice between two transcription styles, and the project's style guide makes it for you. Full verbatim, often just called verbatim, records everything the speaker produced: filler words such as um and uh, false starts, repeated words and stutters. Clean verbatim takes those out but keeps the speaker's meaning and their own wording; it is not a paraphrase or a summary. Here is the same utterance in each style:
- Full verbatim: "So, um, I I think we should, uh, we should leave at six."
- Clean verbatim: "So I think we should leave at six."
Some projects mix the two, for example keeping stutters but dropping filler words, so match the worked examples in the guidelines rather than a textbook definition. If the guidelines do not say which style applies, ask before you submit.
Where to start
Filter the audio and voice listings, then narrow by your language. Check the language variety and the equipment line before applying, since both are hard requirements.
Questions
- How much do audio and voice AI jobs pay?
- Across 104 hourly audio and voice listings on this site on 29 September 2026, rates ran from $6 to $150 an hour with a median midpoint of $28. Mercor's voice-cloning sessions paid $50 to $150, micro1's Audio Engineer $50 to $90, transcription roughly $10 to $35 and audiobook QA $15 to $20.
- Do I need experience for AI voice recording or transcription work?
- Rarely formal experience. Every audio and voice listing on this site was rated entry, junior or mid level. The main requirements are native or near-native command of a specific language variety, careful listening, and for recording roles a quality microphone and a quiet room.
- What is AI voice cloning work?
- You record scripts, typically in one session of about four hours, and the recordings are used to build a synthetic version of your voice for a client. Mercor's voice actor listings pay $50 to $150 an hour and require you to live in the country of the accent. Read the licensing terms carefully before agreeing.
Platforms covered here
Put this into practice
Every listing shows its pay and who it is open to.