
Transcribing research interviews: when AI is enough and when you need a human

SUMMARY SNIPPET
- AI transcription is enough for most research interviews when the audio is clean, two or three people take turns, and you check the draft against the recording.
- A human transcriber is worth paying for when your method analyses how things were said, when the audio or the speakers are hard for speech recognition, or when a quote must be exact.
- For most qualitative studies the best route is hybrid: AI writes the draft, then a person corrects it, applies the verbatim rules and anonymises it.
- Whatever the route, treat recordings as personal research data and check where they are stored, whether they train AI and how they are deleted.
Table of Contents
AI transcription is enough for most research interviews when the audio is clean, two or three people take turns, and you check the draft against the recording. A human transcriber is worth paying for when your method analyses how things were said, when the audio or the speakers are hard for speech recognition, or when a quote must be exact. For most qualitative studies the best route is hybrid: AI writes the draft, then a person corrects it, applies the verbatim rules and anonymises it. Whatever the route, treat recordings as personal research data and check where they are stored, whether they train AI and how they are deleted.
Kadri is a doctoral researcher in Tallinn studying how small manufacturers adopt digital tools. She has 24 semi-structured interviews to run this term, most in Estonian, some in English, and many where participants switch to English for technical terms halfway through a sentence. Her ethics approval requires recordings to stay in the EU and to be deleted when the project ends, and her supervisor wants exact quotes in the thesis. A professional service for every hour is beyond her budget, and typing it all by hand would take most of the semester.
Why is transcribing research interviews still such a bottleneck?
Transcription is where qualitative research slows down. Researchers report five to six hours of work to transcribe one hour of interview verbatim, and up to 60 hours for detailed notation, according to a 2024 comparison of automatic transcription tools in Forum: Qualitative Social Research. In that comparison, the best tool reached 93% word accuracy on an English interview but only just over 70% on a German one, and the authors concluded that every automatic transcript still needs a review against the recording.
Accuracy also depends on who is speaking. A Stanford study published in PNAS found that five commercial systems got about 35% of words wrong for Black American speakers against 19% for white speakers, and a 2024 study of a widely used open-source model found that roughly 1% of transcriptions contained phrases nobody said. The useful question is not "Is AI accurate?" but "Is it accurate enough for this study, these participants and this method?"
Is AI transcription accurate enough for research interviews?
For many studies, yes, as a first draft. Bokhove and Downey (2018) found that most differences between automatic and manual transcripts were easy to fix in a review, and called the output "good enough" for first versions. AI is usually enough when:
- the audio is clean and recorded close to the speakers;
- two or three people take turns rather than talk over each other;
- your analysis works with what was said, as in thematic or applied research;
- you plan a review pass before anything is coded or quoted.
Expect dropped fillers, misheard names and jargon, and the occasional invented phrase. The quickest test is to run one interview you have already transcribed by hand through the tool and compare the two; our transcript accuracy benchmark shows one way to structure it.
How Digiotouch AI handles this: Digiotouch AI transcribes in more than 130 languages, including Estonian, Finnish and French, and detects when a speaker switches language mid-interview, so no language needs to be set in advance. The full transcript, available from the Essential plan, is editable, so the review happens where the draft is. The Free plan gives summaries, not transcripts.
When should you pay for a human transcriber instead?
Pay for a human when the transcript is the data itself, not just a record of the conversation. That usually means:
- Your method analyses how people speak: conversation or discourse analysis needs timed pauses, overlaps and intonation.
- The audio is difficult: focus groups with crosstalk, noisy rooms or a poor phone line.
- Speech recognition serves your participants poorly: strong dialects, children, speech difficulties or a less widely spoken language.
- A mistake would be costly: published quotes, legal or policy evidence, or sensitive accounts.
A human transcriber also handles personal data, so the UK Data Service recommends written transcriber instructions, a non-disclosure agreement and encrypted file transfer.
How Digiotouch AI handles this: Digiotouch AI produces an editable first draft, not conversation-analysis notation, and will not replace a specialist for that work. It shortens the specialist's job instead, because correcting a draft is faster than typing from scratch. Overlapping speech is hard for any automatic tool, Digiotouch AI included, so give crosstalk a human pass.
What is the difference between verbatim and clean verbatim transcription?
Full verbatim keeps every utterance; clean verbatim keeps every meaningful word. Decide before the first interview and write the choice into your transcription guidelines.
- Full verbatim: fillers, repetitions and false starts, plus events such as [laughs] or [pause]. Needed when hesitation is part of the analysis.
- Clean verbatim: fillers and stutters removed, grammar left as spoken. The usual choice for thematic and applied research.
- Edited transcript: tidied prose. Fine for meeting notes, rarely for research data.
Methodologists call this naturalised versus denaturalised transcription, and neither is better in general: the research question decides (Oliver, Serovich and Mason, 2005). AI drafts usually sit closer to clean verbatim, so budget a human pass if your method needs the fillers back.
How Digiotouch AI handles this: Each turn in the transcript is attributed to a speaker and kept in the language it was spoken. Every part of the transcript is editable, so you can bring the draft to your verbatim level and mark passages for anonymisation. The full transcript can be exported as a PDF or Microsoft Word document right from the meeting summary page.
How does a hybrid AI and human workflow work in practice?
AI drafts, a person decides. Seven steps cover most interview studies:
- Record well: quiet room, microphone between you and the participant, names said at the start.
- Generate the AI draft: record into the tool, or upload the file afterwards.
- Review against the audio: fix names, terms and speaker labels, and check closely after long pauses.
- Apply your conventions: add fillers, pauses or overlaps if your method needs them.
- Anonymise: replace identifiers, or mark them in square brackets for later.
- Export and analyse: move the checked transcript into your analysis software.
- Delete on schedule: remove recordings and copies on the date your data management plan sets.
How Digiotouch AI handles this: The iOS and Android apps record in-person interviews on your phone (how in-person recording works), and the desktop app records remote ones from your own device. Existing MP3, MP4, WAV or WebM files become an editable meeting page within minutes, and sessions run up to 5 hours on every plan. Export as PDF or DOCX once the review is done.
Where should research interview recordings be stored, and for how long?
Store them where your ethics approval and data management plan say, and only for as long as they say. Interview recordings are personal data under the General Data Protection Regulation, research interviews often touch the Article 9 special categories such as health or political opinions, and Article 89 expects safeguards such as pseudonymisation. Before uploading, ask any provider, human or automated:
- Where are recordings processed and stored, and can they stay in the EU?
- Are they used to train AI models?
- Does deleting a recording remove the audio itself?
- Is there a data processing agreement for your ethics file?
Tell participants too: name the automated transcription, the storage location and the deletion date in your information sheet and consent form.
How Digiotouch AI handles this: Digiotouch AI stores recordings of EU and EEA users in the EU by default and does not train AI models on your content. Deleting a recording or an account removes the stored media, not just the database entry. Institutions that need data on their own infrastructure can ask about the Enterprise on-premise and private cloud options.
AI, human or hybrid: which transcription route fits your study?
The table sums up the trade-offs for a typical semi-structured interview, including where a human transcriber clearly wins.
| AI only | Human only | Hybrid (AI draft, human review) | |
|---|---|---|---|
| Time per interview hour | Minutes for the draft | Around 5 to 6 hours of work for verbatim | Minutes for the draft, plus a review pass |
| Cost | Lowest | Highest | Low to moderate; mostly researcher time |
| Accents, dialects, speech difficulties | Varies widely; pilot first | Strongest, especially with a specialist | Good if the reviewer knows the language variety |
| Crosstalk and focus groups | Weakest | Strongest | Human pass on overlapping passages |
| Full verbatim and notation | Rarely; fillers are often dropped | Yes, if your guidelines specify it | Added by the reviewer where needed |
| Best for | Thematic and applied research, user interviews | Conversation analysis, difficult audio, high-stakes quotes | Most qualitative interview studies |
Sources: Wollin-Giering et al., Forum: Qualitative Social Research (2024); Bokhove and Downey, Methodological Innovations (2018); UK Data Service, Transcription. Digiotouch AI makes one of the automatic tools in this category, so the table compares routes, not products.
Key takeaways
- AI makes a strong first draft: clean audio and few speakers turn hours of typing into a review.
- Always review against the recording: AI drops fillers, mishears names and occasionally invents words.
- Pay a human when the transcript is the data: conversation analysis, crosstalk, difficult audio, published quotes.
- Hybrid suits most studies: AI draft, human review, consistent verbatim rules, anonymisation, scheduled deletion.
- Treat recordings as research data: check storage location, AI training and deletion before you upload.
Next step
Test the hybrid route on one interview you have already transcribed by hand: upload the recording, compare the AI draft with your own transcript, and time the review. You will need a plan with the full transcript, which starts at Essential.
Related reading
- Copilot vs Digiotouch AI: how accurate is your meeting transcript, really?: a line-by-line accuracy test on three technical recordings.
- Why can't AI meeting tools record in-person meetings?: recording face-to-face conversations on your phone, with speakers labelled by name.
- How to turn a recorded meeting into a shareable summary: what happens when you upload a recording you already have.
Citation index
- Wollin-Giering, S., Hoffmann, M., Höfting, J. and Ventzke, C. (2024). Automatic Transcription of English and German Qualitative Interviews. Forum: Qualitative Social Research, 25(1): qualitative-research.net/index.php/fqs/article/view/4129
- Bokhove, C. and Downey, C. (2018). Automated generation of "good enough" transcripts as a first step to transcription of audio-recorded data. Methodological Innovations: eprints.soton.ac.uk/422043
- Stanford News (2020). Automated speech recognition less accurate for blacks (reporting Koenecke et al., Racial disparities in automated speech recognition, PNAS): news.stanford.edu/stories/2020/03/automated-speech-recognition-less-accurate-blacks
- Koenecke, A., Choi, A. S. G., Mei, K. X., Schellmann, H. and Sloane, M. (2024). Careless Whisper: Speech-to-Text Hallucination Harms. ACM FAccT 2024: arxiv.org/abs/2402.08021
- Oliver, D. G., Serovich, J. M. and Mason, T. L. (2005). Constraints and Opportunities with Interview Transcription: Towards Reflection in Qualitative Research. Social Forces, 84(2): pmc.ncbi.nlm.nih.gov/articles/PMC1400594
- UK Data Service. Transcription (research data management guidance): ukdataservice.ac.uk/learning-hub/research-data-management/format-your-data/transcription
- Regulation (EU) 2016/679, General Data Protection Regulation, EUR-Lex: eur-lex.europa.eu/eli/reg/2016/679/oj/eng
External sources checked on 7 October 2026. Send corrections to notes@digiotouch.ai.
As a first draft, usually yes; as a final transcript, no. Accuracy is highest with clean audio, few speakers and widely supported languages, and drops with crosstalk, strong accents and specialist terms. Published comparisons recommend reviewing every automatic transcript against the recording before coding or quoting it.
Full verbatim records everything as spoken, including fillers, repetitions and false starts. Clean verbatim keeps every meaningful word but removes fillers and stutters. Use full verbatim when how people speak matters to your analysis, and clean verbatim for most thematic research.
Researchers commonly report five to six hours of work for a verbatim transcript of a one-hour interview. Detailed notation for conversation analysis takes far longer, with estimates of up to 60 hours per hour of recording.
Usually, yes. Professional transcription is typically charged per minute of audio, while AI transcription comes with a subscription, so the main cost of a hybrid route is your review time. Digiotouch AI includes the full transcript from the Essential plan; current prices are at digiotouch.ai/en/pricing.
Yes. Digiotouch AI detects when a speaker changes language during a recording and keeps transcribing, with no language to set in advance. Each passage appears in the language it was spoken, and the summary can be translated into any supported language.
Yes. Digiotouch AI accepts MP3, MP4, WAV and WebM files and turns each one into an editable meeting page within minutes. The full transcript is included from the Essential plan.
No. Digiotouch AI does not train AI models on customer content, including interviews you record or upload. Recordings of EU and EEA users are stored in the EU by default, and deleting a recording removes the stored media itself.



