Research data handling for US institutions
Written for IRB applications and campus IT review. The questions on those forms are about who touches a recording and what happens to it, so that is the order this answers them in.
Who transcribes the recording
One named provider, and no human being. Audio goes to AssemblyAI, which returns a machine-generated transcript. No person at Phonotheca and no person at AssemblyAI listens to a recording as part of transcribing it, and there is no contractor pool, no crowd-work platform, and no marketplace of individual transcriptionists involved.
This is usually the question underneath “who will have access to the recordings” on an IRB form. The answer is your research team, plus automated processing by one named vendor under contract.
What happens to the recording afterwards
- The transcript is deleted from the transcription provider once it lands in your workspace.
- Audio and transcripts stay in your workspace until you delete them. Deleting an interview or a project removes both.
- Deleting your account removes your content and your account data.
- Usage records survive deleting an interview, holding a duration and dates and never audio or text, so an allowance cannot be reset by deleting and re-uploading. They are deleted with the account.
Training
Phonotheca does not train any model on your recordings, transcripts, or tags, and does not sell or share them. We build no models and derive no dataset from customer content. US classification frameworks generally gate on training-data use rather than on location, so this is the clause most campus reviews turn on.
Which data classification this suits
Most US institutions sort research data into tiers, whether that is Low and Moderate Risk or Level 1 through Level 3 or a local equivalent. Phonotheca is appropriate for de-identified and low-sensitivity human-subjects recordings, and for identifiable interview audio where your protocol permits a named third-party processor under contract.
It is not appropriate for protected health information under HIPAA, because we do not currently sign a business associate agreement. It is also not in scope for export-controlled research, ITAR, EAR, or controlled unclassified information. We cannot serve those and would rather say so here than during a review.
What we do not have
Phonotheca holds no SOC 2 report, no ISO 27001 certificate, no HIPAA business associate agreement, and no FedRAMP authorisation. Several competitors hold some combination of those. If your process requires one from the vendor itself, we do not meet it today.
Our infrastructure and transcription providers publish their own audits, which cover their part of the pipeline rather than ours.
Where processing happens
In the European Union. Application, storage, and database run in Frankfurt, Germany, and speech-to-text runs in Dublin, Ireland. For most US research this is neither a requirement nor an obstacle, and it is stated here because review forms ask where data is stored.
If your institution has a contractual requirement for onshore storage, write to support@phonotheca.com before you build it into a protocol, because we would rather tell you it is not available than have it surface after approval.
Wording you can paste
For an IRB application describing third-party transcription.
Audio recordings will be transcribed using Phonotheca, a hosted transcription and analysis service, under a signed data processing agreement. Transcription is performed by automated speech recognition with no human review by the vendor or any subcontractor. Recordings and transcripts are accessible only to the research team. The vendor does not use customer recordings or transcripts to train models. The transcript is deleted from the speech recognition provider once retrieved, and recordings are deleted from the workspace at the researcher’s direction.
If a reviewer asks something this does not cover, write to support@phonotheca.com and we will answer in writing so you can attach the reply.
The full statement
The data governance statement carries the complete sub-processor list with countries, the controller and processor split, and retention in full. The privacy notice covers account data and your rights.