> For the complete documentation index, see [llms.txt](https://ailyze.gitbook.io/ailyze-docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://ailyze.gitbook.io/ailyze-docs/prepare-research-data/transcribe-audio-and-video.md).

# Transcribe Audio & Video

This page shows you how to turn audio and video recordings into [transcripts](/ailyze-docs/reference/glossary.md#transcript) you can review, edit, download, and analyze. Evidano transcribes your files, labels the speakers, lets you fix mistakes, and hands off to analysis in one click.

{% hint style="info" %}
**What you need:** An Evidano account and your audio or video files. Your first 60 minutes of audio are free.

**What you will have at the end:** A timestamped, speaker-labeled transcript for each recording, ready to edit, download, or send into analysis.
{% endhint %}

## Step 1: Upload your recordings

From your account page, click **"Transcribe Audio/Video"**. This opens the **New Transcription Project** page. Add files from the **"Local Files"**, **"Google Drive"**, **"Dropbox"**, or **"OneDrive"** tabs, or import Zoom cloud recordings directly.

* Evidano accepts a wide range of audio and video formats (mp3, wav, m4a, mp4, mov, and more). See [Supported File Types](/ailyze-docs/reference/supported-file-types.md) for the full list.
* Each file can be up to 5 GB, and you can select several files at once.

<figure><img src="/files/fy5SUDhmIk0S7XaqwdpE" alt="Uploading audio and video files on the New Transcription Project page"><figcaption></figcaption></figure>

## Step 2: Configure each file

Each uploaded file gets its own settings panel. If you uploaded several files, use **"Apply this settings to all files"** to copy your settings across.

### Pick the Language of Audio (required)

Choose the spoken language under **Language of Audio**. Options are grouped by accuracy (**Best / Better / Good**), plus a **Multilingual** option for recordings that mix two or more languages. For the full list, see [Supported Languages](/ailyze-docs/reference/supported-languages.md).

<figure><img src="/files/Zj0ihWgYuSSgrBOG65cD" alt="Choosing the Language of Audio for an uploaded recording"><figcaption></figcaption></figure>

### Add a translation (optional)

Use **Translate To** to also receive a translated transcript. For multilingual recordings, transcribe first, then translate separately with [Translate Files](/ailyze-docs/prepare-research-data/translate-files.md).

### Optional Settings

Expand **Optional Settings** for finer control. You only see the options your chosen language supports.

**Improve Transcription Accuracy**

* **Custom Vocabulary** — add jargon, names, and acronyms (spell numbers out: "seven", not "7").
* **Custom Spelling/Formatting** — map spoken words to how you want them written, like "sequel" to "SQL".
* **Custom Dictionary** — brand, company, person, and place names.
* **Filler Words** — include "um" and "uh" if you need a word-for-word transcript.
* **Number of Speakers** — 1 to 10, to help Evidano tell speakers apart.
* **Start/Stop Time** — transcribe only part of the recording.

**AI Analysis** (adds insights alongside the transcript)

* **Speaker Labels** (Speaker A/B/C), with optional **Speaker Identification** by **Name** or **Role** (for example, Interviewer and Participant) — detected automatically or from names you provide.
* **Auto Chapters**, **Auto Highlights**, **Entity Detection**, **Sentiment Analysis**, **Topic Detection** (hundreds of predefined topics), and **Multichannel Audio**.

**Redact Sensitive Information**

* **Redact Personally Identifiable Information** — choose from a wide range of personal-information types, such as names, emails, phone numbers, and financial or medical details.

<figure><img src="/files/6wEXKdz77FbDorJhJFzF" alt="The Optional Settings panel with accuracy, AI analysis, and redaction options"><figcaption></figcaption></figure>

{% hint style="info" %}
**Tip for interviews:** turn on **Speaker Labels** and set **Speaker Identification** to **Role** (Interviewer / Participant). It makes the transcript — and your later analysis — much easier to read.
{% endhint %}

## Step 3: Submit your files

Click **"Process All Transcriptions"**. Your first 60 minutes of audio are free; after that, the app charges "USD 1 per 1 hour of audio" (adding a translation costs a little more, based on duration). Evidano shows you the price and asks you to confirm before charging anything. You are then taken to your projects list, and the app tells you: "You will receive an email once completed."

{% hint style="info" %}
**Costs credits:** Transcription is a pay-as-you-go add-on drawn from your wallet. See [Plans, Limits & Credits](/ailyze-docs/reference/plans-limits-and-credits.md).
{% endhint %}

## Step 4: Track progress

On the **Audio/Video Transcription** list, each project shows a live status that moves through stages like **Validating → In Progress → "Initial transcript ready, Enhancing with AI" → "Transcript ready with AI review."** Failed files show **"You will not be charged."**

<figure><img src="/files/2YYwPoD5Vuv73RO5W5XD" alt="The Audio/Video Transcription list showing live status for each project"><figcaption></figcaption></figure>

## Step 5: Review and edit the transcript

Open a project (single-file projects jump straight to the editor). The transcript appears as speaker-grouped paragraphs, each with a timestamp, and every word is clickable to jump to that moment in the audio or video. Video files show the player next to the transcript.

The editor toolbar gives you:

* **Search** — jump between matches (search only; there is no find-and-replace).
* **Rename speaker** — click a speaker name, choose **Rename Speaker**, then **Replace current** or **Replace All**.
* **Manual edit** — fix wording paragraph by paragraph.
* **Edit with AI** — AI-suggested corrections are applied automatically, each with a before/after view and an **Undo**.
* **Clip** — export a short media clip of a paragraph.
* Insight drawers (if you enabled them): **Chapters, Transcription Confidence, Entities, Sentiment, Key Phrases, Categories.**

<figure><img src="/files/hummJvlyCG7owDPD8zKm" alt="Editing a transcript: renaming speakers and reviewing AI corrections"><figcaption></figcaption></figure>

{% hint style="info" %}
Editing unlocks once the transcript reaches **"Transcript ready with AI review."**
{% endhint %}

## Step 6: Download your transcripts

Click **"Download"** (in the project view or the editor) and choose a format: **DOCX** (Word) or **XLSX** (Excel). Multiple files download together as a ZIP. If you added translations, **"Download Translation"** gives you those too. PDF, SRT, and TXT exports are not offered today.

## Step 7: Send transcripts to analysis

From the project view or editor:

* **"Analyze transcripts"** — converts your transcripts to DOCX and opens document analysis with the files already loaded.
* **"Translate transcripts"** — opens [Translate Files](/ailyze-docs/prepare-research-data/translate-files.md).

You can also **"Share"** a transcript with others through a public link (view and edit).

### A realistic example

> **Scenario:** You recorded 12 user interviews on Zoom. Import all 12 cloud recordings, set the language to English, turn on **Speaker Identification** by **Role**, and click **"Process All Transcriptions"**. Review each transcript in the editor (renaming "Speaker A" to "Interviewer" applies across the file), then click **"Analyze transcripts"** to find the [themes](/ailyze-docs/reference/glossary.md#theme) across all 12 at once.

## Next steps

* [Analyze Interview Transcripts](/ailyze-docs/common-research-workflows/analyze-interview-transcripts.md)
* [How to Read Your Results](/ailyze-docs/analyze-data/how-to-read-your-results.md)
* [Translate Files](/ailyze-docs/prepare-research-data/translate-files.md)
* [Transcription Issues](/ailyze-docs/troubleshooting/transcription-issues.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://ailyze.gitbook.io/ailyze-docs/prepare-research-data/transcribe-audio-and-video.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
