Generate Text to Speech Audio (TTS)

Last updated: September 2026

The text to speech tool turns the written text on your tour stops into spoken narration and attaches the audio file to each stop. You choose a language, a tour, which stops you want audio for, and a narrator, then it does the rest while you do something else.

What you'll need

  • Tour stops with audio transcripts. The text to speech generator needs text to generate from, and it looks for this text in the Audio Transcript field (found in the Audio tab). Stops without transcripts appear greyed out, and include an Edit link so you can add the text.

1. Access the tool

There are three ways to access the TTS tool:

  • Click on Tools in the top navigation, then click on Text to Speech.

  • Click on any tour stop, then click on the Audio tab, and then click the link to the built-in TTS tool.

  • Visit /cms/tts.

2. Select a site

The site is the language. Pick the language you're generating audio for. Once you've translated your text, you can then use the Text to Speech tool to transform your translated text into translated audio, using the same steps below.

The voices can speak all of our supported languages, so your tours sound consistent across the whole app.

3. Select a tour

Choose a tour from the dropdown. Its stops appear in a list, with the ones that
have usable text already selected.

You can:

  • Tick and untick individual stops

  • Use the checkbox in the list header to select or clear all of them

  • Click any stop to read the transcript that will be read aloud. Click Hide Transcript to
    close it again.

Below the list you'll see a summary: how many stops will be generated, and
warnings about anything that needs your attention.

4. Choose a voice

The Voice Settings panel shows the narrator for this tour. If the tour has been generated before, it remembers the voice you used last time. Different tours can have different narrators.

If no voice has been set yet, the panel opens with the voice list so you can choose one. You can search by voice name, accent, gender, or attributes like 'narrative' or 'midwest'. Click the play button next to any voice to hear a sample. Click on a voice to select it, then click Done.

Speed adjusts how quickly the narrator reads. Lower is slower and clearer;
0.95 is a good default for museum narration.

5. Generate

In the audio file output section, you’ll see how many stops will be generated, how many credits it will cost, how many credits your plan has remaining this month, and how much money (roughly) this saves you compared to hiring a voice actor.

Click the blue Generate Text to Speech button.

Each stop in the list shows its progress as it goes: Queued, then Generating, then Saved. The panel on the

It can take 2 to 5 minutes per stop to generate the audio.

When it's done, the form unlocks by itself and the list refreshes. The stops you just generated now show a green Has Audio label, and the audio file for each one is saved directly to the stop.

The files are saved into Media, in the audio folder. You can move them into tour-specific folders later, if you prefer.