What Is Gemini 3.1 Flash TTS? The Impact of 70 Languages and 200 Voice Tags
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
The era of "AI speaking with human-like emotion" has finally arrived.
Gemini 3.1 Flash TTS, announced by Google on April 15, 2026, is a speech synthesis AI that speaks fluently in over 70 languages and can freely express everything from anger and joy to whispers.
And the pricing is a fraction of a cent per character — remarkably affordable.
In this article, we'll explain in plain terms what this much-talked-about new model can do and where it can be used in Japan.
TTS (Text-to-Speech) is an AI technology that converts written text into spoken audio.
Think of it as an "all-purpose narrator" that can instantly read aloud any script you hand it.
Various TTS solutions have existed before, but many struggled with being "monotonous and robotic" or "lacking emotional expression."
Gemini 3.1 Flash TTS was introduced as a next-generation model that breaks through these limitations. Announced on April 15, 2026, via Google's official blog, it is available through four channels: the Gemini API for developers, Google AI Studio for hands-on testing, Vertex AI for enterprises, and the video creation tool Google Vids.
Its biggest appeal is a mechanism called "voice tags," which lets you direct the voice's expression in fine detail — just like a stage director. Simply insert tags such as [happy] or [whisper] into your text, and the AI changes its tone of voice accordingly.
The defining feature of Gemini 3.1 Flash TTS is the ability to control emotion and tone using more than 200 "voice tags."
Imagine this scene:
You're directing a voice actor on how to deliver each line.
"Sound sad here." "Deliver this part with power." "Say this section quickly." — All of those nuanced instructions can be realized simply by embedding tags into the text.
Here are some specific examples of voice tags:
[happy], [sad], [angry], [surprised][fast], [slow][whisper], [shouting], [calm]For example, if you input "Today was [excited] an absolutely amazing day! [whisper] But I have a little secret…", the AI will generate audio that sounds energetic in the first half and hushed and conspiratorial in the second. This dramatically expands the expressive range of audiobooks, podcasts, and video narration.
Gemini 3.1 Flash TTS supports over 70 languages.
Of those, 24 languages are designated as "high-quality evaluated languages" with especially high accuracy — and Japanese is among these 24.
Hindi, Arabic, German, French, Spanish, and Portuguese are also included at the top quality level.
Another standout feature is multi-speaker dialogue. By specifying multiple character settings within a single piece of text, you can generate audio of different-voiced characters conversing with each other in one go.
For example, the following kinds of scenes can be created with a single instruction:
Moreover, you can configure a "voice profile" (timbre, tone, accent) for each character and export it for reuse in other projects — meaning you can maintain consistency, such as "the MC voice for this podcast always sounds the same."
Let's take a look at pricing. The API pricing for Gemini 3.1 Flash TTS is $0.50 per million input tokens (approximately ¥75) and $10 per million output tokens (approximately ¥1,500).
A "token" is the unit by which an AI processes text. For Japanese, a rough guideline is 1–2 tokens per character.
Here's an estimate for converting a 10-minute podcast (approximately 2,000 characters) to audio:
To put that in perspective, the popular AI voice service ElevenLabs charges $5 to several hundred dollars per month on paid plans, with per-character pricing sometimes several times to over ten times higher. Gemini 3.1 Flash TTS achieves what creators dream of: high quality at low cost.
In addition, Google AI Studio offers a free tier, so individual developers who want to "just try it out first" can get started with ease.
It has also earned high marks for performance.
On the "Artificial Analysis TTS Leaderboard," which compares the quality of AI voices, Gemini 3.1 Flash TTS achieved an Elo score of 1,211.
That places it 2nd in the industry — a remarkable achievement.
ElevenLabs, a company dedicated solely to voice AI, holds the top spot, but Gemini 3.1 Flash TTS outperformed OpenAI's TTS and all other major models. Moreover, for combining high quality with low cost, it is positioned in the "most attractive quadrant" of that leaderboard.
Another important feature is SynthID — a digital watermarking technology.
All audio generated by Gemini 3.1 Flash TTS has an inaudible "watermark" embedded in it, enabling automatic detection of "this is AI-generated audio" after the fact.
As fake audio scams and misuse become growing social issues, this careful attention to safety is a meaningful inclusion.
Let's organize the three major voice AI models by use case.
Strengths: 70+ language support, fine-grained emotion control with 200+ voice tags, multi-speaker functionality, overwhelmingly low pricing. Free trial available via Google AI Studio.
Recommended for: Multilingual content, educational materials, podcasts, video narration, cost-conscious projects.
Strengths: Industry-leading audio quality, ability to create a "voice clone" from as little as one minute of audio, natural accents across 30+ languages.
Recommended for: Professional audiobook production, recreating the voices of public figures or characters, commercial projects demanding the highest quality.
Strengths: Deep integration with the ChatGPT and GPT-5.4 ecosystem, one-stop workflow from text generation to audio.
Recommended for: Embedding into apps already using the OpenAI API, voice chatbots combined with ChatGPT.
In short, a simple guideline is: Gemini for cost efficiency and multilingual support, ElevenLabs for top quality and voice cloning, and GPT-based models for integration with the OpenAI ecosystem.
Gemini 3.1 Flash TTS represents a major opportunity for Japanese creators and businesses. That's because Japanese has been selected as one of 24 "high-quality evaluated languages."
Many previous voice AIs prioritized English, and Japanese was often only at a "usable" level.
Gemini 3.1 Flash TTS, however, includes Japanese in its highest-quality group and has been fine-tuned accordingly.
This means it is more likely to generate audio that sounds natural to Japanese native speakers.
Here are some concrete use cases expected in the Japanese market:
Japanese Google Workspace users can experience some features through the video creation tool "Google Vids." For serious use, Vertex AI via Google Cloud is the standard route for enterprise adoption.
A. The easiest option is Google AI Studio (ai.google.dev).
If you have a Google account, you can start right away within the free tier.
For serious development, the Gemini API is available, and for enterprise use, Vertex AI is provided.
If you have a Google Workspace subscription, you can also access some features through the video creation tool "Google Vids."
A. It is very natural.
Japanese is included among the "high-quality evaluated languages" (24 languages) out of 70 total, and Google has invested significant effort in fine-tuning it.
Reports indicate that the "accent mismatches" and "mechanical pauses" that were noticeable in older TTS systems have been greatly improved.
That said, the pronunciation of technical terms and proper nouns is not perfect, so we recommend checking important content in advance.
A. Commercial use is permitted.
Pricing is $0.50 per million input tokens and $10 per million output tokens.
For example, a 10-minute podcast (approximately 2,000 characters) can be converted to audio for just a few tens of yen.
Using Vertex AI via Google Cloud also gives you access to enterprise-grade support and data management.
A. ElevenLabs leads the industry in audio quality (ranked 1st on the Artificial Analysis TTS Leaderboard), while Gemini 3.1 Flash TTS ranks 2nd — but at a fraction of the price.
In terms of features, Gemini actually has an edge in some areas, such as support for 70+ languages and 200+ voice tags.
Voice cloning (training the AI on your own voice) is a strength of ElevenLabs, but for emotionally expressive multilingual narration, Gemini offers the best cost-to-performance ratio.
A. Yes, it can be detected.
All audio generated by Gemini 3.1 Flash TTS has a SynthID digital watermark embedded in it.
This is inaudible to human ears, but with the appropriate detection tool, it can be automatically identified as "AI-generated audio."
It is one of Google's safety measures in response to growing concerns about AI-based fraud and deepfakes.
Voice AI is at a turning point: Google has stormed into a market once dominated by dedicated providers, bringing overwhelming multilingual support and cost performance. Since Google AI Studio offers a free tier to get started, podcasters, video creators, and app developers should try it out firsthand to experience the naturalness of Japanese audio for themselves.
This article is a cross-post from AI Friends.