音声エージェントや字幕、録音分析向けに、ストリーミングと録音ファイルの二経路を分けて提供する。Smart transcription(自己訂正の整理・フィラー除去・整形)、カスタム語彙、85 以上の言語自動検出に対応。録音側は話者帰属と単語タイムスタンプも使える(3 人超の話者は実験的)。
- リアルタイム: Live API の
gemini-3.5-transcribe-live - 録音: Interactions API の
gemini-3.5-transcribe
開発者向けは Google AI Studio / Antigravity で public preview。企業向けは Gemini Enterprise Agent Platform でも preview。
性能
| モデル | AA-WER(非ストリーミング) | Median Speed Factor | 価格($/1000 分) |
|---|---|---|---|
| Scribe v2(ElevenLabs) | 2.2% | 53.9 | 3.67 |
| MAI-Transcribe-1.5(Azure) | 2.4% | 194.4 | 6.00 |
| Gemini 3.5 Transcribe | 2.6% | 89.0 | 5.00 |
| GPT Transcribe(OpenAI) | 3.3% | 40.8 | 4.50 |
Artificial Analysis の非ストリーミング指標では上位帯にあり、速度は同帯の一部モデルより遅い一方、価格は中程度。
API 価格の比較
| モデル | 音声入力($/1M tokens) | テキスト出力($/1M tokens) | 目安($/分) |
|---|---|---|---|
gemini-3.5-transcribe | 2.00 | 12.00 | ~0.005 |
gemini-3.5-transcribe-live | 3.50 | 21.00 | ~0.009 |
gpt-4o-transcribe | 2.50 | 10.00 | 0.006 |
gpt-transcribe | — | — | 0.0045 |
録音向けのブレンド単価は OpenAI の gpt-transcribe に近く、Live 側はそれより高い。