Gemini 3.5 Transcribe Arrives Native 85-Language Audio Processing Comes to API and Chrome.
While Google has yet to release its flagship Gemini 3.5 Pro model, the tech giant continues to expand its sub-model ecosystem with the official launch of Gemini 3.5 Transcribe, a specialized model designed for ultra-accurate speech-to-text conversion.
Gemini 3.5 Transcribe serves as the underlying engine powering Rambler, an Android feature first teased at Google I/O 2026 and recently debuted on the Pixel 11 smartphone series. Rambler specializes in listening to disjointed, stuttered, or convoluted spoken speech typical of natural human conversation and automatically restructuring it into clean, coherent text.
Google highlights impressive technical benchmarks and feature sets for the new model:
Low Error Rates: Achieves an industry-leading Word Error Rate (WER) of just 4.0% for live audio streaming and 2.6% for pre-recorded audio files.
Multilingual & Speaker Identification: Supports 85 languages with automatic language detection, alongside dynamic speaker diarization capable of identifying up to 3 distinct speakers.
Customization: Features custom vocabulary mapping to ensure accurate transcription of specialized industry terms, technical jargon, and proper names.
Gemini 3.5 Transcribe is available immediately within the Gemini app (currently exclusive to macOS) and via official developer APIs. Google also confirmed that the model will soon integrate natively into Google Chrome to enable hands-free voice-to-text form filling.
The difference between Gemini 3.5 Transcribe and traditional speech-to-text (STT) transcribers lies in their ability to deliver literal, word-for-word translations, preserving pauses, "um," and sentence beginnings. By directly integrating speech recognition with Gemini's contextual LLM capabilities, Rambler dynamically analyzes human intent, removing unnecessary speech while maintaining the original meaning of spoken thought.
The initial launch of the Gemini desktop app on macOS reflects a strategic effort to engage creative professionals, podcasters, and developers who rely on desktop audio editing and note-taking workflows. Extending the transcription API to desktop operating systems makes Google's AI a core productivity tool in a competitive hardware ecosystem.
The direct integration of Gemini 3.5 Transcribe into Google Chrome signals a broader shift towards voice-centric web navigation. By allowing users to fill out complex forms, compose emails, and interact with web applications using consistently natural speech, Google is lowering accessibility barriers and paving the way for screen-less, immersive browsing.
Source: Google

Comments
Post a Comment