Particle.news
Download on the App Store

Google Releases Gemini 3.5 Transcribe and Adds Agentic Features to Gemini Live

Voice can now produce polished, editable transcripts that connect to automated, multi‑step tasks across Google apps and developer previews are rolling out.

Overview

  • Google announced Gemini 3.5 Transcribe and Gemini Live productivity upgrades in late August, making the new transcription model available to developers in public preview and rolling Live features into trusted‑tester and subscription betas.
  • Gemini 3.5 Transcribe converts raw speech into cleaned, formatted text by removing filler words, handling self‑corrections, applying custom vocabulary, and attributing up to three speakers in recordings.
  • The model supports more than 85 languages, offers real‑time streaming and batch APIs, claims lower word‑error rates and faster time‑to‑final‑transcript versus Google’s prior Chirp 3 engine, and can call other Gemini models to delegate tasks like file analysis or image generation.
  • Gemini Live gained agentic tools — Spark for long‑running, multi‑step workflows, a spoken Daily Brief, hands‑free inbox triage, and Personal Intelligence that remembers context across apps — with phased rollouts limited initially to trusted testers and higher‑tier subscribers.
  • Google says privacy controls restrict training on personal Gmail content and limit access to requested tasks, but the edits Transcribe makes to verbatim speech and the phased availability of features create trade‑offs for users and developers to consider.