Gemini for macOS gets voice-powered transcription, text editing, summarization and image generation


Google is rolling out new natural language voice capabilities for the Gemini app on macOS, allowing users to transcribe speech, rewrite text, summarize content, and generate or edit images without leaving the app they are working in.

The update lets users interact with Gemini from any window on their desktop. By long-pressing the Fn key, users can speak naturally into any window, and Gemini processes the request directly within that app.

Intelligent voice dictation

By default, the feature enables intelligent dictation. Gemini converts spoken words into clean, formatted text by automatically removing filler words such as “um” and “ah,” handling mid-sentence corrections, and inserting the polished text directly at the cursor.

Context-aware AI assistance

Users can also enable Gemini reasoning in the app’s settings. Once enabled, Gemini can understand the context of content on the screen to perform more advanced tasks.

These capabilities include:

  • Summarizing files and documents: Highlight local files, images, or documents and ask Gemini to extract key information or generate summaries.
  • Writing and rewriting text: Highlight text anywhere on the screen and use voice to rewrite, shorten, expand, or change its tone before inserting the updated version back into the document.
  • Generating and editing images: Create new images using voice or edit existing images by referencing visuals already open on the desktop.
Availability

The new natural language voice capabilities are rolling out globally to all users of the Gemini app for macOS in English, with support for additional languages coming later.