Good morning, NOLA.

Design a voice, then direct the scene

Google’s Gemini 3.8 Flash TTS and Flash-Lite TTS turn text-to-speech into a more directed creative tool. Flash can create original voices from natural-language descriptions across more than 100 languages and dialects, then control acting cues, pacing and dialect line by line.

The models can stage two speakers from one script. Voice replication uses a 30-second sample and requires a matching verbal-consent recording from the voice owner.

Google says both models are rolling out through the Gemini API and Google AI Studio, with Flash also rolling out in Gemini Notebook. For a practical starting point, tool builder Simon Willison made a bring-your-own-key playground for narration and multi-speaker conversations. His two-pelican demo took about 20 seconds to generate 1 minute and 18 seconds of audio.

Muse follow-up: glasses, Mac and email

Meta’s Muse is an agent designed to handle everyday tasks by connecting to services such as email and calendars. The latest update brings Muse to Meta smart glasses, where spoken requests could log meals, book appointments or help with purchases. Meta expects that integration in the coming months.

Muse computer work is also coming to Mac, with the agent intended to operate apps and continue queued jobs after the user walks away. Meta says Muse will soon get an email address, allowing people to add it to threads or forward messages for handling.

Talk through work on your phone

OpenAI is bringing voice-triggered workflows to ChatGPT mobile. Plus and Pro subscribers can use the Work tab to draft documents and email or summarize Slack messages. Free and Go users get plugins and connected apps.

Users can switch between voice and text, then start a conversation on mobile and resume it on desktop.

Also

Share this brief →

Pass it on

Share this brief