AI dubbing from model to product

I led the 0→1 design of AI dubbing for rednote, connecting the creator experience with viewer playback. Alongside the product work, I joined ML quality reviews. In a subsequent internal evaluation, the share of text-to-speech (TTS) outputs rated excellent increased by 13.16 pp.

RoleLead product designer (sole designer)
Duration8 months
CollaborationPM · ML · Engineering · Legal & Compliance
Companyrednote (Xiaohongshu)
StatusShipped · Experiment running

Identifying the opportunity

AI dubbing was part of making rednote, a community where most content is in Chinese, accessible to overseas users. Text, images, video subtitles, and live captions already helped people cross language barriers. Voice was the next layer of multimodal translation.

Comment translation in rednote
Text translation
Image translation in rednote
Image translation
Translated subtitles on a video
Video captions
Translated captions and comments in a livestream
Live captions
Blurred spoken-video screen with a centered white audio waveform
What about voice?

In January 2026, our ML team introduced a voice model that could power AI dubbing on rednote. My PM and I still needed to understand how it performed on real videos, whether creators would welcome the feature, and what was feasible for the first release.

How could we build a first release that worked for creators and viewers?

Scoping the first launch

Working with ML, we tested the model on videos colleagues had already posted on rednote. It worked better with clear, normal-paced Mandarin narration, but results varied across videos. We also partnered with UXR to understand creators’ expectations and concerns. Creators were curious about AI dubbing but wanted to hear how their voices would sound.

These findings led us to start with a small group of creators whose videos worked well with the model.

Creator message list with a new AI dubbing invitation
System message inviting a creator to try AI dubbing
AI dubbing introduction and authorization page
Private preview of the creator's AI-dubbed video
AI dubbing management page with the global setting and dubbed videos

Pushing back on another entry point

During a design review, my PM suggested adding another entry point to encourage more creators to enable AI dubbing. Creators often revisit their own videos, so the proposal was to place a creator authorization switch next to the viewer dubbing setting in the long-press panel. If creators discovered the viewer setting while watching dubbed videos, they would know where to look on their own videos.

How I pushed back

The panel treated everyone as a viewer

Even when creators watched their own videos, the long-press panel still treated them as viewers. Adding a creator authorization switch would mix a publishing decision into that viewer context. I mocked up both in the same panel to make the confusion visible.

We agreed to keep creator authorization out of the playback panel.

Product proposal with the viewer dubbing setting and creator authorization switch in one panel
Product proposal
Final long-press panel with only the viewer dubbing setting
Final decision

That decision also made playback simpler. Viewers saw one dubbing setting in the long-press panel, and their language choice carried across eligible videos instead of resetting on every video.

Viewer playback flow

Eligible video playing with its original audio
Original audio
Playback settings with AI dubbing available and turned off
Playback settings
AI dubbing language menu with English available
AI dubbing settings
Video playing with a disclosure that the English dub is on
English dub on
A different eligible video inheriting the English dubbing preference
Preference carries over

Viewer playback flow

Eligible video playing with its original audio
Original audio
Playback settings with AI dubbing available and turned off
Playback settings
AI dubbing language menu with English available
AI dubbing settings
Video playing with a disclosure that the English dub is on
English dub on
A different eligible video inheriting the English dubbing preference
Preference carries over

Reframing with data

~11%invitation-to-page conversion
58%opted in after reaching the page

More than half of creators who reached the authorization page opted in, but only around 11% of invited creators reached it. Separately, some creators still asked what AI dubbing would sound like after seeing the introduction page. We needed to improve both how creators discovered the feature and how the page explained it.

Helping creators discover, understand and control

I stepped through the product from a creator’s point of view and looked at where they already published and managed videos. That led me to propose new entry points and bring discovery, authorization, and control into the same flow.

Creator profile with the AI dubbing education sheetPublishing settings with an AI dubbing authorization entryCreator Center with a new AI dubbing entry

Surface AI dubbing in creator workflows

I proposed entry points on creator profiles, in publishing, and inside Creator Center. All three led to the same authorization flow.

Revised feature page with an Original and AI-dubbed comparison

Explain AI dubbing with a shared sample

To make AI dubbing easier to understand, I replaced the private preview with a shared Original and AI-dubbed comparison. A shared sample offered a more consistent way to demonstrate the difference.

Publishing settings with AI dubbing enabled for the current videoScope menu for turning off one video or all videos

Balance simple authorization with per-video control

One authorization covered eligible videos, while per-video controls let creators disable dubbing before publishing or edit the post later if the result did not work.

Raising the quality bar

Broader discovery would only help if the voice experience was ready. As the team prepared for a wider rollout, refining model quality became a major part of the work. The ML team shared an English dub that preserved the speaker’s Sichuan accent. Product and I felt the accent was exaggerated and made the output sound unnatural.

I proposed adding an experience review for voice quality and worked with my PM, language experts, and the ML team to turn that judgment into a repeatable process.

  1. Step 01

    Make the gap observable

    We paired each source video with the latest model output, then traced each issue to ASR, translation, TTS, or audio mixing.

  2. Step 02

    Turn reactions into criteria

    We turned reactions like “this sounds unnatural” into dimensions the team could tag and score. Within TTS, I helped define accent naturalness alongside ease of understanding and voice similarity.

  3. Step 03

    Review across disciplines

    My PM led the process while language experts, the ML team, and I listened together, tagged issues, and aligned on the scoring standard.

  4. Step 04

    Re-evaluate each iteration

    I joined multiple rounds of scoring as the ML team iterated, using the same criteria to judge whether each version improved the listening experience.

In a subsequent internal human evaluation, the share of TTS outputs rated excellent increased by 13.16 percentage points from the prior round. The review also gave the team a consistent way to find quality issues and assess each new version across the dubbing pipeline.

Outcomes and learnings

The first AI dubbing experience launched for creators and viewers. By the time I left, the revised product experience and an updated model were running together in a new experiment.

AI opened a new way for videos to travel across languages. It also introduced a new trust challenge for the platform. Creators needed confidence in how their voice would be used by AI, while viewers needed confidence in the AI-generated content the platform delivered. I saw an opportunity for experience design to help people understand what AI was doing, set expectations for its output, and stay in control of how it was used.

Designing AI experiences means shaping quality and earning trust.

Lauren