Qwen logo

Qwen composer UX: multimodal + menu, Auto/Thinking & voice

Updated August 25, 2026

Qwen Studio keeps the default bar calm. The product mostly lives in +: uploads, image and video, search, research, sites, slides. Chat models sit in the header; image models show up after you pick Create Image. Auto, Thinking, and Fast are effort, not model names. The waveform button starts Voice or Video Chat. The mic is dictation. Two voice paths, same bar, no labels.

Calm default

Open chat.qwen.ai on a new chat with the + menu closed.
Open chat.qwen.ai on a new chat with the + menu closed.

What works

  • The bar is empty except for Ask Qwen, +, Auto, mic, and the waveform button. First visit is not a mode catalog.
  • Qwen3.7-Plus in the header names the default brain without opening a picker.
  • Projects and All chats in the sidebar give structure without crowding the composer.

What we would push on

  • How can I help you? is generic chat copy. Nothing says Qwen does images, video, or slides until you dig into +.
  • Mic and waveform sit side by side with no labels. New users will not know one is dictation and one is a live voice session.

Product bet

Alibaba is betting the default loop stays familiar chat. Multimodal and agent jobs stay nested so casual Q&A does not look like a studio on day one.

Tradeoff

DecisionBenefitCost
Calm bar with multimodal modes only in +Low cognitive load on first sendImage, video, and dev work are invisible until +

Takeaway

If + holds your real product surface, keep the default bar typeable and uncluttered. Label voice dictation vs voice chat before users tap the wrong control.

Qwen3.7-Plus vs Max

Click Qwen3.7-Plus in the header to open the model menu.
Click Qwen3.7-Plus in the header to open the model menu.

What works

  • Qwen3.7-Plus and Qwen3.8-Max each get a sentence about what they are for, not just a version number.
  • Checkmark on 3.7-Plus makes the current choice obvious.
  • Expand more models at the bottom promises a longer list without dumping ten rows on first open.
  • Model Comparison toggle in the menu is an A/B path for people who care about benchmarks, not a permanent header chip.

What we would push on

  • Plus vs Max is still raw naming. Casual users need a plain word for why Max costs more.
  • No latency, context window, or price hint on either row.

Product bet

Qwen ships version numbers as the brand, same as GLM. The picker is the launch surface for 3.8-Max while 3.7-Plus stays the safe default.

Tradeoff

DecisionBenefitCost
Header model menu with comparison toggle and collapsed long tailFlagship visible; power users can comparePlus/Max jargon; no cost or speed preview

Takeaway

Keep version IDs if that is your brand, but add one plain line per row about speed, cost, or what unlocks on Max.

Pattern: Model Selection UIRows mix version IDs with one-line job copy and a Model Comparison toggle at the top of the menu.

Pattern: Cost Transparency

+ menu: the real product

Click + in the composer to open the attach and modes menu.
Click + in the composer to open the attach and modes menu.

What works

  • Upload attachment lists file, image, video, audio upfront. Users know multimodal is in scope.
  • Create Image and Create Video sit next to Web search and Deep Research. Generation and research are peers, not buried settings.
  • Web Dev and Slides name deliverables, not model features.
  • Tools row with a toggle at the bottom looks like a master switch for tool calling without leaving +.

What we would push on

  • Eight rows plus More and Tools is a lot for one flyout. Kimi puts artifact jobs on home pills; Qwen hides them here.
  • No capture for Deep Research, Web Dev, or Slides in this walkthrough. Users cannot tell if those modes rewrite the bar or just add a chip.

Product bet

Qwen is selling a multimodal studio inside one chat shell. The + menu is the SKU list: media, search, research, sites, decks.

Tradeoff

DecisionBenefitCost
Single + flyout for attach, generative modes, search, and dev jobsOne attach surface; calm default barHeavy first open; no home-row advertising like Kimi

Takeaway

When + is your mode catalog, spell formats on Upload and group generative vs research vs dev so the list scans fast.

Inline attachment preview

Upload an image via +; keep the composer open with the thumbnail visible.
Upload an image via +; keep the composer open with the thumbnail visible.

What works

  • The thumbnail sits inside the bar with an X to remove it. You see what you attached before send.
  • Ask Qwen placeholder stays. The bar still reads as typeable chat, not a upload-only form.
  • Send upgrades to a black up-arrow once input is ready, which is clearer than the idle waveform state.

What we would push on

  • No file name or type label on the thumbnail. A portrait and a PDF look the same at a glance.
  • Auto effort dropdown stays visible. Unclear if attachment changes which model or mode runs.

Tradeoff

DecisionBenefitCost
In-bar thumbnail with dismiss, generic placeholder unchangedVisual confirm before send; bar stays familiarNo metadata on the chip; effort picker unexplained with files

Takeaway

Show attachments inside the composer with a remove control. Add file type or filename when the thumbnail alone is ambiguous.

Auto, Thinking, Fast

Click Auto on the right side of the composer.
Click Auto on the right side of the composer.

What works

  • Three choices with plain names. Auto, Thinking, Fast are effort labels, not Qwen3.x version IDs.
  • Checkmark on Auto shows the default without extra copy.
  • Nested under the bar, same slot as Kimi thinking effort and Claude effort controls.

What we would push on

  • No subtext on what Thinking costs in time or tokens. Fast vs Auto is a guess.
  • Header still says Qwen3.7-Plus while effort is separate. Two levers, one sentence of guidance would help.

Product bet

Auto keeps the free-feeling loop cheap. Thinking is the upsell for hard prompts without forcing everyone through a reasoning model on every send.

Tradeoff

DecisionBenefitCost
Effort menu separate from header model pickerReasoning depth without retooling the whole barNo outcome copy or latency preview on Thinking

Takeaway

Put thinking effort in the composer, not the model menu. One line per option beats Auto and Thinking alone.

Voice Chat vs Video Chat

Click the black waveform button on the right of the composer.
Click the black waveform button on the right of the composer.

What works

  • Two labeled rows: Voice Chat with a headset icon, Video Chat with a camera icon. The choice is explicit.
  • Lives on the waveform button, not the mic. Dictation and live session are separate entry points.

What we would push on

  • Neither row names the omni model or a time limit until you are already in session.
  • Side-by-side mic and waveform with no helper text will still cause mis-taps.

Product bet

Qwen3.5 Omni and Qwen2.5-Omni are products, not mic settings. A dedicated launcher keeps realtime audio/video out of the text composer state machine.

Tradeoff

DecisionBenefitCost
Waveform button opens Voice vs Video, mic stays for dictationRealtime sessions do not hijack the text barTwo voice icons; limits hidden until connect

Takeaway

Split dictation from voice chat at the control level. Name both paths on the bar or in a tooltip.

Voice Chat session

Start Voice Chat from the waveform menu.
Start Voice Chat from the waveform menu.

What works

  • Full-screen session with a soft orb, I'm listening, and Qwen3.5 Omni at the bottom. Feels like a phone call, not chat.
  • Timer badge shows 00:02 / 10:00. Users see the cap before they settle in.
  • Close, settings, and mic mute sit in predictable corners.

What we would push on

  • Settings opens voice picker on first visit with no hint that personas exist.
  • Ten minutes is short for a work conversation. The timer is honest but tight.

Tradeoff

DecisionBenefitCost
Immersive voice UI with visible session capClear mode switch; time limit upfrontPersona choice deferred to settings; 10:00 may feel stingy

Takeaway

Voice sessions deserve their own screen and a visible timer. Surface voice choice before connect if personas matter.

Pattern: Voice Input

Voice personas

In Voice Chat, open settings to reach Select voice.
In Voice Chat, open settings to reach Select voice.

What works

  • Search Voice and Filter scale when the list is long.
  • Each persona gets a name and a personality blurb, not only a gender tag.
  • Play on a row plus a detail popover with Supported Language list makes Gold feel like a character pack, not a system voice.
  • Start a new chat as primary action commits the voice to a fresh thread.

What we would push on

  • Copy runs poetic ("hint of coffee and old books"). Fun, but hard to pick a voice for a meeting.
  • Role-playing tag on Gold is accurate. Still an odd default catalog tone for a general assistant.

Product bet

Voice Chat is partly entertainment. Persona depth and multilingual support sell omni models to consumers who want character, not just transcription.

Tradeoff

DecisionBenefitCost
Searchable persona catalog with playful copy and language matrixDifferentiation vs flat TTS; global language list visibleHard to pick quickly; tone skews role-play

Takeaway

If you ship personas, add a neutral default and filter by use case, not only search and vibe copy.

Inline dictation

Click the mic in the composer and speak.
Click the mic in the composer and speak.

What works

  • The bar expands into a card with live transcript ("Hello! Hello! Hello!") and a waveform.
  • X cancels, checkmark confirms. Same confirm pattern as mobile voice keyboards.
  • Dictation stays on the home composer. You never leave How can I help you?

What we would push on

  • Waveform button still visible in idle shots beside mic. The two voice entry points remain easy to confuse.

Tradeoff

DecisionBenefitCost
In-place dictation card with transcript and confirm/cancelEdit before send; no mode pageCompetes with Voice Chat on the same bar

Takeaway

Dictation should expand the composer with transcript and explicit confirm. Keep it off the same button as realtime voice chat.

Pattern: Voice Input

Create Image mode

From +, choose Create Image.
From +, choose Create Image.

What works

  • Create Image chip lands in the bar with an X. Mode is visible before send.
  • Qwen-Image 2.0 and 16:9 dropdowns sit beside the chip. Model and format are set pre-prompt.
  • Gallery below shows finished work (crochet doll, Christmas scene, pencil sketch, frogs). The job is obvious.
  • Explore more links out without cluttering the bar.

What we would push on

  • Placeholder stays Ask Qwen. Kimi and GLM rewrite copy to the job when a mode is armed.
  • Header model is still Qwen3.7-Plus while image model is Qwen-Image 2.0. Two pickers, no bridge copy.

Product bet

Image gen is a first-class mode, not a plugin. In-bar chip plus gallery competes with Midjourney-style browse while keeping chat as the shell.

Tradeoff

DecisionBenefitCost
Create Image chip with image-model and aspect dropdowns plus example galleryScope set before generate; browse plus typeDual model pickers; generic placeholder

Takeaway

When image mode is on, show the chip, model, aspect, and examples together. Change placeholder copy to a create job.

Use Prompt on gallery cards

In Create Image mode, hover a gallery card and click Use Prompt.
In Create Image mode, hover a gallery card and click Use Prompt.

What works

  • Card reveals the actual prompt ("A close-up, professionally composed photograph of a hand-crocheted yarn doll…"). Users learn what good input looks like.
  • Use Prompt is a one-click fill, not copy-paste from alt text.
  • Blur overlay keeps the gallery pretty while still teaching structure.

What we would push on

  • Hover-only on desktop does not help touch users.
  • Only one card in the row shows the overlay in this capture. Unclear if others are clickable the same way.

Tradeoff

DecisionBenefitCost
Hover card with full prompt text and Use Prompt CTATeaches prompt shape; lowers blank-page anxietyHover dependent; affordance easy to miss

Takeaway

Show the prompt on example cards and let users inject it. Always visible beats hover-only.

Pattern: Prompt StartersFeatured images expose the full generation prompt on hover with a Use Prompt button, not only a silent thumbnail.

Qwen-Image 2.0 vs 3.0

In Create Image mode, open the Qwen-Image dropdown in the bar.
In Create Image mode, open the Qwen-Image dropdown in the bar.

What works

  • Image models are separate from chat models. Qwen-Image 3.0 and 2.0 live in the mode bar, not the header.
  • Checkmark on 2.0 shows the active image engine without leaving Create Image.

What we would push on

  • Rows have no quality, speed, or style copy. 3.0 vs 2.0 is a version guess.
  • Chat header still reads Qwen3.7-Plus. Three model names on one screen.

Tradeoff

DecisionBenefitCost
Nested image-model picker inside Create Image chip rowChat and image engines stay decoupledNo job copy; model name overload

Takeaway

Split chat and image model pickers, but say what 3.0 improves in one line inside the menu.

Aspect ratio in the bar

In Create Image mode, open the aspect ratio dropdown.
In Create Image mode, open the aspect ratio dropdown.

What works

  • Five ratios: 1:1, 3:4, 4:3, 16:9, 9:16. Covers feed, portrait, and slide shapes.
  • Checkmark on 16:9 matches the selected label on the closed dropdown.
  • Format is chosen before the prompt, same pattern as Kimi Design Adaptive control.

What we would push on

  • Text labels only. Kimi uses rectangle icons so 9:16 vs 16:9 is scannable.

Takeaway

Put aspect ratio in the mode bar for image gen. Icons help when the list is mostly numbers.

Web search mode

From +, choose Web search.
From +, choose Web search.

What works

  • Web search chip in the bar with an X. Search is armed visibly, not a hidden globe default.
  • Suggested queries under the bar (Best budget laptops, Learn Spanish online…) give one-tap starters.
  • Fast effort dropdown stays on the right. Search mode does not swap the whole composer chrome.

What we would push on

  • Starters are generic SEO queries. Perplexity ties starters to recency and citations upfront.
  • Deep Research is a separate + row. Relationship between search, thinking effort, and deep research is unclear.

Product bet

Web search is a mode chip, not always-on browse. Starters nudge first search without opening a research product.

Tradeoff

DecisionBenefitCost
Removable Web search chip plus query starters under the barMode obvious; low-friction first queryGeneric starters; overlap with Deep Research row in +

Takeaway

When search is a mode, show a removable chip and starters that match your trust story, not generic head terms.

Copy this

  • Multimodal attach list in + with explicit file, image, video, audio formats
  • Auto / Thinking / Fast effort menu separate from header model versions
  • Waveform button for Voice Chat vs Video Chat, mic for dictation only
  • Voice session with visible timer (00:02 / 10:00) and immersive orb UI
  • Searchable voice persona catalog with language matrix on detail
  • Create Image chip with Qwen-Image model picker and aspect ratio in the bar
  • Gallery Use Prompt that injects the full generation prompt
  • Web search as a removable chip with suggested queries under the bar

Skip this

  • Mic and waveform side by side with no labels
  • Plus vs Max model rows without cost, speed, or plain-language why
  • Eight-row + menu with no home pills for Slides, Web Dev, or Deep Research
  • Poetic voice personas as the only catalog tone
  • Generic Ask Qwen placeholder while Create Image mode is armed
  • Three model names visible (chat header, image engine, effort) with no map

How others design the composer

How other products handle the same job, and what each tradeoff reveals.

Compare composer UX across products

ChatGPT, Claude, Perplexity, and Gemini side by side: default bar, tools, cost, and patterns to copy.

Full comparison

Useful for a critique or spec? Share it.

Original gallery pages: Tool Switching in Composer