Qwen composer UX: multimodal + menu, Auto/Thinking & voice
Qwen Studio keeps the default bar calm. The product mostly lives in +: uploads, image and video, search, research, sites, slides. Chat models sit in the header; image models show up after you pick Create Image. Auto, Thinking, and Fast are effort, not model names. The waveform button starts Voice or Video Chat. The mic is dictation. Two voice paths, same bar, no labels.
Calm default

What works
- The bar is empty except for Ask Qwen, +, Auto, mic, and the waveform button. First visit is not a mode catalog.
- Qwen3.7-Plus in the header names the default brain without opening a picker.
- Projects and All chats in the sidebar give structure without crowding the composer.
What we would push on
- How can I help you? is generic chat copy. Nothing says Qwen does images, video, or slides until you dig into +.
- Mic and waveform sit side by side with no labels. New users will not know one is dictation and one is a live voice session.
Product bet
Alibaba is betting the default loop stays familiar chat. Multimodal and agent jobs stay nested so casual Q&A does not look like a studio on day one.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Calm bar with multimodal modes only in + | Low cognitive load on first send | Image, video, and dev work are invisible until + |
Takeaway
If + holds your real product surface, keep the default bar typeable and uncluttered. Label voice dictation vs voice chat before users tap the wrong control.
Pattern: Tool Switching in Composer
Pattern: Model Selection UI
Qwen3.7-Plus vs Max

What works
- Qwen3.7-Plus and Qwen3.8-Max each get a sentence about what they are for, not just a version number.
- Checkmark on 3.7-Plus makes the current choice obvious.
- Expand more models at the bottom promises a longer list without dumping ten rows on first open.
- Model Comparison toggle in the menu is an A/B path for people who care about benchmarks, not a permanent header chip.
What we would push on
- Plus vs Max is still raw naming. Casual users need a plain word for why Max costs more.
- No latency, context window, or price hint on either row.
Product bet
Qwen ships version numbers as the brand, same as GLM. The picker is the launch surface for 3.8-Max while 3.7-Plus stays the safe default.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Header model menu with comparison toggle and collapsed long tail | Flagship visible; power users can compare | Plus/Max jargon; no cost or speed preview |
Takeaway
Keep version IDs if that is your brand, but add one plain line per row about speed, cost, or what unlocks on Max.
Pattern: Model Selection UIRows mix version IDs with one-line job copy and a Model Comparison toggle at the top of the menu.
Pattern: Cost Transparency
Inline attachment preview

What works
- The thumbnail sits inside the bar with an X to remove it. You see what you attached before send.
- Ask Qwen placeholder stays. The bar still reads as typeable chat, not a upload-only form.
- Send upgrades to a black up-arrow once input is ready, which is clearer than the idle waveform state.
What we would push on
- No file name or type label on the thumbnail. A portrait and a PDF look the same at a glance.
- Auto effort dropdown stays visible. Unclear if attachment changes which model or mode runs.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| In-bar thumbnail with dismiss, generic placeholder unchanged | Visual confirm before send; bar stays familiar | No metadata on the chip; effort picker unexplained with files |
Takeaway
Show attachments inside the composer with a remove control. Add file type or filename when the thumbnail alone is ambiguous.
Auto, Thinking, Fast

What works
- Three choices with plain names. Auto, Thinking, Fast are effort labels, not Qwen3.x version IDs.
- Checkmark on Auto shows the default without extra copy.
- Nested under the bar, same slot as Kimi thinking effort and Claude effort controls.
What we would push on
- No subtext on what Thinking costs in time or tokens. Fast vs Auto is a guess.
- Header still says Qwen3.7-Plus while effort is separate. Two levers, one sentence of guidance would help.
Product bet
Auto keeps the free-feeling loop cheap. Thinking is the upsell for hard prompts without forcing everyone through a reasoning model on every send.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Effort menu separate from header model picker | Reasoning depth without retooling the whole bar | No outcome copy or latency preview on Thinking |
Takeaway
Put thinking effort in the composer, not the model menu. One line per option beats Auto and Thinking alone.
Voice Chat session

What works
- Full-screen session with a soft orb, I'm listening, and Qwen3.5 Omni at the bottom. Feels like a phone call, not chat.
- Timer badge shows 00:02 / 10:00. Users see the cap before they settle in.
- Close, settings, and mic mute sit in predictable corners.
What we would push on
- Settings opens voice picker on first visit with no hint that personas exist.
- Ten minutes is short for a work conversation. The timer is honest but tight.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Immersive voice UI with visible session cap | Clear mode switch; time limit upfront | Persona choice deferred to settings; 10:00 may feel stingy |
Takeaway
Voice sessions deserve their own screen and a visible timer. Surface voice choice before connect if personas matter.
Pattern: Voice Input
Voice personas

What works
- Search Voice and Filter scale when the list is long.
- Each persona gets a name and a personality blurb, not only a gender tag.
- Play on a row plus a detail popover with Supported Language list makes Gold feel like a character pack, not a system voice.
- Start a new chat as primary action commits the voice to a fresh thread.
What we would push on
- Copy runs poetic ("hint of coffee and old books"). Fun, but hard to pick a voice for a meeting.
- Role-playing tag on Gold is accurate. Still an odd default catalog tone for a general assistant.
Product bet
Voice Chat is partly entertainment. Persona depth and multilingual support sell omni models to consumers who want character, not just transcription.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Searchable persona catalog with playful copy and language matrix | Differentiation vs flat TTS; global language list visible | Hard to pick quickly; tone skews role-play |
Takeaway
If you ship personas, add a neutral default and filter by use case, not only search and vibe copy.
Pattern: Progressive Disclosure
Inline dictation

What works
- The bar expands into a card with live transcript ("Hello! Hello! Hello!") and a waveform.
- X cancels, checkmark confirms. Same confirm pattern as mobile voice keyboards.
- Dictation stays on the home composer. You never leave How can I help you?
What we would push on
- Waveform button still visible in idle shots beside mic. The two voice entry points remain easy to confuse.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| In-place dictation card with transcript and confirm/cancel | Edit before send; no mode page | Competes with Voice Chat on the same bar |
Takeaway
Dictation should expand the composer with transcript and explicit confirm. Keep it off the same button as realtime voice chat.
Pattern: Voice Input
Create Image mode

What works
- Create Image chip lands in the bar with an X. Mode is visible before send.
- Qwen-Image 2.0 and 16:9 dropdowns sit beside the chip. Model and format are set pre-prompt.
- Gallery below shows finished work (crochet doll, Christmas scene, pencil sketch, frogs). The job is obvious.
- Explore more links out without cluttering the bar.
What we would push on
- Placeholder stays Ask Qwen. Kimi and GLM rewrite copy to the job when a mode is armed.
- Header model is still Qwen3.7-Plus while image model is Qwen-Image 2.0. Two pickers, no bridge copy.
Product bet
Image gen is a first-class mode, not a plugin. In-bar chip plus gallery competes with Midjourney-style browse while keeping chat as the shell.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Create Image chip with image-model and aspect dropdowns plus example gallery | Scope set before generate; browse plus type | Dual model pickers; generic placeholder |
Takeaway
When image mode is on, show the chip, model, aspect, and examples together. Change placeholder copy to a create job.
Use Prompt on gallery cards

What works
- Card reveals the actual prompt ("A close-up, professionally composed photograph of a hand-crocheted yarn doll…"). Users learn what good input looks like.
- Use Prompt is a one-click fill, not copy-paste from alt text.
- Blur overlay keeps the gallery pretty while still teaching structure.
What we would push on
- Hover-only on desktop does not help touch users.
- Only one card in the row shows the overlay in this capture. Unclear if others are clickable the same way.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Hover card with full prompt text and Use Prompt CTA | Teaches prompt shape; lowers blank-page anxiety | Hover dependent; affordance easy to miss |
Takeaway
Show the prompt on example cards and let users inject it. Always visible beats hover-only.
Pattern: Prompt StartersFeatured images expose the full generation prompt on hover with a Use Prompt button, not only a silent thumbnail.
Qwen-Image 2.0 vs 3.0

What works
- Image models are separate from chat models. Qwen-Image 3.0 and 2.0 live in the mode bar, not the header.
- Checkmark on 2.0 shows the active image engine without leaving Create Image.
What we would push on
- Rows have no quality, speed, or style copy. 3.0 vs 2.0 is a version guess.
- Chat header still reads Qwen3.7-Plus. Three model names on one screen.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Nested image-model picker inside Create Image chip row | Chat and image engines stay decoupled | No job copy; model name overload |
Takeaway
Split chat and image model pickers, but say what 3.0 improves in one line inside the menu.
Pattern: Model Selection UI
Pattern: Tool Switching in Composer
Aspect ratio in the bar

What works
- Five ratios: 1:1, 3:4, 4:3, 16:9, 9:16. Covers feed, portrait, and slide shapes.
- Checkmark on 16:9 matches the selected label on the closed dropdown.
- Format is chosen before the prompt, same pattern as Kimi Design Adaptive control.
What we would push on
- Text labels only. Kimi uses rectangle icons so 9:16 vs 16:9 is scannable.
Takeaway
Put aspect ratio in the mode bar for image gen. Icons help when the list is mostly numbers.
Pattern: Tool Switching in Composer
Pattern: Input Mode Toggle
Web search mode

What works
- Web search chip in the bar with an X. Search is armed visibly, not a hidden globe default.
- Suggested queries under the bar (Best budget laptops, Learn Spanish online…) give one-tap starters.
- Fast effort dropdown stays on the right. Search mode does not swap the whole composer chrome.
What we would push on
- Starters are generic SEO queries. Perplexity ties starters to recency and citations upfront.
- Deep Research is a separate + row. Relationship between search, thinking effort, and deep research is unclear.
Product bet
Web search is a mode chip, not always-on browse. Starters nudge first search without opening a research product.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Removable Web search chip plus query starters under the bar | Mode obvious; low-friction first query | Generic starters; overlap with Deep Research row in + |
Takeaway
When search is a mode, show a removable chip and starters that match your trust story, not generic head terms.
Copy this
- Multimodal attach list in + with explicit file, image, video, audio formats
- Auto / Thinking / Fast effort menu separate from header model versions
- Waveform button for Voice Chat vs Video Chat, mic for dictation only
- Voice session with visible timer (00:02 / 10:00) and immersive orb UI
- Searchable voice persona catalog with language matrix on detail
- Create Image chip with Qwen-Image model picker and aspect ratio in the bar
- Gallery Use Prompt that injects the full generation prompt
- Web search as a removable chip with suggested queries under the bar
Skip this
- Mic and waveform side by side with no labels
- Plus vs Max model rows without cost, speed, or plain-language why
- Eight-row + menu with no home pills for Slides, Web Dev, or Deep Research
- Poetic voice personas as the only catalog tone
- Generic Ask Qwen placeholder while Create Image mode is armed
- Three model names visible (chat header, image engine, effort) with no map
How others design the composer
How other products handle the same job, and what each tradeoff reveals.
Compare composer UX across products
ChatGPT, Claude, Perplexity, and Gemini side by side: default bar, tools, cost, and patterns to copy.
Kimi puts Slides, Docs, and Design on home pills that rewrite the bar. Qwen keeps the same calm shell and packs image, video, research, and dev into +.
Read teardownChatGPT parks tools in + but voice mode is one product. Qwen splits dictation, Voice Chat on Qwen3.5 Omni, and Video Chat on Qwen2.5-Omni.
Read teardownGemini ties image model and aspect into the composer for image mode. Qwen matches that pattern with Qwen-Image 2.0/3.0 and ratio dropdowns plus a Use Prompt gallery.
Read teardownUseful for a critique or spec? Share it.
Original gallery pages: Tool Switching in Composer

