Qwen output UX: images, video, search & selection
Qwen threads treat media and search as continuations, not side quests. Upload a portrait and the model writes a long image prompt before generate. Finished images get Create Video and Edit overlays. Say "animate it" and the bar picks up a Create Video chip with the source thumbnail. Web search answers ship with Searched the web, inline domains, a Thinking and Search sidebar, and selection actions (Copy, Ask Qwen, Explain, Translate). The UI looks trustworthy. In this capture it also confidently cites a fictional Apple MacBook Neo for 2026, which is the trust gap the chrome hides.
Photo to image prompt

What works
- The assistant reply is a full art-direction brief (yarn doll, lighting, texture) instead of jumping straight to pixels.
- User sees the plan in chat before spend. Easier to edit text than regenerate blindly.
- Create Image chip and Qwen-Image 2.0 / 16:9 stay pinned in the composer through the reply.
What we would push on
- The prompt is long and single-block. No highlight of what changed from the upload.
- Send stays greyed while mode chips are set. Unclear if user must approve the text prompt first.
Product bet
Showing prompt expansion builds trust for image spend. Users feel the model understood the reference photo before GPU time runs.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Visible text prompt step before image generate | Edit scope before render; teaches prompt quality | Extra scroll; slower time-to-image |
Takeaway
For reference-image gen, show the rewritten prompt in-thread and let users edit it before render.
Pattern: Progressive Disclosure
Image in thread

What works
- Image renders inline at readable size with share, regenerate, and overflow on the message.
- Create Video and Edit float on the image. Next steps are on the artifact, not buried in +.
- Thumbs, share, and download also sit top-right on hover for quick feedback and export.
What we would push on
- Composer still shows 16:9 while the output image is square. Format chip and result can disagree.
- AI-generated content may not be accurate disclaimer is easy to miss under the bar.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Inline image with Create Video / Edit overlays plus message actions | Chains image to video without re-upload | Aspect chip can lie; disclaimer fades into footer |
Takeaway
Put chain actions on the image itself. Match aspect label to what actually rendered.
Animate with reference

What works
- Source thumbnail appears in the bar above animate it. Context is visible before send.
- Create Video chip replaces Create Image automatically. Mode follows intent.
- Short verb prompt works. User did not re-describe the doll.
What we would push on
- No motion controls (duration, camera move) in the bar yet, only 16:9 later.
- Chip swap is smart but silent. A one-line "Switching to video from this image" would teach the pattern.
Product bet
Image-to-video is the upsell path. Reference thumbnail plus auto chip swap keeps users in one thread instead of opening a video app.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Reference thumbnail in bar with auto Create Video chip | Natural language handoff; no re-upload | Limited motion controls at prompt time |
Takeaway
When a reply references prior media, pin the thumbnail and swap mode chips to match the next job.
Pattern: Tool Switching in Composer
Video generation progress

What works
- Gradient card with spark icon and 4% label. Users know video is rendering, not stuck.
- Create Video chip and 16:9 stay in the composer so settings survive the wait.
- User bubble animate it stays minimal. Focus is on the job card.
What we would push on
- No ETA or stage name beyond percent. Long video jobs will feel opaque.
- Send button greys out. Cannot queue a follow-up until render finishes.
Takeaway
Show percent for video gen, but add stage copy or ETA when renders run longer than a few seconds.
Pattern: Progressive Disclosure
Preview, download, publish

What works
- Modal player with 00:00 / 00:05 timer. Short clip length is obvious before download.
- Download and Publish are equal primary actions under the player.
- Create Video chip and 16:9 remain in the bottom bar for another pass without closing context.
What we would push on
- Publish destination is not named on the button. Trust requires knowing where it goes.
- Square video in a 16:9 modal adds letterboxing. Same aspect mismatch as image output.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Modal preview with Download and Publish plus persistent mode bar | Clear export path; easy retry | Publish target vague; aspect mismatch visible |
Takeaway
Treat video like a deliverable: preview, duration, download, and name the publish target on the button.
Web search answer

What works
- Lead sentence states intent: I'll search for information about the best budget laptops…
- Searched the web row is tappable. Signals retrieval happened.
- Inline grey domain chips after claims (pcworld.com, pcmag.com, cnet.com) mirror Perplexity-style trust chrome.
- Structured list with bold product names and price line reads scannable.
What we would push on
- Answer cites 2026 reviews and an Apple MacBook Neo at $599. That product does not exist. The UI still looks authoritative.
- Domains are not numbered footnotes. Hard to match a claim to a specific source click.
- Disclaimer under the bar is tiny compared to the confident list above it.
Product bet
Search mode competes on Perplexity-shaped answers. Citations and Searched the web sell freshness even when the model fills gaps with fiction.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Chat-formatted search answer with inline domain chips | Familiar reading flow; looks researched | Hallucination wears a citation costume; weak claim-to-source link |
Takeaway
If you show domain chips, tie each claim to a numbered source and surface when retrieval failed or dates look wrong.
Selection actions

What works
- Floating menu: Copy, Ask Qwen, Explain, Translate with a language submenu.
- Translate lists many locales in one scroll (English US, Arabic, German, Spanish, Persian, French, Hindi, Italian, Japanese, Korean…).
- Selection works on search answers, not only plain chat prose.
What we would push on
- Ask Qwen vs Explain overlap. Users will not know which to pick.
- No Check sources or Cite selection action despite search mode being on.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Generic selection menu with translate submenu on search output | Reuse chat refinement patterns on research answers | Missing search-specific actions; Ask vs Explain redundant |
Takeaway
Extend selection menus for search threads with source-check actions, not only translate and explain.
Follow-ups and source badges

What works
- Three follow-up pills continue the research thread (under $500, battery life, what to look for).
- Message actions row includes copy, feedback, share, regenerate, and globe badges with +17.
- Web search chip persists in the composer so the next turn stays in search mode.
What we would push on
- Follow-ups repeat the 2026 framing. Bad suggestions reinforce bad facts.
- +17 globes are not clickable domains in this view. Count impresses more than it informs.
Tradeoff
| Decision | Benefit | Cost |
|---|---|---|
| Follow-up pills plus source count badges on search answers | Keeps research session going; signals many sources | Suggestions can amplify errors; badge count without list |
Takeaway
Follow-ups should deepen verification, not repeat the same shaky premise. Make source badges open the drawer.
Copy this
- Prompt expansion in-thread before image generate from a reference photo
- Create Video and Edit overlays on finished images
- Reference thumbnail plus auto Create Video chip on animate it
- Video progress card with percent while mode chips stay in the composer
- Preview modal with Download and Publish for short clips
- Searched the web row plus inline domain chips on search answers
- Thinking and Search sidebar with numbered snippets and +N overflow
- Selection menu with Translate submenu on research output
- Follow-up pills that match an armed Web search composer chip
Skip this
- Citation chips on answers that cite fictional products
- 16:9 selected while square image or letterboxed video renders
- Publish button with no destination named
- Ask Qwen and Explain both on selection without guidance
- Follow-ups that double down on wrong dates or products
- +17 source globes that do not open a source list
How others output, artifacts & refinement
How other products handle the same job, and what each tradeoff reveals.
Compare output & artifacts UX
ChatGPT, Claude, Perplexity, and Gemini side by side: refinement, artifacts, export, and patterns to copy.
Perplexity numbers citations and centers sources. Qwen keeps chat-shaped answers with grey domain chips and an optional Thinking and Search drawer.
Read teardownChatGPT chains image and video in-thread but search is not a composer chip with starters. Qwen arms Web search in the bar and suggests head-term queries below it.
Read teardownGemini output teardown covers Guides and Workspace exports. Qwen focuses on generative chain (photo to prompt to image to video) plus search refinement menus.
Read teardownUseful for a critique or spec? Share it.
Original gallery pages: Output & Refinement
