Connect with me

Founding Product Designer  ·  AI · 0→1 · Mobile  ·  09/2025 – now

Shellmind is an AI advisor for first-time home renovators. You upload a contractor's quote; it checks prices, specs and hidden risks, and tells you what to do about them. I'm the founding, and only, product designer.

This is the story of V2. V1's AI was right, and 60% of people still left without a decision. Producing an analysis and helping someone decide turned out to be two different jobs.

TeamCEO–PM, 4 engineers, 1 marketer
PlatformWeChat Mini-Program · PC web
ScopeResearch, IA, interaction, the AI's output rules, design system, coded prototypes
<40 →65%+Report completion

Drop-off before acting fell from about 60% to about 30%.

90 →5 minQuote review

What an expert friend used to do over an evening.

4,690Organic users, 30-day V2 launch

Contributed to a $700K seed round.

Shellmind hero: the quote analysis on a phone — Don't sign this quote as it stands — on a blue gradient.
V2 on the WeChat Mini-Program: the verdict comes first, the evidence one tap below.

Where I started

I bet that showing more AI detail would earn more trust. I was wrong.

No product, no process when I joined. I picked the one question that could change our direction: does showing more of the AI's work make people trust it more?

I thought yes. V1 put the full analysis on display for about 200 beta users — and the answer came back no. That wrong bet is where V2 starts.

V1 report: overall rating Poor, grade D, then Key Price Analysis. V1 report: a long price table and View all 26 items. V1 report: ten missing specs listed in full.

V1, the "show everything" bet, as 200 beta users saw it.

01 · Research

The AI was right. 60% still left without a decision.

The baseline

Only 40% took any action

Mobile V1, a beta with about 200 invited users. The analysis was complete and mostly correct; 90% got through it, 65% opened the report, 40% acted.

So the goal wasn't a better report. It was getting a first-time homeowner from a quote to a confident next sentence: a question to ask, or a number to push for.

Sankey diagram of V1's 200 beta users: 100% upload, 90% reach analysis, 65% review the report, 40% take action, 14% return.

Method

12 think-alouds, one honest question

"Is this clear?" gets a polite yes. "What will you ask him?" shows whether the analysis has become the person's own judgment. Confidence broke at every handoff, not only in the report.

/02Wait

"It's just spinning… is it stuck?"

"Almost done" for a minute or two, and no word on whether leaving was safe.

/03Report

"Too much text. Where's the verdict?"

A grade, then long tables in jargon, all with equal weight.

/04Act

"How do I use this to negotiate?"

It ended on three generic questions. Nothing said what to ask for.

02 · Reframe

The unit wasn't the report. It was the decision.

One test at every stage: after this screen, can the person reach the next one on their own? One redesign became four stages.

the decision
/01Understandthe verdict
/02Verifythe evidence
/03Acton it
Journey map with four stages: start and provide context, wait for AI analysis, understand the report, act on the findings, with user questions under each.
The real journey had four stages. V1 lost people at each handoff.

/01Start

V1 assumed you arrived with a quote in hand.Upload stays the hero; a chat catches everyone else.

/02Wait

A full-screen "Almost done", for minutes.The wait lives in the chat — free to leave.

/03Report

Everything the AI found, in the AI's order.Keep everything, reorganise around the decision.

/04Act

V1 ended on three generic questions.Not advice — the actual words to send.

/01Start

From a tool page to an advisor

"How do I start if I don't have a quote yet?"

Upload stays the hero. An always-on chat catches everyone else. It answers first, asks only what changes the advice, and saves your profile visibly.

What it cost: AI spend on chats that never become an upload. Contained — every chat keeps a way back: "I got a quote".

Both paths lead to the same place: upload or chat, then analysis, then a clear verdict.

Two doors, one analysis.

Home · upload stays the hero

V2 home with an Upload quote button, topic chips and a composer.

Upload · try a sample

V2 upload sheet with a Try a sample link.

Add info · the chats quotes leave out

V2 add supporting info: quote files and vendor chats.
Answer firstA real price range before any questions.
Saved, visiblyThree taps become a profile — no "Want me to save this?".

Then ask · three quick taps

Three quick taps about the home.

The whole flow · recorded

/02Wait

From "is it working?" to "you can leave"

"It's just spinning… is it stuck?"

The wait lives in the chat, as a status card. Steps, time left, "you're free to leave" — and it asks only what changes the verdict, mid-check.

What it cost: the check pauses to ask, and a card can scroll away. Contained — progress saved, status in the header even folded, answers editable later.

V1 · "Almost done", full screen

Explored A · task dashboard

Explored option A: a task dashboard with a question sheet covering it.

Explored B · card stack

Explored option B: a turtle with a stack of task cards.

Both full-screen options hold you hostage. That's why the wait moved into the chat.

Free to leaveAbout 5 min for 6 pages — it says so up front.
Done folds awayThe report entry stays in the thread; Home brings you back.

Running · in the conversation

Done · folded into the thread

V2 done: the card folds and the report entry stays in the thread.

Asked mid-check · which total is right

Editable later · re-prices in seconds

Ask only what changes the verdict, when it matters — and the answers stay editable.

/03Report

From everything the AI found to what you should do

"I read the whole thing, but I still didn't know what to do."

Keep everything. Reorganise it around the decision. Verdict, then the price to aim for, then what's worth arguing about. That changed the AI's output rules, not just the layout.

What it cost: evidence sits one tap down. Contained — every verdict carries its reason, and every expand row says how much is behind it.

Six flat report sections become three layers: decision, evidence, deep dive.

Six flat sections become three layers: decision, evidence, deep dive.

Verdict first · with its reason

V2 verdict: Don't sign this contract yet, with the reason and next step.

Price guidance · recorded scroll

Worth arguing · biggest first

V2 where you're overpaying: 18 items, biggest first.
Drag the dividerThe same analysis, reorganised around the decision.
V1 → V2Grade and tables become a verdict with a next step.
V2 report: the verdict comes first — Don't sign this contract yet. V1 report: a grade, then tables grouped by analysis type. V1V2
Two modes: Verdict Seekers who want 60 seconds, Detail Auditors who check every number, with V1 in the valley between.

Two reading modes. Cutting fails the auditors; V1's tables failed the seekers. Reorganising serves both.

/04Act

Don't just read the report. Know what to say next.

"How do I use this to negotiate?"

Not advice. The actual words. An opening offer, an "Ask for" line on every red flag, and a script written to be sent as-is.

What it cost: the AI now speaks for the person. Contained — numbers come as ranges, the reason sits beside every line, and scripts are copied, never sent for you.

Ask for, in writingEvery red flag ends in a sentence you can send.
Copied, never sentYou stay the sender — the script is yours to edit.

Red flags · most serious first

What to say · opening offer + script

If they push back · three ready replies

If the vendor pushes back: three ready replies.

The math · as a sheet

Next steps and a cost breakdown spreadsheet.

V1 ended with a download button — V2 ends with the user's next sentence, and replies for when the vendor pushes back.

03 · System

When the system isn't sure, it says so, and never loses your work

You're offline, with one Retry.Offline

Said once, in the reply. One Retry.

Recoverable
Page 2 is too blurry: retake just that page.Page too blurry

One clear photo, never a re-upload.

Trustworthy data
Which total is right? Pick one to analyze.Total doesn't add up

Nothing concluded from unconfirmed data.

Trustworthy data
Running slower than usual, you can still leave.Slow analysis

Slow isn't failed. Leaving is safe.

Honest status
Analysis didn't finish: your file is saved, run it again.Analysis failed

File saved. Run again, no re-upload.

Recoverable

04 · From design to build

Figma is the single source of truth

Tokens and screens live in Figma; behaviour ships as coded pages engineers can click. Scroll — or run each one live.

Tokens

Design system playground

49 hex values, 72 color roles, live contrast checks — pick a brand color and watch every role re-derive.

Components, every state

Button spec

Three faces — glass, ink, outlined — every state, and the rules written down beside them.

The working prototype

Coded prototype

The flow engineers clicked instead of reading: upload a quote, answer the mid-check questions, open the report.

Tokens 问龟 Shellmind design system playground: brand color with live contrast checks, theme toggle, 49 hex values and 72 color roles.
Components, every state Button spec page: glass hero CTA, ink primary, outlined secondary, with the rules — one ink primary per screen, press is instant.
The working prototype The coded prototype: the phone, interactive — upload a quote, answer the mid-check questions, open the report.

05 · Impact

Three measures, because they answer three different questions

<40 →65%+Did they finish the report?

Reached the end of the core report path within one session.

~40 →75%Do they know what to do?

In moderated sessions, could name a specific next step unprompted.

~60 →30%Did they reach an action?

Drop-off before an action, read alongside events like copying a script.

Results for V2 as a whole — no single screen gets the credit. In its first 30 days V2 reached 4,690 organic users, contributing to a $700K seed round.

06 · Takeaways

My takeaways

1. Reframing the problem was the real design work

The brief said "make the report better." The real unit was the decision — and that call wasn't visual, it was definitional.

2. Every decision should name what it gives up

Saying the cost out loud, and how I'd know I was wrong, is what made my calls easy to back.

3. Design the moments between screens, not just the screens

People left between screens: before a quote, during an unsafe-feeling wait, after a report that stopped short of the next sentence.

What I'd do differently

  • Measure what happens after people leave the app.
  • Check time estimates against real runs, early.
  • Watch the chat-to-upload rate from day one.

The AI is usually right. People still have to decide whether to trust it. That's the moment I love designing for.

Manya ZhuFounding Product Designer, Shellmind
Next