PickMe logo
PickMeAI Product Visibility Lab

Will AI pick your product?

Crash-test how shopping agents discover, understand, and rank a product—then improve the verified evidence and replay the same test.

100 human-style cases Live agent evidence trace Controlled fix-and-rerun loop
The behaviour shift

Shoppers stopped speaking in keywords.

THEN · SEARCH

“compression sleeves women”

Keyword match → product grid → human compares pages.

NOW · AGENTIC SHOPPING

“What can I wear for tired legs during long nursing shifts?”

Intent → clarification → retrieval → evidence comparison → recommendation.

The AI agent is becoming the first shelf a product must earn.

The blind spot

A suitable product can disappear four different ways.

01

Not retrieved

The page does not express the shopper's language.

02

Not understood

Important attributes or use cases are unclear.

03

Not verified

The agent sees a claim but cannot support it.

04

Outranked

A competitor provides clearer evidence of fit.

Search analytics show the outcome. Merchants still cannot see where the agent lost confidence.

The product

PickMe makes the recommendation journey testable.

01

Submit

Product URL + natural buyer intent

02

Retrieve

Search a real fashion evidence corpus

03

Stress

Run discovery and 100 cases in parallel

04

Diagnose

See rank, score, competitors and evidence gaps

05

Retest

Edit verified metadata and replay the suite

DiscoverUnderstandVerifyCompareRecommend
Agent architecture

One request. Two specialised branches.

GPT-5.6 TERRA · HIGH

Discovery agent

  • Clarifies known and missing requirements
  • Inspects 25 evidence candidates
  • Builds a seven-checkpoint path
  • Ranks, scores and proposes grounded fixes
GPT-5.6 LUNA · LOW

Shopper simulator

  • Runs four parallel batches of 25
  • Tests six real-world writing patterns
  • Varies disclosure and dialogue stage
  • Returns independent ranks and top picks
Adversarial coverage

One need. One hundred ways humans might express it.

100

Controlled cases

Balanced across four independently processed batches.

6

Writing patterns

Simple, Singlish, shorthand, constraints, ambiguity and context shifts.

5

Dialogue stages

From vague opening to clarification, preference shift and refusal.

“need smth for my legs, shift damn long leh” should not be a completely different product universe.
Observability

See what the agent checks—not just its final answer.

PickMe streams safe reasoning summaries, environment actions, inspected evidence, rank movement and batch progress as the run happens.

  • No fake countdown
  • No private raw chain-of-thought
  • Every final claim links back to supplied evidence
01 · ASK SHOPPER clarify fit and compression needs
02 · SEARCH [calf sleeves nursing standing]
03 · INSPECT 25 candidates retrieved
04 · COMPARE target lacks numerical fit ranges
05 · RANK target moves #7 → #4
Actionable result

From “the AI said no” to an evidence-level diagnosis.

72/100
Rank #4 because the listing names the use case—but omits measurable fit, material and care evidence.

The merchant sees the leaderboard, competitor advantages, 100 outcomes, seven discovery checkpoints, and publishable fixes.

Controlled optimisation

Change the evidence—not the exam.

BASELINE

Weak or incomplete listing

Run 100 fixed shopper messages and save every definition.

ScoreRankTop-5 coverage
VALIDATION RERUN

Evidence-preserving metadata edit

Replay the exact prompts. Only product evidence and resulting outcomes change.

Same intentSame casesSame corpus

PickMe never guarantees a higher rank. If a competitor is still the better fit, the remaining gap stays visible.

Live demo path

Compression sleeves: broad intent → measurable evidence.

01 · BASELINE

Ask naturally

“What can I wear to help with tired legs during long nursing shifts?”

02 · OBSERVE

Watch both agents

Follow retrieval, comparison, 100 cases and rank movement live.

03 · FIX

Preserve verified facts

Expose calf sizes, compression by style, care and supported use cases.

04 · REPLAY

Measure the delta

Compare score, primary rank and Top-5 coverage on the same suite.

The product page changes live. The validation standard does not.

Built for a credible demo

Real products, real competitors, deployable architecture.

5

Editable products

Amazon-style storefront and evidence pages.

5K

Deployable corpus

Stratified competitors in a 5.9 MB SQLite FTS5 index.

826K

Full local index

Unique Amazon Fashion records for large-corpus experiments.

Live

Streaming stack

Next.js, OpenAI Responses API, structured output and NDJSON.

Next.js 16React 19TypeScriptSQLite FTS5Vercel-ready
Why PickMe

Not AI SEO. AI-commerce observability.

Stress before launch

Test messy shopper language before real agents decide what to surface.

Diagnose the actual gap

Separate retrieval, evidence, comparison and correct-exclusion failures.

Improve responsibly

Turn supported facts into clearer metadata without inventing claims.

When AI becomes the shelf, brands need a way to test whether the shelf understands them.

PickMeTest · Diagnose · Fix · Replay