Show HN: I put a $2.43 necklace on 3 outfits. VLMs priced it at $19 to $104
BrianneLee011
18 points
24 comments
July 28, 2026
Related Discussions
Found 5 related stories in 379.4ms across 15,236 title embeddings via pgvector HNSW
- Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models adam_rida · 310 pts · July 23, 2026 · 48% similar
- Show HN: Halo – open-source, tamper-evident runtime evidence for AI agents brian_kuan · 29 pts · July 07, 2026 · 48% similar
- Show HN: Vibe Coding a $20k /Year Enterprise Logistics Platform ryanckulp · 29 pts · May 15, 2026 · 47% similar
- Show HN: Oodle.ai – $10 per million agent traces kirankgollu · 29 pts · July 14, 2026 · 46% similar
- Show HN: Free tool to see how much AI bots are costing your site plaintosapp · 18 pts · May 11, 2026 · 45% similar
Discussion Highlights (7 comments)
BrianneLee011
I wanted to see how much environmental and framing cues distort object valuation in vision-language models. I bought a $2.43 chain necklace and $0.71 earrings on Temu, photographed them across three outfits (tailored blazer, party dress, recycling yard flannel), plus an isolated flat-lay control, and ran ~1,500 stateless API sessions across 6 models (Claude Fable 5, GPT-5.6, GPT-4o, Grok 4.5, Kimi K3, DeepSeek V4). A few interesting findings: - The Halo Multiplier: Models priced the exact same physical necklace anywhere from $18.80 to $103.90 depending on attire (3.6× halo). - Isolation Controls (F2): Using a flat-lay control (S4) unmasked two opposite mechanisms: Claude’s bias is formal inflation (formal attire inflates value above baseline), while Kimi’s bias is casual deflation (yard attire depresses value below baseline). - Post-hoc Material Stories (F6): Models invent visual evidence to justify their priors—GPT-5.6 and Kimi started describing the base metal as "gold-plated" or "gold vermeil" almost exclusively under formal framing. - Denial without Correction (F7): When asked sequentially if clothing changed its answer, Claude admitted it 100% of the time, while GPT-4o denied it 82% of the time despite exhibiting a 3.9x text halo. The full dataset (N=4,604 analysis rows), evaluation scripts, and protocol specs are in the repo. I’d love to hear feedback on the experimental design or ideas for follow-up behavioral probes!
nemomarx
This is an interesting test, but it does seem to me that visual models being able to price things was already kinda unrealistic? It makes me think of the calorie guessing use case. you can't tell the difference between materials and ingredients in a photo, so how will the model? especially "in situ" as part of an outfit or in a finished meal. maybe they could do it if you placed them on a blank table or background to avoid context? I assume that's the control you mentioned
Invictus0
The photos are terrible, you can barely see the necklace at all. i doubt any human could accurately price a generic necklace from 5 ft away either
Retr0id
It looks like there are exactly 4 "stimulus" images. Fair enough, but I think you'd need a larger and more diverse dataset to form any real conclusions.
reedf1
Jewelry is a veblen good whose price is mostly based on provenance. I'm not sure any classifier could do anything but guess on surrounding context. Insurance classifiers are much less naive than you'd expect, btw. I have designed a few.
GuB-42
A bit off topic, but why are the original author comments flagged to death?
reedf1
SW50cmVzdGluZyBleHBlcmltZW50IC0gZ3JlYXQgZm9yIGEgaGFja2VyIG5ld3MgZGlzY3Vzc2lvbg==