Companion study: Which answer did the 17-year-old write? (Lee, 2026)
Methodology & Findings
This repository contains a six-model Vision-Language Model (VLM) study evaluating context-driven valuation bias. The experiment analyzes how visual and textual contexts inflate or deflate price estimates across frontier models.
- Sample Size: N = 4,604 trials
- Models Evaluated: 6 frontier Vision-Language Models
- Scoring Method: Dual-anchor (receipt vs. retail) analysis
Project Structure
-
scripts/: Python execution scripts for running model evaluations. -
analysis/: Data processing and metric evaluation scripts. -
human-baseline/: Experimental protocol for human-rater baseline comparison.
Further Reading
- Dear OpenAI, (Substack) — Visual safety filters and cross-modal valuation anomaly analysis.
- Six AI Models. Six Ways to Fail. (Substack) — Operational breakdown of model behavioral signatures under pressure.
- What If AI Uses Information You Never Intended It to Use? (Substack) — Structural evaluation of the algorithmic halo effect.
- Can your outfit increase a necklace's value by 43×? (Substack) — Empirical context multipliers across frontier VLMs.