In January 2024, Pinterest searches for "mob wife" went from nothing to viral in three weeks. Nobody had used the phrase before. The look itself was older than most of the people Googling it. By April, Zara had nine SKUs you could call mob wife. By July the search interest collapsed.
That was a micro trend with a six-month half-life. Buyers who placed orders in February got nine months of inventory for a six-week sell-through. Buyers who skipped it missed the only fashion sub-genre that mattered that quarter.
The dominant industry tools for fashion forecasting watch Instagram and tell you what's already popular. They cannot tell you what's coming, when it will peak, how long it will sustain, or what to do with the information. Most folks haven't touched the surface.
Fashion may be the most cultural space we have. Forecasting it is one of the hardest multimodal reasoning problems in AI. The field doesn't have a world model for it yet. We're trying to build one.
What a world model has to do
The thing being modeled is cultural taste, which evolves continuously in a high-dimensional space we only see through partial signals. Designers show looks. Influencers post outfits. Search queries spike. Magazines name the moment. None of these alone tells you what aesthetic is forming underneath.
A useful world model does four things at once. It encodes the latent state from multimodal observations. It rolls the dynamics forward in time. It reasons hierarchically, from mega cultural shifts down to single-attribute styling moments. And it reasons about counterfactuals when external events perturb the system.
None of these are solved for fashion right now.
Step 1: a joint embedding
Modern joint vision-language models like Google's SigLIP form the foundation of any world-model architecture. Image and text into a single latent space where same-concept observations land near each other. LeCun's JEPA framing (Joint Embedding Predictive Architecture) describes the world-model paradigm as exactly that: joint embedding plus predictive dynamics on top.
Generic joint embeddings don't capture fashion. Silhouette, material, color, pattern, and cultural moment compose every look in ways that general image-text alignment can't differentiate. A domain-specific embedding is not optional.
Our open-source work: MODA
We open-sourced MODA, a fashion embedding model that beats Marqo's FashionSigLIP on LookBench image-to-image retrieval. Fine R@1 of 67.68 vs FashionSigLIP's 63.84. A 3.84-point gain. It's the latent-state encoder for everything downstream.
The MODA repo also ships the first open-source end-to-end benchmark for fashion search: 253,685 purchase-grounded H&M queries across 105,542 products. We had to build the benchmark ourselves. There wasn't one.
Do take a look at our work (Shameless promotion)
huggingface.co/HopitAI
github.com/hopit-ai/Moda
Why fashion needs reasoning and forecasting together
Most multimodal AI work picks one. Reasoning (VQA, captioning, retrieval) or forecasting (time series, sequential modeling). Fashion forces both, plus more.
The data is multimodal across at least four sources: runway, social images, search-intent traces, e-commerce. Each has different resolutions, lags, and reliability. Naming is generative. When Pinterest searches for "mob wife" spike, no one had ever called it that before. The system has to figure out what to call it. The dynamics are non-stationary, adversarial, cyclical, multi-horizon. The TikTok-era cycle is shorter than the Instagram era, which was shorter than the WWD era. Y2K returns. Cottagecore returns. A fast-fashion buyer needs six-week confidence. A wholesale buyer needs six-month peak timing. A capsule-line strategist needs a two-year early warning.
No single AI sub-discipline solves this. World-model architectures do, in principle, because they integrate all four requirements into one system.
How we're approaching it - Two-step architecture, faithful to JEPA.
Step one is the MODA fashion embedding. The joint embedding layer.
Step two is a substrate of latent aesthetic factors derived from three decades of designer collections. The predictive dynamics layer. Each factor is a coherent dimension: minimalist tailoring, opulent maximalism, athleisure layering, dark academia. When a query arrives, we project its constituent attributes onto the substrate and roll its weighted blend of factor curves forward in time. The rollout is the forecast.
The integration covers runway, Instagram, Pinterest, and Google Trends. Runway is the upstream lever. The trickle-down literature has been describing this signal since the 1960s, but most modern fashion AI tools ignore it. We built the system around it.
What's working
On a labeled backtest of 25 historical microtrends (tomato girl, mob wife, balletcore, clean girl, gorpcore, the whole 2020-2024 catalog):
Substrate signal alone: 76% peak-month accuracy within ±12 months.
Multi-signal ensemble: 96%, no per-trend tuning.
For four of the trends (balletcore, whimsigoth, tomato girl, gorpcore), the system flagged emergence 18 to 29 months before consumer awareness peaked. Earlier than the entire ready-to-wear buying cycle.
The MODA embedding beating FashionSigLIP shows the joint embedding piece works. The substrate composition rollout shows the predictive dynamics piece works. Multi-signal integration cancels orthogonal error and adds 20+ points to single-signal accuracy.
What isn't working
A SigLIP-text-to-IG-image visual neighborhood clustering attempt failed completely. FashionSigLIP's text encoder doesn't know aesthetic names. "Tomato girl" and "y2k cyber" return cosine similarity of 0.07 to 0.18 against random IG fashion images. Barely above the 0.03 to 0.10 noise floor. Zero of 25 queries clustered cleanly. Generic vision-language embeddings collapse where fashion-specific semantics are needed.
Our lifecycle classifier is strong at post-peak detection (around 80%) but weak at at-peak detection for short-cycle micros (around 12%). That's a substrate-vs-consumer-phase artifact: for long-lead trends, the substrate has already moved to Decline by the time consumer press hits Peak. The Sproles 5-state vocabulary fits. The per-state thresholds aren't fully calibrated yet.
The validation is brutal. 25 hand-curated historical microtrends is a tiny dataset. Five of those 25 have no clean cluster identity. Tomato girl, indie sleaze, latte makeup, preppy revival, eclectic grandpa. All real cultural moments. None formally clustered anywhere.
There is no real benchmark
This is the deeper problem. The field has no shared evaluation standard for trend forecasting. WGSN, Heuritech, EDITED are proprietary, unreplicable, untestable. Academic work has been on small private datasets that don't generalize.
We're building benchmarks because we have to. The MODA fashion-search benchmark is one. The 25-microtrend labeled forecasting backtest is another, expanding to 100+ trends next quarter. Both will be released. The field can't make real progress while every claim is unverifiable.
Culture is way harder than enterprise
Most multimodal AI builders work on enterprise problems where the data is clean. Documents, transactions, supply chains, code, customer support. Everything is structured, bounded, named, logged. There are SKU IDs and account IDs and ticket IDs. Financial ground truth at the end of every quarter.
Culture has none of that. No canonical IDs for vibe shifts or aesthetic moments. Naming is emergent and adversarial. Reality is cyclical: Y2K Revival 2022 vs Y2K 2002 are the same trend and not the same. The ground truth lags by years. You find out a trend mattered by what sold next season, and most fashion brands don't share sell-through data.
Building infrastructure that reasons about culture is qualitatively harder than building infrastructure that reasons about enterprise. That's probably why most folks haven't tried.
What we'll open source
Most of what we build will be released.
MODA fashion embedding (released).
MODA fashion search benchmark (released).
Substrate forecasting engine (next month).
Lifecycle classifier (Sproles 5-state, type-aware thresholds, horizon-aware actions).
Sustain-vs-fizzle classifier for the 1-2mo fast-fashion horizon.
Long-lead detector for the 1-year aesthetic-strategist horizon.
25-microtrend labeled forecasting backtest (expanding to 100+).
All notebooks and validation harnesses.
Data infrastructure stays internal. Scraping, vendor agreements, production pipes. The science part should be public.
Next: Causality
When the engine says an aesthetic will peak in 24 months, a merchandiser deserves to know why. The Barbie movie? A specific Resort collection? A cultural undercurrent that started building in São Paulo and Seoul before it arrived in New York?
We're adding Robust Synthetic Control for per-trend counterfactual estimation. Difference-in-Differences on a curated ledger of catalyst events. Causal Forests for heterogeneous treatment effects.
We haven't built the entire world model for fashion yet. We've shipped the first version of the joint-embedding layer (MODA) and shown the predictive-dynamics layer works on a small labeled backtest. The full architecture, validated at scale and explained causally, is the next part of work.
V1 of the forecasting engine ships next month. The causality work runs concurrent. The SME validation study runs after.
If you're building anything adjacent - culture forecasting, multimodal world models, fashion data infrastructure - we'd like to know. Talk to us.
We have posted this blog on substack too - https://hopitai.substack.com/p/why-fashion-needs-a-world-model