Guide · explorable
Rank vs Resolution
AI digital twins predict group outcomes well and individual outcomes poorly. That is not two findings — it is one instrument seen at two resolutions. This is an illustration; the numbers are from published research. To measure a real brand from public evidence, use the Brand Spectrometer.
It was never “synthetic vs human.” It’s rank vs resolution.
Bain reports synthetic panels replicate conjoint outcomes at ~90% (aggregate). A pre-registered study of 19 experiments finds the individual twin correlation is ~.20, and 76% of people differ significantly from their own twin. Both are right. They are one instrument seen from two distances: it preserves the population average and discards individual detail — because that is what a low-rank representation does. So the real question is never the technology; it’s: what rank does this decision need, and what rank does this instrument deliver?
The same mirror, two accuracies Panel 1 · drag the rank
Each teal dot is a real person; its amber dot is the AI twin, joined by a line. Drag from a high-rank instrument (rich, individual) to a low-rank one and watch the twins collapse onto a few demographic centres: the group average stays right while the individual match falls apart.
rich / individual low
generic / demographic
At the low-rank end the twins land on ~.20 — the funhouse-mirror finding. Nothing “broke”; the instrument was built for a lower-rank job than an individual question needs.
Three instruments, three questions Panel 2 · hover or tap a card
Most confusion comes from treating these as rivals. Each answers a different question and fails in a different place — hover a card (or tap on mobile) to reveal where it fails.
Simulate the respondent
Role-play a representative person and predict what they’d do to a stimulus you haven’t shipped. Answers a what-if. Strong when validated (Stanford r ≈ .85).
Where it fails ↓Fails when a forecast is filed as “data,” or the question was never about a future stimulus.
Read the reflection
Point an instrument at a real artifact — an ad, a pack, a page — and record what its brand signal casts across 8 dimensions. Answers a what-is. No person simulated.
Where it fails ↓Fails when asked to forecast behaviour or stand in for a person it never measured.
Revealed data
What people actually did — bought, when, how much, where. High-N, individual-linked if it’s card-level, immune to social-desirability.
Where it fails ↓Thin on the why: low per-person dimensionality; it records behaviour, not meaning.
The five funhouse distortions are one problem Panel 3 · tap each
Peng et al. name five ways a digital twin distorts. They aren’t five separate fixes — they’re five faces of one mechanism: rank collapse. Tap each to see what it is and why it is the same problem.
What it is:
Why it’s rank collapse:
What this shows. Aggregate accuracy and individual failure are one
finding, not two: a low-rank instrument keeps the population average and
drops per-person structure, so its useful resolution is only ever what
the decision needs of it. That single mechanism is also the five funhouse distortions.
What it does NOT claim. The scatter and the endpoint numbers are
illustrative; exact figures are per the sources below. Reading the reflection measures
perception — what a real artifact emits and how it is read across eight
dimensions — it does not predict desire, demand, or who will buy. If
someone says a perception reading forecasts next quarter’s sales, they have slid it back
into the forecasting lane. Companion reading: “The Mirror Is the Problem, Not the Reflection.”
Related explorables: spectral metamerism
(why one impression underdetermines the brand) and
the eight-dimensional profile.
Sources. Peng et al. (2025),
Digital Twins as Funhouse Mirrors (arXiv:2509.19088) — 19 experiments,
individual r ≈ .20, 76% differ, the five distortions;
Bain
& Company (2026) — synthetic panels ~90% conjoint replication (aggregate);
Park et al. (2024),
Generative Agent Simulations of 1,000 People (arXiv:2411.10109) — biographical
interviews raise individual accuracy;
Hewitt et al. (Stanford)
— respondent simulation r ≈ .85.
From illustration to measurement
Rank collapse is why a single number can be both accurate and useless: right about the crowd, wrong about the person. A real reading doesn’t pretend otherwise — it reports the resolution it actually has, reads real public artifacts rather than simulating people, and declares a difference only when the signal clears its noise floor. That discipline is the Brand Spectrometer. For what the eight dimensions are and why eight, take the CMO path; for the measurement theory, the empirical path on the Guide.