The reframe

It was never “synthetic vs human.” It’s rank vs resolution.

Bain reports synthetic panels replicate conjoint outcomes at ~90% (aggregate). A pre-registered study of 19 experiments finds the individual twin correlation is ~.20, and 76% of people differ significantly from their own twin. Both are right. They are one instrument seen from two distances: it preserves the population average and discards individual detail — because that is what a low-rank representation does. So the real question is never the technology; it’s: what rank does this decision need, and what rank does this instrument deliver?

The same mirror, two accuracies Panel 1 · drag the rank

Each teal dot is a real person; its amber dot is the AI twin, joined by a line. Drag from a high-rank instrument (rich, individual) to a low-rank one and watch the twins collapse onto a few demographic centres: the group average stays right while the individual match falls apart.

high
rich / individual
low
generic / demographic
Real person AI twin Demographic centre
Group-average match .90 stays high at any rank
Individual match .85 rich instrument

At the low-rank end the twins land on ~.20 — the funhouse-mirror finding. Nothing “broke”; the instrument was built for a lower-rank job than an individual question needs.

Illustrative animation. Endpoints: aggregate ~.90 (Bain 2026); individual ~.20 (Peng et al. 2025); rich-instrument individual gain from Park et al. (2024).

Three instruments, three questions Panel 2 · hover or tap a card

Most confusion comes from treating these as rivals. Each answers a different question and fails in a different place — hover a card (or tap on mobile) to reveal where it fails.

Forecast

Simulate the respondent

Role-play a representative person and predict what they’d do to a stimulus you haven’t shipped. Answers a what-if. Strong when validated (Stanford r ≈ .85).

Where it fails ↓

Fails when a forecast is filed as “data,” or the question was never about a future stimulus.

Measure

Read the reflection

Point an instrument at a real artifact — an ad, a pack, a page — and record what its brand signal casts across 8 dimensions. Answers a what-is. No person simulated.

Where it fails ↓

Fails when asked to forecast behaviour or stand in for a person it never measured.

Behaviour

Revealed data

What people actually did — bought, when, how much, where. High-N, individual-linked if it’s card-level, immune to social-desirability.

Where it fails ↓

Thin on the why: low per-person dimensionality; it records behaviour, not meaning.

The five funhouse distortions are one problem Panel 3 · tap each

Peng et al. name five ways a digital twin distorts. They aren’t five separate fixes — they’re five faces of one mechanism: rank collapse. Tap each to see what it is and why it is the same problem.

What it is:

Why it’s rank collapse:

So how do you “address the five”? Don’t buy five expensive fixes. Match the instrument’s rank to the decision’s rank. And note: all five are properties of simulating a respondent. Reading a real artifact has no respondent to distort — that is a different instrument answering a different question.

What this shows. Aggregate accuracy and individual failure are one finding, not two: a low-rank instrument keeps the population average and drops per-person structure, so its useful resolution is only ever what the decision needs of it. That single mechanism is also the five funhouse distortions.

What it does NOT claim. The scatter and the endpoint numbers are illustrative; exact figures are per the sources below. Reading the reflection measures perception — what a real artifact emits and how it is read across eight dimensions — it does not predict desire, demand, or who will buy. If someone says a perception reading forecasts next quarter’s sales, they have slid it back into the forecasting lane. Companion reading: “The Mirror Is the Problem, Not the Reflection.” Related explorables: spectral metamerism (why one impression underdetermines the brand) and the eight-dimensional profile.

Sources. Peng et al. (2025), Digital Twins as Funhouse Mirrors (arXiv:2509.19088) — 19 experiments, individual r ≈ .20, 76% differ, the five distortions; Bain & Company (2026) — synthetic panels ~90% conjoint replication (aggregate); Park et al. (2024), Generative Agent Simulations of 1,000 People (arXiv:2411.10109) — biographical interviews raise individual accuracy; Hewitt et al. (Stanford) — respondent simulation r ≈ .85.

From illustration to measurement

Rank collapse is why a single number can be both accurate and useless: right about the crowd, wrong about the person. A real reading doesn’t pretend otherwise — it reports the resolution it actually has, reads real public artifacts rather than simulating people, and declares a difference only when the signal clears its noise floor. That discipline is the Brand Spectrometer. For what the eight dimensions are and why eight, take the CMO path; for the measurement theory, the empirical path on the Guide.