Methodology
A population you can check, then ask
How the synthetic population is built from public statistics, how every build is validated, what an answer contains, and which known methods all of it stands on.
How it works
The Census Bureau's 2025 population estimates give every state's population by single year of age, sex, race and Hispanic origin. That joint is reproduced exactly.
American Community Survey microdata is raked to that skeleton and to published tables, so education, work, income, family, citizenship, disability, insurance, commute and housing carry the survey's real joint structure. Health and behaviour come from the NHIS by demographic cell.
Every build holds out published tables it never fitted and reports the error, per state. A number is only claimed with its held-out score.
A representative sample of profiles answers the question, ten at a time, each as itself. The interval narrows as the sample grows; the reasons people give are collected across the answers.
What an answer contains
Not one model’s opinion. A representative sample of synthetic people, each answering as themselves, with everything needed to judge the result.
The share choosing each option, or the mean, with a 95% interval that narrows as more profiles are asked.
The running estimate batch by batch, so you watch it settle rather than trust a single number.
A handful of the people asked explain their answer in their own words, across the answers, with a short reading of what drives the split.
The answer by sex, age, education, employment and state, from the sample's own attributes, with the biggest gaps called out.
Synthetic profiles and their individual answers. Tangible, and never a real person.
Exactly how the question was put to each person, the population it was asked of, the accuracy record for that kind of question, and the assumptions made along the way.
Validation comes first
A number is only worth what its check is worth. Every kind of question carries its own accuracy record, checked against real surveys, and every answer says which. Every check, with its number.
Every state's population by single year of age, sex, race and origin is reproduced within sampling error. Not approximately: cell by cell.
Published tables are held out of the calibration and predicted afterwards. Employment by sex and age across all states: median error 2.8%. Marital status, veterans, disability, insurance, citizenship, occupation, industry, commute and enrollment are checked the same way.
On 63 attitude and self-report questions from the General Social Survey, simulated answers sit a median 0.23 total-variation distance from real answers within sex, age and education groups: 0.17 on yes-or-no questions, 0.24 on choices, 0.27 on scales and 0.30 on amounts, where the model still gives rounder numbers than people do. The misses are reported with the hits.
Where a kind of question has not been checked against a real survey yet, the answer page says so in the same sentence that gives the interval.
Three ways to get an answer
Recruit a panel, prompt a model, or ask a calibrated population. They differ in what the answer is grounded in.
| A survey panel | Personas from a prompt | census.world | |
|---|---|---|---|
| Who answers | Recruited respondents, weighted afterwards | Characters a model invents from a description | One synthetic profile per resident, built from census and survey records |
| Representative of | The panel, after weighting | Whatever the prompt implies | The population, cell by cell: state, age, sex, race, origin, education, work, income, household, health |
| Sampling interval | Yes | No | Yes, on every number, narrowing as more profiles are asked |
| Checked against | Itself | Usually a handful of past surveys | Official tables held out of the build, and real survey answers by demographic group |
| Reproducible | No: a new field, a new sample | No: a new prompt, a new crowd | Yes: versioned builds, the same person in every build of a lineage |
| Time | Weeks | Minutes | Minutes, with the interval shown as it forms |
| Cost | Per respondent | Per call | The first hundred profiles of every question are free; full runs by arrangement |
Standing on known methods
Nothing here is a new idea. The population is built with methods statisticians have used for decades, applied end to end and checked at every step.
- 1940Iterative proportional fittingDeming and Stephan's method for adjusting a table to known margins; the basis of raking in survey weighting ever since.In census.world todayAmerican Community Survey records are raked to every state's census cell counts and to published tables, so each state's education, work and income structure matches what the Census Bureau reports.
- 1996Synthetic populations for microsimulationBeckman, Baggerly and McKay's synthetic baseline populations for transport models: survey records expanded to a whole population that matches census constraints.In census.world todayOne synthetic profile per resident, at any scale up to one-to-one, carrying the survey's joint structure rather than independently drawn attributes.
- 2009Held-out validation of synthetic populationsThe practice, common since the population-synthesis literature of the 2000s, of scoring a synthetic population on tables it was not fitted to.In census.world todayEvery build predicts employment, marital status, veteran status, disability, insurance, citizenship, occupation, industry, commute and enrollment tables it never saw, and reports the error per state.
- 2023Language models as simulated respondentsArgyle and colleagues' finding that a language model conditioned on a real respondent's demographics reproduces survey response patterns, and the wave of work testing where that holds and where it fails.In census.world todayEach profile answers in character from its own record; the answers are checked against the General Social Survey by demographic group, and the page says how far off that kind of question tends to be.