Methods
How the numbers are made
Anyone who needs a defensible number can use this app — a reporter on deadline, a student writing a paper, a researcher checking a hypothesis, or a policy analyst briefing an audience. Every estimate has to survive scrutiny, whatever the setting. Here is exactly what happens between your question and the figure on screen.
Every number comes from the data, not the AI
The AI in this app does one job: it reads your question and routes it to the right survey variable. It never computes an estimate. Every figure you see is calculated by an R server using the actual microdata from the survey's publisher.
Survey weights
Estimates use the weights the survey publishers provide so results represent the U.S. population, not just the people who happened to answer. For the General Social Survey that means the NORC weight (wtssnr for 2004 onward, wtssall for earlier years). For the Current Population Survey ASEC, the person-level weight from Census.gov public-use files.
Uncertainty intervals
Each estimate ships with a 95% confidence interval computed with the survey package in R. If a subgroup is too small to support a reliable estimate, the server refuses to produce one and tells you why — a narrow answer is better than a wrong one.
Full provenance
Every result includes the exact question wording, the variable name, the years covered, the weight used, the unweighted sample size, the dataset release, a codebook citation, and the reproducible R code that produced the numbers. A fact-checker can rerun every figure independently.
Refusals are honest
If a question maps to no known variable, a variable wasn't asked in the years you picked, or the sample is too small, the app says so plainly and asks you to revise. It will never fill the gap with an invented number.
One engine for people and developers
The Ask page and the developer API preview use the same controlled research engine. Language models help interpret questions and explain results, but numerical findings come only from approved statistical computation on the source data — never from generated text.
Ask a question and see the provenance for yourself.