Verasight Report

Probability-Based Samples Can No Longer Serve as the Heuristic for Accurate and Representative Surveys

Peter K. Enns and Joey Marshall

Published Jul 21, 2026

Updated

Probability-Based Samples Can No Longer Serve as the Heuristic for Accurate and Representative Surveys

We work for Verasight, a survey firm that conducts both probability and nonprobability sampling. We have also both relied extensively on probability surveys in other contexts—Enns in his academic research and Marshall in his prior work for the Pew Research Center and the U.S. Census Bureau. We believe the industry needs to revise how it thinks about probability surveys.

For decades, probability-based samples were the gold standard of survey research. For this reason, scholars often view probability samples as a signal of data accuracy. Unfortunately, this sample type is no longer an informative data quality heuristic. As Michael Bailey explains in Polling at a Crossroads, “Almost 70 years into the random sampling era, the paradigm has begun to show its age. Changes in society have made it harder and harder to implement surveys according to the tenets of the paradigm” (Bailey 2024, 46).

Of course, many probability-based surveys produce accurate and representative data. As do many surveys that use nonprobability methods. It is easy to find examples of both sample types that match known population benchmarks (e.g., Enns et al. 2024; Santillan et al. 2024; Tamanna et al. 2025).

We do not question the potential upside of various sample types. High-quality probability and nonprobability surveys remain an important and excellent source of data. Our point is simply that when deciding which surveys to use or analyze, the sample type should not be the determining factor. Knowing if a survey used a probability sample no longer provides a clear signal about whether the sample accurately represents the population of interest.

Quick Background (Some Key Terms)

In a random sample, each individual in the population surveyed has an equal chance of being selected into the sample. The random sample is a special type of probability-based sample. In a probability sample, everyone in the population has a known, nonzero probability of being sampled.

While many factors contribute to the overall accuracy and representativeness of surveys, probability (and thus random) samples will produce accurate estimates of the population—as well as calculable uncertainty around these estimates—when assumptions hold.

A key assumption of probability surveys is that those who take the survey are, on average, equivalent to those who were invited but did not take the survey. Those who are invited but do not take the survey reflect survey (or unit) nonresponse. If those not responding are missing at random, the nonresponse can be ignored.

However, if those who do not take the survey differ systematically from those who take the survey, we have nonignorable nonresponse. If the nonresponders differ on aspects that correlate with what we care about, our estimates will be biased and the survey will not represent the population being studied.

Poststratification weights do not solve the nonignorable nonresponse problem. These weights adjust the sample so proportions in each demographic group in the survey match the actual proportions in the population (often measured in a census or other large government survey). But this weighting also assumes nonignorable nonresponse on any variables or characteristics not included in the weights.

To see why, suppose we conduct a political survey and individuals with lower education levels are less likely to take the survey. Further, suppose that the lower education individuals who do take the survey tend to be much more politically interested than lower education individuals who did not take the survey. These response patterns will lead us to overestimate political interest among the lower education group.

What’s more, this overestimate of political interest will increase in the full sample when we weight the data by education. The increased bias results because weighting increases the influence of respondents with lower education levels, since they are underrepresented in the survey. But these lower education respondents are not representative of the lower education population as a whole. Weighting in this way will increase the influence of non-representative respondents (i.e., lower education respondents who are much more politically interested than those who did not take the survey).

Why We Cannot Assume the Core Assumption of Probability-Based Sampling Holds

In the 1990s, response rates to telephone surveys were 36% (Kennedy and Hartig 2019). When more than one-in-three individuals contacted took the survey, it seemed plausible that the nonresponders were not systematically different from those who took the survey. But response rates have plummeted since the 1990s.

Current response rates for the most rigorous probability-based samples range from about 1.6 to 3.4%.1 These response rates mean that out of every 100 people invited, only 2 or 3 typically take the survey. It seems extremely unlikely that the 2 to 3 who choose to take the survey are not systematically different from the 97 to 98 who do not take the survey.

Indeed, one of us experienced this concern directly after fielding an expensive probability-based survey just over two weeks before the 2020 U.S. presidential election.2 The first question in the survey asked respondents, “On a scale from zero to 10, please indicate how likely it is that you will vote in the presidential elections being held in November.” The second question asked, “Who do you plan to vote for in the 2020 election for president?”

The weighted results from the probability-based survey had Biden winning the two-party vote by 20.8 percentage points—an overestimate of his actual winning margin by more than 16 percentage points. If we limit the sample to those who indicated they had voted or were extremely likely to vote, the Biden vote share overestimate grew to more than 20 percentage points.

Of course, late-deciding voters and not knowing which respondents will actually vote can lead vote intentions in pre-election surveys to differ from the final election outcome (e.g., Jamieson et al. 2023; Shapiro 2020). But these factors cannot account for an overestimate of more than 16 to 20 percentage points just over two weeks prior to the election.

Evidence of nonignorable nonresponse in probability-based surveys extends beyond this single example. Enns and Rothschild (2021) analyzed 355 surveys of likely voters from the two months preceding the 2020 election and found that surveys with probability-based samples produced an average raw error of 6%—more than nonprobability samples and well beyond the expected margin of sampling error. Further, instead of error centered around the final vote, all 81 probability-based samples overestimated Biden’s vote share.3 We would not expect this pattern of results if nonresponse was ignorable in probability surveys.

While low single digit response rates suggest nonresponse cannot be ignored in probability-based surveys, nonignorable nonresponse also exists when response rates are much higher.

Tyler et al. (2026) analyzed the 2020 American National Election Study (ANES); a probability-based survey with an impressive 36.7% response rate. Yet, Tyler et al. conclude, “We document evidence of nonignorable unit nonresponse; that is, Trump supporters were less likely to respond and less likely to complete the 2020 ANES if they did respond.” These concerns appear more severe with the 2024 ANES (Enns et al. 2025). Nonignorable nonresponse can create bias in even the most expensive probability-based surveys with the highest response rates.

The final two Iowa Polls of the 2024 presidential election, which missed the election outcome by 9 and 16 points, respectively, offer further evidence that probability-based surveys do not prevent extreme outlier results. Of course, nonresponse bias can affect all types of probability surveys—not just election surveys (see, e.g., Guyot et al. 2023; Kontto et al. 2025; Marello et al. 2026; Stein et al. 2025). Based on an extensive analysis, Wertheimer (2026) recently concluded the University of Michigan’s probability-based consumer sentiment survey is “broken.”

What About Studies that Compare Probability and Nonprobability Surveys?

Some people acknowledge the above concerns, but argue that these concerns are less prevalent in probability-based surveys than in other sampling approaches. These claims rely on outdated findings or—ironically—draw incorrect inferences by generalizing from an unrepresentative sample of surveys.

Bilgen and Dutwin (2026) recently wrote, “the overwhelming evidence shows that probability methods are consistently more accurate.” The most recent evidence they cite that directly speaks to the accuracy of probability and nonprobability samples is Cornesse et al.’s (2020) review article. This review article relied on studies using data from 1999-2000 (Berrens et al. 2003) to 2012 (MacInnis et al. 2018).4 The 2012 random digit dial (RDD) response rate in the MacInnis et al study, the most recent study in Cornesse et al. (2020), was over 15%—5 to 10 times the range of typical response rates today.5 What’s more, the nonprobability surveys they analyzed relied on outdated sampling methods (e.g., “natural traffic on the website”) and outdated incentives (e.g.,  “a chance to win prizes”) (MacInnis et al. 2018, 732). These were important studies, but we should not make inferences about contemporary sample type accuracy based on studies conducted between 1999 and 2012.

Mercer and Lau (2023) conducted a more recent comparison of three probability and three nonprobability samples from 2021. Across most of the 28 benchmarks in their study, the probability samples performed well (though they substantially overestimated voter turnout and the average absolute error of unweighted demographic variables was more than 4 times the expected margin of sampling error). Their study should not be used, however, to make generalizations about probability versus nonprobability samples. Two of the nonprobability samples they used rely on sample aggregators/marketplaces, which introduce data transparency concerns (Enns and Rothschild 2022). Further, all three nonprobability samples rely on simple quota-based sampling (education, race/ethnicity, and gender). For decades, we have known about the accuracy of more sophisticated nonprobability sampling approaches (e.g., Rivers 2007). We cannot generalize from three samples that do not follow best practices. Doing so ignores the heterogeneity across nonprobability sampling methods and ignores the nonprobability surveys that are as accurate or more accurate than probability surveys when compared to a range of population benchmarks, including election outcomes, number of children born, vaccination status, and disease status (e.g., Enns and Rothschild 2021; Enns et al. 2024, 2025; Santillan et al. 2024; Tamanna et al. 2025).

Solutions

In an important review of survey research for PNAS Nexus, Jamieson et al. (2023) wrote, “Probability samples minimize the risk of systematic bias.” Statements like this no longer reflect the survey landscape. Intuition suggests that the 2 to 3 people out of every 100 who take a probability-based survey are unlikely to be equivalent to the 97 to 98 who were invited but did not take the survey. The data support this intuition. As Dominitz and Manski (2026, 60) explain, “polling has sought to adhere to an ideal promoted in the statistical literature on survey research… However, the ideal is essentially never achieved in practice.”

As emphasized above, this does not mean that all probability-based surveys are problematic. Accurate and representative probability surveys exist and there can be compelling reasons for using probability-based sampling—including in our own work. Further, nonprobability samples must also address nonignorable nonresponse, as well as the fact that in nonprobability samples the probability of being included in the survey is zero for some members of the population. For this reason, we cannot assume that nonprobability samples represent the population on unmeasured characteristics. But significant nonignorable nonresponse leaves probability samples in the same position. In both cases, potential for systematic bias exists. We need additional information to assess survey data.

We have two recommendations for assessing survey samples. These recommendations apply equally to surveys using probability and nonprobability samples.

Always ask:

1.)   How is nonrandom nonresponse minimized?

Response rates may be part of the answer, but as we saw above, recent evidence of nonignorable nonresponse exists with response rates as high as 37%. More important than response rate, we need to know how various aspects of the respondent recruitment process, such as mode of outreach, number of contacts, invitation messaging, mode of survey, and incentives, help ensure that those who do not take the survey are equivalent to those who do.

2.)   How do prior surveys from that firm or survey organization using the same or similar sampling approaches compare to known population benchmarks?

Comparing survey results with administrative benchmarks can be an important tool for evaluating probability and nonprobability samples—as well as the two sample types in combination (Enns et al. 2025), but this is often difficult to do in practice. Following appropriate caveats (e.g., Jamieson et al. 2023; Shapiro 2020), election outcomes offer a highly visible benchmark for surveys. Unfortunately, many probability-based surveys have stopped asking vote intentions, eliminating this opportunity for independent assessments of survey accuracy. The high costs of probability-based surveys further complicate benchmark studies. Absent major funding, it is impossible for researchers to assess if recent probability-based surveys align with known population benchmarks if these organizations do not regularly publish this information.

Conclusions

High quality surveys continue to provide critical data across all sectors of society. Yet, users of survey data must remember that in the current era of polling, sampling methods alone are not enough to assess whether a survey accurately represents the population of interest. More generally, thinking in terms of probability or nonprobability samples may not be helpful. Many prominent surveys now effectively combine both sample types, showing that the probability/nonprobability distinction represents a false dichotomy. As Freese and Jin (2025, 111-112) explain, “The boundary between probability samples and nonprobability samples has become increasingly blurred with the deterioration of the ability to execute probability samples fully successfully in countries like the United States.” We believe survey research will be stronger as the focus turns away from outdated heuristics and instead assesses data quality and accuracy directly

Endnotes

1 Many probability-based surveys do not report their response rates. We greatly appreciate those that do. The response rate range is based on the following surveys: 2026 NYT/Siena, 2026 NYT/Siena, 2024 Pew Research Center, 2025 AP-NORC, 2025 Pew Research Center, and 2025 AP-NORC. The correct response rate for probability-based panels must incorporate the original response rate into the panel, the panel retention rate, and the specific survey response rate.

2 The field dates were 10/14-10/21/2020. Survey details available here: https://doi.org/10.25940/ROPER-31118042.

3 We performed a similar analysis of 2024 election polls (Enns et al. 2025). Probability samples were slightly (0.9 percentage points) more accurate than nonprobability samples, but this pattern reversed with nonprobabilty samples becoming more accurate (1 percentage point) once we controlled for data transparency by focusing on surveys that are members of the AAPOR Transparency Initiative and/or archive their data with the Roper Center.

4 MacInnis et al. (2018) is the most recently published study in Cornesse et al. (2020). Bilgen and Dutwin also focus on LLM fraud in surveys. More recent research demonstrates that this type of fraud in high quality nonprobability surveys is exceptionally rare (Gordon et al. 2026, Rothchild et al. 2026). Bilgen and Dutwin do not report any studies that have investigated potential LLM fraud in probability surveys, so we do not have empirical evidence on the prevalence of LLM fraud in probability surveys.

5 The probability-based internet response rate was 2%, right in the middle of today’s response rates.

References

Bailey, Michael A. 2024. Polling at a Crossroads: Rethinking Modern Survey Research. New York: Cambridge University Press.

Berrens, Robert P., Alok K. Bohara, Hank Jenkins-Smith, Carol Silva, and David L. Weimer. 2003. “The Advent of Internet Surveys for Political Research: A Comparison of Telephone and Internet Samples.” Political Analysis 11(1): 1-22.

Bilgen, Ipek and David Dutwin. 2026. “Using Nonprobability Survey Samples Can Be a Dangerous Gamble.” Research Brief. NORC at the University of Chicago (Feb. 6).

Cornesse, Carina, Annelies G. Blom, David Dutwin, Jon A. Krosnick, Edith D. de Leeuw, Stéphane Legleye, Josh Pasek, Darren Pennay, Benjamin Phillips, Joseph W. Sakshaug, Bella Struminskaya, and Alexander Wenz. 2020. “A Review of Conceptual Approaches and Empirical Evidence on Probability and Nonprobability Sample Survey Research.” Journal of Survey Statistics and Methodology 8(1): 4-36.

Dominitz, Jeff and Charles F. Manski. 2025. “Using Total Margin of Error to Account for Non-Sampling Error in Election Polls: The Case of Nonresponse.” Journal of the American Statistical Association. 121(553): 60-71.

Enns, Peter K. and Jake Rothschild. 2021. “Revisiting the ‘gold standard’ of polling: new methods outperformed traditional ones in 2020” 3Streams. (Mar 18).

Enns, Peter K., Amelia Goranson, Jake Rothschild, and Gretchen Streett. 2025. “The Future of Election Polls” 3Streams. (May 13).

Enns, Peter K., Colleen L. Barry, James N. Druckman, Sergio Garcia-Rios, David C. Wilson, and Jonathon P. Schuldt. 2024. “The Need for a Recurring Large-Scale Benchmarking Survey to Continually Evaluate Sampling Methods and Administration Modes: Lessons from the 2022 Collaborative Midterm Survey” ArXiv.

Freese, Jeremy and Olivia Jin. 2025. “Online Nonprobability Samples.” Annual Review of Sociology 51: 109-128.

Gordon, Andrew, David Rothschild, Felipe M. Affonso, Justin Sulik, David J. Hauser, Karine Pepin, and Simon Jones. 2026. “AI Agent Prevalence and Data Quality Across Multiple Online Sample Providers.” PsyArXiv.

Guyot, Madeleine, Ingrid Pelgrims, Raf Aerts, Hans Keune, Roy Remmen, Eva M. De Clercq, Isabelle Thomas, and Sophie O. Vanwambeke. 2023. “Non-Response Bias in the Analysis of the Association Between Mental Health and the Urban Environment: A Cross-Sectional Study in Brussels, Belgium.” Archives of Public Health 81:129.

Jamieson, Kathleen Hall, Arthur Lupia, Ashley Amaya, Henry E. Brady, René Bautista, Joshua D. Clinton, Jill A. Dever, David Dutwin, Daniel L. Goroff, D. Sunshine Hillygus, Courtney Kennedy, Gary Langer, John S Lapinski, Michael Link, Tasha Philpot, Ken Prewitt, Doug Rivers, Lynn Vavreck, David C. Wilson, Marcia K. McNutt. 2023. “Protecting the integrity of survey research” PNAS Nexus (2)3:1-10.

Kennedy, Courtney and Hannah Hartig. 2019. “Response rates in telephone surveys have resumed their decline.” Pew Research Center (Feb. 27).

Kontto, Jukka, Hanna Tolonen, and Anne H. Salonen. 2025. “Using Administrative Register Data for Adjusting Non-Response Bias in the Finnish Gambling Harms Survey.” BMC Public Health 25(1): 1807.

MacInnis, Bo, Jon A. Krosnick, Annabell S. Ho, and Mu-Jung Cho. 2018. “The Accuracy of Measurements with Probability and Nonprobability Survey Samples: Replication and Extension.” Public Opinion Quarterly 82(4): 707-744.

Marello, Madeline Marie, Christy Denckla, and Maja O’Connor. 2026. “Recruitment and Retainment of Bereaved Spouses: Non-Response to a Danish Registry-Sampled Longitudinal Survey on Mental Health After Loss.” Quality & Quantity 60(2): 6073-6089.

Mercer, Andrew, and Arnold Lau. 2023. “Comparing Two Types of Online Survey Samples.” Pew Research Center. (September 7).

Rivers, Douglas. 2007. “Sampling for Web Surveys.” Paper prepared for the 2007 Joint Statistical Meetings, Salt Lake City, UT, August 1, 2007.

Rothschild, David M., Soubhik Barari, Trent D. Buskirk, Andrew Gordon, and D. Sunshine Hillygus. 2026. “Reply to Westwood: Questioning the Empirical Evidence that AI Survey Contamination Is Real and Substantial.” SocArXiv.

Santillana, Mauricio, Ata A. Uslu, Tamanna Urmi, Alexi Quintana-Mathe, James N. Druckman, Katherine Ognyanova, Matthew Baum, Roy H. Perlis, and David Lazer. 2024. “Tracking COVID-19 Infections Using Survey Data on Rapid At-Home Tests.” JAMA Network Open 7(9): e2435442.

Shapiro, Robert Y. 2020. “Despite the 2020 election results, you can still trust polling. Mostly.” Washington Post. (Dec. 3).

Urmi, Tamanna, Binod Pant, George Dewey, Alexi Quintana-Mathe, Iris Lang, James Druckman, Katherine Ognyanova, Matthew Baum, Roy Perlis, Christoph Riedl, David Lazer, and Mauricio Santillana. 2025.  "Characterizing population-level changes in human behavior during the COVID-19 pandemic in the United States." PNAS 122(37): 1-12.

Tyler, Matthew, D. Sunshine Hillygus, Matthew DeBell, Ted Brader, Shanto Iyengar, Daron Shaw, and Nicholas A. Valentino. 2026. “Why are surveys struggling to estimate vote shares?” American Journal of Political Science. (early view).

Wertheimer, Joel. 2026. “Is the Vibecession Real - Or Is the Survey Broken?” Silver Bulletin. (June 29).

Author Affiliations

Peter K. Enns: Professor of Government, Professor of Public Policy, and Robert S. Harrison Director of the Cornell Center for Social Sciences, Cornell University and Chief Data Scientist, Verasight

Joey Marshall: Vice President of Data Science, Verasight. Prior to joining Verasight, Joey was a Data Scientist at the U.S. Census Bureau and a Research Associate at Pew Research Center.

Acknowledgements

We thank Jamie Druckman, Mujahed Islam, Kabir Khanna, Jake Rothschild, Jon Schuldt, Kathleen Weldon, and David Wilson for helpful comments on prior versions of this article.

Suggested Citation

Enns, Peter K. and Joey Marshall. 2026. “Probability-Based Samples Can No Longer Serve as the Heuristic for Accurate and Representative Surveys” Verasight Report. (July 21).

About Verasight

Founded by academic researchers, Verasight enables leading institutions to survey any audience of interest (e.g., engineers, doctors, policy influencers). From academic researchers and media organizations to Fortune 500 companies, Verasight is helping our client stay ahead of trends in their industry. Learn more about how Verasight can support your research. Contact us at contact@verasight.io.

Verasight Report

Probability-Based Samples Can No Longer Serve as the Heuristic for Accurate and Representative Surveys

Peter K. Enns and Joey Marshall

Published Jul 21, 2026

Updated

Probability-Based Samples Can No Longer Serve as the Heuristic for Accurate and Representative Surveys

We work for Verasight, a survey firm that conducts both probability and nonprobability sampling. We have also both relied extensively on probability surveys in other contexts—Enns in his academic research and Marshall in his prior work for the Pew Research Center and the U.S. Census Bureau. We believe the industry needs to revise how it thinks about probability surveys.

For decades, probability-based samples were the gold standard of survey research. For this reason, scholars often view probability samples as a signal of data accuracy. Unfortunately, this sample type is no longer an informative data quality heuristic. As Michael Bailey explains in Polling at a Crossroads, “Almost 70 years into the random sampling era, the paradigm has begun to show its age. Changes in society have made it harder and harder to implement surveys according to the tenets of the paradigm” (Bailey 2024, 46).

Of course, many probability-based surveys produce accurate and representative data. As do many surveys that use nonprobability methods. It is easy to find examples of both sample types that match known population benchmarks (e.g., Enns et al. 2024; Santillan et al. 2024; Tamanna et al. 2025).

We do not question the potential upside of various sample types. High-quality probability and nonprobability surveys remain an important and excellent source of data. Our point is simply that when deciding which surveys to use or analyze, the sample type should not be the determining factor. Knowing if a survey used a probability sample no longer provides a clear signal about whether the sample accurately represents the population of interest.

Quick Background (Some Key Terms)

In a random sample, each individual in the population surveyed has an equal chance of being selected into the sample. The random sample is a special type of probability-based sample. In a probability sample, everyone in the population has a known, nonzero probability of being sampled.

While many factors contribute to the overall accuracy and representativeness of surveys, probability (and thus random) samples will produce accurate estimates of the population—as well as calculable uncertainty around these estimates—when assumptions hold.

A key assumption of probability surveys is that those who take the survey are, on average, equivalent to those who were invited but did not take the survey. Those who are invited but do not take the survey reflect survey (or unit) nonresponse. If those not responding are missing at random, the nonresponse can be ignored.

However, if those who do not take the survey differ systematically from those who take the survey, we have nonignorable nonresponse. If the nonresponders differ on aspects that correlate with what we care about, our estimates will be biased and the survey will not represent the population being studied.

Poststratification weights do not solve the nonignorable nonresponse problem. These weights adjust the sample so proportions in each demographic group in the survey match the actual proportions in the population (often measured in a census or other large government survey). But this weighting also assumes nonignorable nonresponse on any variables or characteristics not included in the weights.

To see why, suppose we conduct a political survey and individuals with lower education levels are less likely to take the survey. Further, suppose that the lower education individuals who do take the survey tend to be much more politically interested than lower education individuals who did not take the survey. These response patterns will lead us to overestimate political interest among the lower education group.

What’s more, this overestimate of political interest will increase in the full sample when we weight the data by education. The increased bias results because weighting increases the influence of respondents with lower education levels, since they are underrepresented in the survey. But these lower education respondents are not representative of the lower education population as a whole. Weighting in this way will increase the influence of non-representative respondents (i.e., lower education respondents who are much more politically interested than those who did not take the survey).

Why We Cannot Assume the Core Assumption of Probability-Based Sampling Holds

In the 1990s, response rates to telephone surveys were 36% (Kennedy and Hartig 2019). When more than one-in-three individuals contacted took the survey, it seemed plausible that the nonresponders were not systematically different from those who took the survey. But response rates have plummeted since the 1990s.

Current response rates for the most rigorous probability-based samples range from about 1.6 to 3.4%.1 These response rates mean that out of every 100 people invited, only 2 or 3 typically take the survey. It seems extremely unlikely that the 2 to 3 who choose to take the survey are not systematically different from the 97 to 98 who do not take the survey.

Indeed, one of us experienced this concern directly after fielding an expensive probability-based survey just over two weeks before the 2020 U.S. presidential election.2 The first question in the survey asked respondents, “On a scale from zero to 10, please indicate how likely it is that you will vote in the presidential elections being held in November.” The second question asked, “Who do you plan to vote for in the 2020 election for president?”

The weighted results from the probability-based survey had Biden winning the two-party vote by 20.8 percentage points—an overestimate of his actual winning margin by more than 16 percentage points. If we limit the sample to those who indicated they had voted or were extremely likely to vote, the Biden vote share overestimate grew to more than 20 percentage points.

Of course, late-deciding voters and not knowing which respondents will actually vote can lead vote intentions in pre-election surveys to differ from the final election outcome (e.g., Jamieson et al. 2023; Shapiro 2020). But these factors cannot account for an overestimate of more than 16 to 20 percentage points just over two weeks prior to the election.

Evidence of nonignorable nonresponse in probability-based surveys extends beyond this single example. Enns and Rothschild (2021) analyzed 355 surveys of likely voters from the two months preceding the 2020 election and found that surveys with probability-based samples produced an average raw error of 6%—more than nonprobability samples and well beyond the expected margin of sampling error. Further, instead of error centered around the final vote, all 81 probability-based samples overestimated Biden’s vote share.3 We would not expect this pattern of results if nonresponse was ignorable in probability surveys.

While low single digit response rates suggest nonresponse cannot be ignored in probability-based surveys, nonignorable nonresponse also exists when response rates are much higher.

Tyler et al. (2026) analyzed the 2020 American National Election Study (ANES); a probability-based survey with an impressive 36.7% response rate. Yet, Tyler et al. conclude, “We document evidence of nonignorable unit nonresponse; that is, Trump supporters were less likely to respond and less likely to complete the 2020 ANES if they did respond.” These concerns appear more severe with the 2024 ANES (Enns et al. 2025). Nonignorable nonresponse can create bias in even the most expensive probability-based surveys with the highest response rates.

The final two Iowa Polls of the 2024 presidential election, which missed the election outcome by 9 and 16 points, respectively, offer further evidence that probability-based surveys do not prevent extreme outlier results. Of course, nonresponse bias can affect all types of probability surveys—not just election surveys (see, e.g., Guyot et al. 2023; Kontto et al. 2025; Marello et al. 2026; Stein et al. 2025). Based on an extensive analysis, Wertheimer (2026) recently concluded the University of Michigan’s probability-based consumer sentiment survey is “broken.”

What About Studies that Compare Probability and Nonprobability Surveys?

Some people acknowledge the above concerns, but argue that these concerns are less prevalent in probability-based surveys than in other sampling approaches. These claims rely on outdated findings or—ironically—draw incorrect inferences by generalizing from an unrepresentative sample of surveys.

Bilgen and Dutwin (2026) recently wrote, “the overwhelming evidence shows that probability methods are consistently more accurate.” The most recent evidence they cite that directly speaks to the accuracy of probability and nonprobability samples is Cornesse et al.’s (2020) review article. This review article relied on studies using data from 1999-2000 (Berrens et al. 2003) to 2012 (MacInnis et al. 2018).4 The 2012 random digit dial (RDD) response rate in the MacInnis et al study, the most recent study in Cornesse et al. (2020), was over 15%—5 to 10 times the range of typical response rates today.5 What’s more, the nonprobability surveys they analyzed relied on outdated sampling methods (e.g., “natural traffic on the website”) and outdated incentives (e.g.,  “a chance to win prizes”) (MacInnis et al. 2018, 732). These were important studies, but we should not make inferences about contemporary sample type accuracy based on studies conducted between 1999 and 2012.

Mercer and Lau (2023) conducted a more recent comparison of three probability and three nonprobability samples from 2021. Across most of the 28 benchmarks in their study, the probability samples performed well (though they substantially overestimated voter turnout and the average absolute error of unweighted demographic variables was more than 4 times the expected margin of sampling error). Their study should not be used, however, to make generalizations about probability versus nonprobability samples. Two of the nonprobability samples they used rely on sample aggregators/marketplaces, which introduce data transparency concerns (Enns and Rothschild 2022). Further, all three nonprobability samples rely on simple quota-based sampling (education, race/ethnicity, and gender). For decades, we have known about the accuracy of more sophisticated nonprobability sampling approaches (e.g., Rivers 2007). We cannot generalize from three samples that do not follow best practices. Doing so ignores the heterogeneity across nonprobability sampling methods and ignores the nonprobability surveys that are as accurate or more accurate than probability surveys when compared to a range of population benchmarks, including election outcomes, number of children born, vaccination status, and disease status (e.g., Enns and Rothschild 2021; Enns et al. 2024, 2025; Santillan et al. 2024; Tamanna et al. 2025).

Solutions

In an important review of survey research for PNAS Nexus, Jamieson et al. (2023) wrote, “Probability samples minimize the risk of systematic bias.” Statements like this no longer reflect the survey landscape. Intuition suggests that the 2 to 3 people out of every 100 who take a probability-based survey are unlikely to be equivalent to the 97 to 98 who were invited but did not take the survey. The data support this intuition. As Dominitz and Manski (2026, 60) explain, “polling has sought to adhere to an ideal promoted in the statistical literature on survey research… However, the ideal is essentially never achieved in practice.”

As emphasized above, this does not mean that all probability-based surveys are problematic. Accurate and representative probability surveys exist and there can be compelling reasons for using probability-based sampling—including in our own work. Further, nonprobability samples must also address nonignorable nonresponse, as well as the fact that in nonprobability samples the probability of being included in the survey is zero for some members of the population. For this reason, we cannot assume that nonprobability samples represent the population on unmeasured characteristics. But significant nonignorable nonresponse leaves probability samples in the same position. In both cases, potential for systematic bias exists. We need additional information to assess survey data.

We have two recommendations for assessing survey samples. These recommendations apply equally to surveys using probability and nonprobability samples.

Always ask:

1.)   How is nonrandom nonresponse minimized?

Response rates may be part of the answer, but as we saw above, recent evidence of nonignorable nonresponse exists with response rates as high as 37%. More important than response rate, we need to know how various aspects of the respondent recruitment process, such as mode of outreach, number of contacts, invitation messaging, mode of survey, and incentives, help ensure that those who do not take the survey are equivalent to those who do.

2.)   How do prior surveys from that firm or survey organization using the same or similar sampling approaches compare to known population benchmarks?

Comparing survey results with administrative benchmarks can be an important tool for evaluating probability and nonprobability samples—as well as the two sample types in combination (Enns et al. 2025), but this is often difficult to do in practice. Following appropriate caveats (e.g., Jamieson et al. 2023; Shapiro 2020), election outcomes offer a highly visible benchmark for surveys. Unfortunately, many probability-based surveys have stopped asking vote intentions, eliminating this opportunity for independent assessments of survey accuracy. The high costs of probability-based surveys further complicate benchmark studies. Absent major funding, it is impossible for researchers to assess if recent probability-based surveys align with known population benchmarks if these organizations do not regularly publish this information.

Conclusions

High quality surveys continue to provide critical data across all sectors of society. Yet, users of survey data must remember that in the current era of polling, sampling methods alone are not enough to assess whether a survey accurately represents the population of interest. More generally, thinking in terms of probability or nonprobability samples may not be helpful. Many prominent surveys now effectively combine both sample types, showing that the probability/nonprobability distinction represents a false dichotomy. As Freese and Jin (2025, 111-112) explain, “The boundary between probability samples and nonprobability samples has become increasingly blurred with the deterioration of the ability to execute probability samples fully successfully in countries like the United States.” We believe survey research will be stronger as the focus turns away from outdated heuristics and instead assesses data quality and accuracy directly

Endnotes

1 Many probability-based surveys do not report their response rates. We greatly appreciate those that do. The response rate range is based on the following surveys: 2026 NYT/Siena, 2026 NYT/Siena, 2024 Pew Research Center, 2025 AP-NORC, 2025 Pew Research Center, and 2025 AP-NORC. The correct response rate for probability-based panels must incorporate the original response rate into the panel, the panel retention rate, and the specific survey response rate.

2 The field dates were 10/14-10/21/2020. Survey details available here: https://doi.org/10.25940/ROPER-31118042.

3 We performed a similar analysis of 2024 election polls (Enns et al. 2025). Probability samples were slightly (0.9 percentage points) more accurate than nonprobability samples, but this pattern reversed with nonprobabilty samples becoming more accurate (1 percentage point) once we controlled for data transparency by focusing on surveys that are members of the AAPOR Transparency Initiative and/or archive their data with the Roper Center.

4 MacInnis et al. (2018) is the most recently published study in Cornesse et al. (2020). Bilgen and Dutwin also focus on LLM fraud in surveys. More recent research demonstrates that this type of fraud in high quality nonprobability surveys is exceptionally rare (Gordon et al. 2026, Rothchild et al. 2026). Bilgen and Dutwin do not report any studies that have investigated potential LLM fraud in probability surveys, so we do not have empirical evidence on the prevalence of LLM fraud in probability surveys.

5 The probability-based internet response rate was 2%, right in the middle of today’s response rates.

References

Bailey, Michael A. 2024. Polling at a Crossroads: Rethinking Modern Survey Research. New York: Cambridge University Press.

Berrens, Robert P., Alok K. Bohara, Hank Jenkins-Smith, Carol Silva, and David L. Weimer. 2003. “The Advent of Internet Surveys for Political Research: A Comparison of Telephone and Internet Samples.” Political Analysis 11(1): 1-22.

Bilgen, Ipek and David Dutwin. 2026. “Using Nonprobability Survey Samples Can Be a Dangerous Gamble.” Research Brief. NORC at the University of Chicago (Feb. 6).

Cornesse, Carina, Annelies G. Blom, David Dutwin, Jon A. Krosnick, Edith D. de Leeuw, Stéphane Legleye, Josh Pasek, Darren Pennay, Benjamin Phillips, Joseph W. Sakshaug, Bella Struminskaya, and Alexander Wenz. 2020. “A Review of Conceptual Approaches and Empirical Evidence on Probability and Nonprobability Sample Survey Research.” Journal of Survey Statistics and Methodology 8(1): 4-36.

Dominitz, Jeff and Charles F. Manski. 2025. “Using Total Margin of Error to Account for Non-Sampling Error in Election Polls: The Case of Nonresponse.” Journal of the American Statistical Association. 121(553): 60-71.

Enns, Peter K. and Jake Rothschild. 2021. “Revisiting the ‘gold standard’ of polling: new methods outperformed traditional ones in 2020” 3Streams. (Mar 18).

Enns, Peter K., Amelia Goranson, Jake Rothschild, and Gretchen Streett. 2025. “The Future of Election Polls” 3Streams. (May 13).

Enns, Peter K., Colleen L. Barry, James N. Druckman, Sergio Garcia-Rios, David C. Wilson, and Jonathon P. Schuldt. 2024. “The Need for a Recurring Large-Scale Benchmarking Survey to Continually Evaluate Sampling Methods and Administration Modes: Lessons from the 2022 Collaborative Midterm Survey” ArXiv.

Freese, Jeremy and Olivia Jin. 2025. “Online Nonprobability Samples.” Annual Review of Sociology 51: 109-128.

Gordon, Andrew, David Rothschild, Felipe M. Affonso, Justin Sulik, David J. Hauser, Karine Pepin, and Simon Jones. 2026. “AI Agent Prevalence and Data Quality Across Multiple Online Sample Providers.” PsyArXiv.

Guyot, Madeleine, Ingrid Pelgrims, Raf Aerts, Hans Keune, Roy Remmen, Eva M. De Clercq, Isabelle Thomas, and Sophie O. Vanwambeke. 2023. “Non-Response Bias in the Analysis of the Association Between Mental Health and the Urban Environment: A Cross-Sectional Study in Brussels, Belgium.” Archives of Public Health 81:129.

Jamieson, Kathleen Hall, Arthur Lupia, Ashley Amaya, Henry E. Brady, René Bautista, Joshua D. Clinton, Jill A. Dever, David Dutwin, Daniel L. Goroff, D. Sunshine Hillygus, Courtney Kennedy, Gary Langer, John S Lapinski, Michael Link, Tasha Philpot, Ken Prewitt, Doug Rivers, Lynn Vavreck, David C. Wilson, Marcia K. McNutt. 2023. “Protecting the integrity of survey research” PNAS Nexus (2)3:1-10.

Kennedy, Courtney and Hannah Hartig. 2019. “Response rates in telephone surveys have resumed their decline.” Pew Research Center (Feb. 27).

Kontto, Jukka, Hanna Tolonen, and Anne H. Salonen. 2025. “Using Administrative Register Data for Adjusting Non-Response Bias in the Finnish Gambling Harms Survey.” BMC Public Health 25(1): 1807.

MacInnis, Bo, Jon A. Krosnick, Annabell S. Ho, and Mu-Jung Cho. 2018. “The Accuracy of Measurements with Probability and Nonprobability Survey Samples: Replication and Extension.” Public Opinion Quarterly 82(4): 707-744.

Marello, Madeline Marie, Christy Denckla, and Maja O’Connor. 2026. “Recruitment and Retainment of Bereaved Spouses: Non-Response to a Danish Registry-Sampled Longitudinal Survey on Mental Health After Loss.” Quality & Quantity 60(2): 6073-6089.

Mercer, Andrew, and Arnold Lau. 2023. “Comparing Two Types of Online Survey Samples.” Pew Research Center. (September 7).

Rivers, Douglas. 2007. “Sampling for Web Surveys.” Paper prepared for the 2007 Joint Statistical Meetings, Salt Lake City, UT, August 1, 2007.

Rothschild, David M., Soubhik Barari, Trent D. Buskirk, Andrew Gordon, and D. Sunshine Hillygus. 2026. “Reply to Westwood: Questioning the Empirical Evidence that AI Survey Contamination Is Real and Substantial.” SocArXiv.

Santillana, Mauricio, Ata A. Uslu, Tamanna Urmi, Alexi Quintana-Mathe, James N. Druckman, Katherine Ognyanova, Matthew Baum, Roy H. Perlis, and David Lazer. 2024. “Tracking COVID-19 Infections Using Survey Data on Rapid At-Home Tests.” JAMA Network Open 7(9): e2435442.

Shapiro, Robert Y. 2020. “Despite the 2020 election results, you can still trust polling. Mostly.” Washington Post. (Dec. 3).

Urmi, Tamanna, Binod Pant, George Dewey, Alexi Quintana-Mathe, Iris Lang, James Druckman, Katherine Ognyanova, Matthew Baum, Roy Perlis, Christoph Riedl, David Lazer, and Mauricio Santillana. 2025.  "Characterizing population-level changes in human behavior during the COVID-19 pandemic in the United States." PNAS 122(37): 1-12.

Tyler, Matthew, D. Sunshine Hillygus, Matthew DeBell, Ted Brader, Shanto Iyengar, Daron Shaw, and Nicholas A. Valentino. 2026. “Why are surveys struggling to estimate vote shares?” American Journal of Political Science. (early view).

Wertheimer, Joel. 2026. “Is the Vibecession Real - Or Is the Survey Broken?” Silver Bulletin. (June 29).

Author Affiliations

Peter K. Enns: Professor of Government, Professor of Public Policy, and Robert S. Harrison Director of the Cornell Center for Social Sciences, Cornell University and Chief Data Scientist, Verasight

Joey Marshall: Vice President of Data Science, Verasight. Prior to joining Verasight, Joey was a Data Scientist at the U.S. Census Bureau and a Research Associate at Pew Research Center.

Acknowledgements

We thank Jamie Druckman, Mujahed Islam, Kabir Khanna, Jake Rothschild, Jon Schuldt, Kathleen Weldon, and David Wilson for helpful comments on prior versions of this article.

Suggested Citation

Enns, Peter K. and Joey Marshall. 2026. “Probability-Based Samples Can No Longer Serve as the Heuristic for Accurate and Representative Surveys” Verasight Report. (July 21).

About Verasight

Founded by academic researchers, Verasight enables leading institutions to survey any audience of interest (e.g., engineers, doctors, policy influencers). From academic researchers and media organizations to Fortune 500 companies, Verasight is helping our client stay ahead of trends in their industry. Learn more about how Verasight can support your research. Contact us at contact@verasight.io.