Epidemiologists spend a large amount of their working hours seeing graphs that the majority of the public will never see at a monitoring center that falls halfway between a typical university computer lab and something that belongs in a government building. Hospital occupancy rates and patient counts are too late. These researchers monitor signs, such as search query numbers in areas with historically inadequate disease reporting infrastructure, wastewater virus concentrations from cities across three nations, and genomic sequencing data indicating which viral mutations are increasing and in which direction. The graphs that move two months before anything appears in a clinic are the most important ones.
Over the past five years, the systems performing this function have significantly improved. Not in a smooth, linear manner, but rather in the erratic, occasionally unexpected way that any technology advances under performance pressure. A decade ago, sources that would have seemed incidental to epidemic forecasting are today used by the best prediction algorithms. During the COVID-19 pandemic, wastewater surveillance proved to be one of the most accurate leading indicators since it tracks viral RNA released by sick populations before the majority of them have ever sought medical attention. Clinical case numbers usually increased one to two weeks after wastewater concentrations in a particular metropolitan region increased dramatically. Public health systems require this latency, but they seldom have it.
The AI component is particularly noticeable in variant prediction. When medical professionals sequence enough patient samples to identify a distinct circulation, standard surveillance detects a new strain. By then, the variation has frequently already spread to several different areas. The predictive models operate in a different way; they continuously scan the whole landscape of circulating mutations for the combination of alterations that, according to historical data, may result in an advantage in immune escape or transmission. Months before a candidate variety of concern achieves the threshold for official WHO identification, some of the more advanced systems can identify it. The alternative is to wait for the clinical signal to arrive, but it is not a guaranty and the false positive rate is a serious issue.
The history of digital epidemiology, which uses search trends, social media activity, and movement data to identify illness, is more convoluted. Before it was deactivated, Google Flu Trends, an early and ambitious attempt to anticipate influenza surges using search data, famously overestimated flu activity for a number of years. Researchers who came after learned that search data is too noisy and too sensitive to media coverage to be a trustworthy stand-alone indication, not that the strategy was flawed. Before producing any form of alarm, the systems currently in use carefully assess and check digital signals against environmental and clinical data, treating them as one layer among several.
The dual-use feature that public health professionals are increasingly willing to openly articulate is what makes this moment intriguing—and a little disturbing to follow attentively. Theoretically, a more transmissible pathogen could be engineered using the same machine learning architecture that can identify which viral changes will spread most efficiently. Scientists don’t discount this theory. This is a current biosecurity policy issue that is being discussed in scholarly publications and in discussions with national security agencies across several nations. Most of the developers of these predictive tools have a keen understanding of the problem and are in favor of governance structures. There is now no satisfactory response to the question of whether those frameworks can stay up with the technology.

Unusual spillover events in Southeast Asia and sub-Saharan Africa, a cluster of mutations in a circulating influenza lineage that appears unfavorable in terms of vaccine match, and wastewater data from multiple urban centers showing elevated novel pathogen signatures that have not yet been fully characterized are all contributing factors to the current warning flags in the predictive systems—the signals that epidemiologists are closely monitoring. Each of them does not on its own represent a stated threat. When combined, they represent precisely the type of convergence that the systems are intended to detect at an early stage. To be honest, we’re still figuring out if that’s a real reason for increased readiness or just the statistical background noise of a world with far more sensitive detectors than it had before COVID.
