Respiratory disease outbreaks burden American healthcare systems with over one million hospitalizations annually, yet current surveillance systems lag 1–2 weeks behind real-time conditions, preventing timely intervention. We present a machine learning early warning system that combines Google search trends with traditional epidemiological data using ensemble voting algorithms to predict outbreak timing across multiple respiratory pathogens. Unlike prior digital surveillance systems focused on retrospective evaluation or single-pathogen settings, this work presents a unified, prospectively deployed early warning framework that detects both outbreak onsets and peaks across multiple respiratory pathogens at the state level in real time. The system applies anomaly detection and transfer learning to monitor syndromic influenza-like illnesses, and hospitalizations caused by respiratory syncytial virus or influenza, simultaneously, across all 50 states. During operational real-time deployment from August 2024 through the 2024–2025 season, the system detects 98.0% of outbreak onsets and 97.0% of outbreak peaks, with average lead times of approximately 5 and 2 weeks, respectively, and positive predictive values exceeding 82%. This framework transforms reactive public health responses into proactive epidemic preparedness by reducing historical timing uncertainty from 10–20 weeks to consistent 2–6 week prediction windows, providing a scalable approach for monitoring both seasonal outbreaks and emerging respiratory threats.
Epidemic models face a critical challenge: surveillance systems capture only a fraction of infections (often < 10%). We reveal two fundamental problems. First, when models ignore underdetection entirely—treating detected cases as complete—parameter errors exceed 1000% despite visually reasonable fits. Second, when models explicitly account for underdetection by including case detection ratios as unknown parameters, structural identifiability analysis proves transmission rates and detection ratios become mathematically confounded—rendering infinite epidemiologically distinct scenarios equally plausible from case data alone. Integrating even a single population-level seroprevalence measurement resolves both problems by independently constraining cumulative exposure. Through Bayesian inference on synthetic SIR data, we demonstrate that this approach reduces parameter uncertainty by orders of magnitude, enabling accurate inference of transmission dynamics, peak timing, and outbreak size under realistic noise. Our framework establishes serological surveillance integration as both a mathematical necessity and a strategic investment for pandemic preparedness.
Background Generative artificial intelligence (AI) use has been suggested to have adverse mental health consequences but a causal relationship has not been examined.
ObjectiveTo simulate a randomised controlled trial of AI use in a work, school or personal context by applying target trial emulation to multiple waves of data from a nationally representative survey. MethodsWe conducted a target trial emulation using non-probability survey data from three waves of a nationally representative survey conducted between 18 June 2024 and 8 January 2025. Participants aged ≥18 years reported generative AI use frequency at baseline. High-frequency use was defined as multiple times per week or more. The primary outcome was depressive symptom severity measured using the Patient Health Questionnaire 9-item (PHQ-9) at follow-up. Generalised causal forests assessed heterogeneity of treatment effects. FindingsAmong 19 099 participants assessed at baseline, 2862 (15.0%) reported AI use at least multiple times per week. A subset of 3109 (16.3%) returned for follow-up. In the primary weighted analysis, high-frequency use was not significantly associated with change in PHQ-9 score at follow-up (mean difference −0.18, 95% CI −0.94 to 0.59; p=0.65). Multiple sensitivity analyses using alternate outcome definitions also did not identify significant causal effects. Generalised causal forests yielded no significant evidence of heterogeneity of effect (p=0.81). ConclusionsIn an emulated randomised trial among US adults, generative AI use was not associated with subsequent depressive symptoms. This result does not support the premise that AI use causes greater depressive symptoms, although adverse outcomes among vulnerable individuals cannot be excluded. Clinical implicationsAI use is unlikely to cause increased depressive symptoms among most US adults. Continued monitoring should clarify potential risks among vulnerable populations.
Introduction Long-acting monoclonal antibodies for respiratory syncytial virus (RSV) prevention, including nirsevimab and the recently approved clesrovimab, are recommended for all infants less than 8 months of age entering their first RSV season, with administration advised from October through March in most of the continental United States.1,2 This recommendation is based on historical RSV seasonality data and operational simplicity;3 however, the granular data available in the post-COVID era shows the epidemiology of RSV is highly variable.4 While current guidance acknowledges flexibility based on local epidemiology,1 state Medicaid and Vaccines for Children programs, which cover most eligible infants, generally adhere to this fixed administration window. However, the fraction of RSV burden falling outside this window has not been quantified using contemporary national surveillance data. Therefore, we analyzed national syndromic and hospital surveillance data to quantify RSV activity outside the October-March window and characterize geographic variability.
Vaccine strain selection for seasonal influenza A(H3N2) depends on knowing which hemagglutinin (HA) substitutions are most likely to erode neutralizing antibody recognition, yet published antigenic site sets disagree substantially on which positions matter most. We applied interpretable gradient-boosted tree models with SHAP-based site attribution to two complementary hemagglutination inhibition (HI) datasets to produce a more consolidated ranking of candidate antigenic positions. Models trained on a Neher/Bedford benchmark dataset recover the canonical cluster-transition sites established by prior analyses. Moreover, after filtering the WIC dataset for confounding factors, our models recover the majority of positions from four major prior reference sets (Koel, Neher/Bedford, Harvey, and Shah) and improve concordance between rankings derived from the Neher/Bedford and WIC datasets. Rankings from our models also agree more strongly with models trained to predict sampling time or passage identity than with standard evolutionary metrics used to detect diversifying selection. Our results show that interpretable sequence-based models can provide a more integrative ranking of candidate antigenic positions across different data sources and modeling approaches. This work should aid efforts to prioritize H3N2 substitutions for epidemic surveillance.
The effective reproduction number,R, is a predominant statistic for tracking infectious disease spread and informing health policies. An estimatedR = 1is universally interpreted as a stability threshold distinguishing epidemic growth (R > 1) from control (R < 1). We demonstrate that this interpretation frequently fails becauseRtypically averages over groups with heterogeneous characteristics. We find thatR = 1conceals valuable early-warning signals of resurgence and misclassifies complex dynamics as noise, generating false positive stability thresholds that diminish predictive and policymaking value. We further illustrate that a popular alternative transmissibility definition (using next-generation matrices) overcorrects this issue, producing false negative stability signals by amplifying stochastic variation. We address these limitations by adapting a recently developed statistic,E, derived fromRusing experimental design theory. We show thatEtightly constrains the set of scenarios consistent with stability, while remaining robust to noise and establishE = 1as a more practical and meaningful real-time threshold.
Machine learning models are increasingly used in clinical research to predict patient outcomes, yet many clinicians lack the training to critically appraise these studies. This article provides a conceptual introduction to machine learning for the pediatric hospitalist with no prior computational experience. We focus on the most common application in clinical medicine: supervised learning, where models learn from data with known outcomes to make predictions about new unseen patients. Core tasks such as classification and regression are explained, along with intuitive models like decision trees and advanced methods like ensembles. Essential concepts for critical appraisal, including overfitting and leakage, the challenge of interpretability, and data bias, are highlighted. We emphasize the importance of model validation and the distinction among prediction, interpretation, and causation. The article concludes by deconstructing a published pediatric study to illustrate these principles in practice, equipping the reader to better understand and evaluate research that uses machine learning. Our goal is to equip pediatric hospitalists with the foundational knowledge to become informed consumers and potential contributors within the machine learning ecosystem, ensuring that this technology augments, rather than replaces, clinical judgment.
Google Trends reports how frequently specific queries are searched on Google over time. It is widely used in research and industry to gain early insights into public interest. However, its data generation mechanism introduces missing values, sampling variability, noise, and trends. These issues arise from privacy thresholds mapping low search volumes to zeros, daily sampling variations causing discrepancies across historical downloads, and algorithm updates altering volume magnitudes over time. Data quality has recently deteriorated, with more zeros and noise, even for previously stable queries. We propose a comprehensive statistical methodology to preprocess Google Trends search information using hierarchical clustering, smoothing splines, and detrending. We validate our approach by forecasting U.S. influenza hospitalizations up to three weeks ahead with several statistical and machine learning models. Compared to omitting exogenous variables, our results show that preprocessed signals enhance forecast accuracy, while raw Google Trends data often degrades performance in statistical models.
American immigration attitudes Matthew A. Baum, Hong Qu, Benjamin Shair, David Lazer, Katherine Ognyanova, Roy H. Perlis, Mauricio Santillana CHIP50. March 29, 2026
Abstract
Key Takeaways • Partisan polarization dominates immigration attitudes, with a 67-percentage-point gap between Republicans (78%) and Democrats (11%) on Trump's immigration approval, far exceeding divisions along any other demographic dimension. • Personal stakes vary dramatically by race: 42% of Hispanic Americans worry about family or friend deportation compared to 20% of White Americans, reflecting differential proximity to enforcement consequences. • Personal connections to undocumented immigrants differ sharply by race: 31% of Hispanic Americans personally know someone undocumented versus 15% of White Americans. • Gender differences are systematic across enforcement questions, with men showing 14 percentage points higher Trump immigration approval than women. • Birthright citizenship retains 59% national support with 79% Democrat and 39% Republican support, representing one of few immigration questions with substantial cross-partisan agreement. • Education shows systematic patterns: those with some high school or less show 30% Trump immigration approval versus 43% among graduate degree holders, a 13-point educational gradient, with immigration approval rising with education level. • ICE enforcement approval shows a 60-percentage-point partisan gap (69% Republican, 9% Democrat), paralleling divisions on Trump's overall immigration handling. • Rural residents show 43% Trump immigration approval compared to 34% in urban areas, an 8-point geographic divide. • Geographic patterns tend to mirror presidential voting patterns, with states that vote reliably Republican tending to be most supportive of aggressive immigration tactics and policies by the Trump administration. • State variation on Trump immigration approval varies across states by as much as 22 percentage points (from a low of 28% in Hawaii to a high of 50% in Idaho).
The global oil tanker shipping network emerges from individual ship and fleet decisions driven by economic, environmental, and operational factors. However, most existing shipping network analysis rely on static, time-aggregated representations, overlooking critical temporality connecting individual vessel routing strategies with both operational efficiency and global cargo flows. To address this gap, we introduce a dual-scale framework complementing sequential motif analysis—capturing recurrent patterns in vessel movement sequences—with Dynamic Mode Decomposition (DMD), extracting temporal dynamics from vessel trajectories to global cargo flows. Using tanker movement data across four vessel classes, we demonstrate that vessels exhibiting diverse regional exploration patterns spend up to 50% more time transporting rather than seeking cargo, indicating greater economic and environmental efficiency. At the system scale, DMD analysis reveals distinct seasonality with an average peak-to-trough amplitude of 16%. Major import regions show synchronous annual demand cycles, while export regions exhibit anti-synchronicity. These temporal patterns, invisible to static analysis, reveal performance differences that enable route optimization for both economic and environmental benefits.
ImportanceGenerative artificial intelligence (AI) has rapidly entered mainstream use in the US, but its association with mental health has not been characterized. ObjectiveTo examine the associations of the extent and type of generative AI use among US adults with negative affective symptoms in a large, nationally representative sample. Design, Setting, and ParticipantsThis survey study used data from a 50-state US internet nonprobability survey conducted between April and May 2025. Survey respondents were aged 18 years and older. Data were analyzed in August 2025. ExposureParticipants self-reported generative AI and social media use. Main Outcomes and MeasuresThe outcome of interest, negative affect, was measured using the Patient Health Questionnaire 9-item (PHQ-9). ResultsThere were 20 847 unique participants, with mean (SD) age 47.3 (17.1) years and 10 327 (49.5%) female, 10 386 (49.8%) male, and 134 (0.6%) nonbinary participants; 2152 participants (10.3%) reported using AI at least daily, including 1053 participants (5.1%) who reported daily use and 1099 participants (5.3%) who reported use multiple times per day. Among participants who used daily or more frequently, 1033 (48.0%) reported use for work, 246 (11.4%) for school, and 1875 (87.1%) for personal applications. In survey-weighted regression models, daily or more frequent AI use was significantly more common among men, younger adults, those with higher education and income, and those in urban settings. Greater AI use was associated with greater levels of depressive symptoms in sociodemographic-adjusted regression models: (daily use: β = 1.08 [95% CI, 0.55-1.62]; multiple times per day: β = 0.86 [95% CI, 0.35-1.37]) compared with nonuse, and with greater likelihood of reporting at least moderate depressive symptoms (odds ratio [OR], 1.29 [95% CI, 1.15-1.46]); similar patterns were observed for anxiety and irritability. The highest estimates were observed among individuals using AI for personal use (β = 0.31 [95% CI, 0.10-0.52]) and those aged 25 to 44 years (β = 1.22 [95% CI, 0.70-1.74]) or 45 to 64 years (β = 1.38 [95% CI, 0.72-2.05]). Conclusions and Relevance This survey study found that AI use was significantly associated with greater depressive symptoms, with magnitude of differences varying by age group. Further work is needed to understand whether these associations are causal and explain heterogeneous effects.
BackgroundAntidepressants are among the most prescribed medications in the USA, yet challenges in access to mental health treatment persist. MethodsWe conducted a cross-sectional survey study using data from a national non-probability internet-based panel weighted to approximate national demographics (age, gender, race and ethnicity, education, US census region, and urbanicity) based on 2020 US Census data. Data were collected between 10 April and 27 May 2025 from 30 810 adults residing in the USA. The primary outcomes were self-reported current and past antidepressant and psychotherapy use, and support for or opposition to potential federal restrictions on antidepressant prescribing. Logistic regression models estimated demographic and treatment-related features associated with these outcomes. FindingsAmong 30 115 respondents with complete antidepressant data, 16.6% reported current antidepressant use, and of 30 098 respondents with psychotherapy data, 10.4% reported current psychotherapy. Use of both treatments was significantly greater among White respondents compared with all other racial groups. When asked about potential federal restrictions on doctors prescribing antidepressants, 16.4% of respondents supported and 48.0% opposed such regulation, with lesser opposition among those of male gender (OR 0.69, 95% CI 0.65 to 0.73), and greater opposition among those with lifetime antidepressant treatment (OR 2.37, 95% CI 2.21 to 2.54). ConclusionsAntidepressant and psychotherapy use remains unevenly distributed across demographic groups. A significant proportion of adults in every US state oppose efforts to restrict access to antidepressant prescribing, reflecting broad public support for maintaining access to treatment. Clinical ImplicationsFindings from this study suggest that restrictive policies on antidepressant prescribing are unlikely to align with public sentiment and may risk exacerbating existing inequities in care.
Large and persistent sociodemographic disparities in rates of mental health treatment in the United States have been reported, but whether these differences reflect institutional mistrust or limited social support remains unclear. This study described current treatment use among American adults with moderate-to-severe depressive or anxiety symptoms and examined whether trust in health care institutions and availability of emotional support were associated with lack of treatment. A cross-sectional analysis was conducted using data from a nationally distributed, web-based opinion survey of 9733 American adults with moderate-to-severe depressive or anxiety symptoms (Patient Health Questionnaire-9 score ≥10 and/or Generalized Anxiety Disorder-2 score ≥3). The survey was fielded April 10th–28th, 2025, using quota sampling for age, gender, race, ethnicity, education, U.S. census region, and urbanicity; post-stratification weights approximated the U.S. adult population. The primary outcome was no current mental health treatment (neither antidepressant nor psychotherapy use). Weighted logistic regression estimated odds ratios for treatment absence by sociodemographic characteristics, trust in physicians and hospitals, scientists and researchers, the Centers for Disease Control and Prevention, pharmaceutical companies, and emotional support. Among 9733 adults with elevated symptoms, 66.3 % reported no current treatment. Racial and ethnic minority groups, men, and those born outside the United States had higher odds of being untreated, while public insurance predicted lower odds. Lower trust in doctors and hospitals, lower trust in science, and lack of emotional support each independently predicted treatment absence, but inclusion of these variables did not meaningfully attenuate sociodemographic disparities.
Background Influenza and respiratory syncytial virus (RSV) are major contributors to the burden of seasonal influenza-like illnesses (ILI) in the US. The prevention and treatment of ILI varies substantially across age groups and in cost and administration schedule. Clearly identifying the times when healthcare resources are most needed to mitigate the effects of seasonal RSV and influenza outbreaks will improve public health responses before and during ILI seasons. Methods We implemented stacked-regression linear models to infer the contribution of each of these diseases to seasonal ILI syndromic indicators. We further implemented anomaly-detection algorithms on data from the US Centers for Disease Control and Prevention National Syndromic Surveillance Program to identify the timing of onsets and peaks of RSV, influenza, and COVID-19. Findings A total of 148 state-ILI seasons were analyzed. In 114 out of 148 (77.0%) of analyzed seasons, volume of RSV emergency department (ED) visits peaked before influenza ED visits. The median time difference between peaks of RSV and peaks of influenza was +3.0 weeks. The timing of RSV and influenza onsets were found to occur more synchronously in the 2023-2024 and 2024-2025 ILI seasons. Interpretations RSV epidemics frequently reach peak volume before influenza epidemics across the US. Healthcare professionals and public health authorities should anticipate increases in RSV cases and hospitalizations at the start of the annual ILI season and establish infrastructure and planning to handle incoming surges of both RSV and influenza appropriately. Funding No specific funding was provided for this study. Summary Epidemics of respiratory pathogens such as influenza or RSV drive the influenza-like illness season in the US. We show that RSV epidemics peak before influenza epidemics in most states, with about a one to three week difference separating the epidemics.