TruaceTracing the truth around AITuesday, August 25, 2026
TRV-2026-0864Version 1 · Certified

Written 2026-08-24 06:06:13 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0864
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-08-24T06:06:13.685603Z
status: published
lens: g_space
sector: health
headline: Identifying risk factors for marijuana use among male high school students using machine learning: Implications for public health
dek: Objectives To identify risk and protective factors associated with lifetime marijuana use among male high school students through an interpretable machine learning model, providing evidence to support early and targeted public health interventions. Study design Cross-sectional analysis of 2023 Youth Risk Behavior Surveillance System (YRBS) data for boys in grades 9-12 across the United States. Methods The final analytical sample included 8285 boys after excluding missing outcomes. Thirty-six predictors spanning…
gain_title: An interpretable logistic regression model trained on YRBS data predicted lifetime marijuana use among male high school students with high accuracy, enabling early identification of at-risk boys for targeted school and community prevention.
problem_title: (none)
trace_subject: (none)
gain_reading: An interpretable logistic regression model trained on YRBS data predicted lifetime marijuana use among male high school students with high accuracy, enabling early identification of at-risk boys for targeted school and community prevention.
gain_evidence: The optimized Logistic Regression model achieved strong performance (AUC = 0.9034; Accuracy = 0.8582; F1 = 0.8928) | providing evidence to support early and targeted public health interventions | This study applies an interpretable machine learning approach, grounded in systematic algorithm benchmarking, that provides actionable insights to identify high school boys with patterns associated with marijuana use, supporting early, focused prevention in schools and communities
problem_reading: (none)
problem_evidence: (none)
quick_read: A cross-sectional study of 8285 male high school students from the 2023 Youth Risk Behavior Surveillance System used 36 demographic, behavioral, and psychosocial predictors to model lifetime marijuana use, reported by 28.4% of participants. After benchmarking 17 algorithms, an optimized logistic regression model with SHAP and LIME explanations achieved AUC 0.9034 and accuracy 0.8582.

The work matters because it translates interpretable machine learning into actionable public health signals, highlighting co-occurring substance use, social media use, age, grade, bullying and sexual violence exposure, and academic achievement as key factors. As of the August 2026 publication date, results remain associational from a single US survey of boys, leaving uncertainty about causality, generalizability to other groups, and effectiveness of subsequent interventions.
limitation: Findings are based on a cross-sectional survey of only male high school students in the US with missing outcomes excluded, limiting causal inference and generalizability beyond this population.
tag: Evidence-backed gain
key_points: Analysis used 2023 Youth Risk Behavior Surveillance System data for 8285 boys in grades 9-12 after excluding missing outcomes. | Thirty-six predictors spanning demographics, substance use, lifestyle, psychosocial stressors, mental health, and household environment were retained. | Seventeen algorithms were benchmarked; Logistic Regression was selected for predictive performance and interpretability, with SHAP and LIME for explanation. | Lifetime marijuana use was reported by 28.4% of participants; top predictors included electronic vapor product use, cigarette smoking, alcohol consumption, and frequent social media use.
rundown: Researchers analyzed 2023 YRBS data for boys in grades 9-12, retaining 36 predictors after optimizing missing data thresholds and determining ensemble feature importance using Logistic Regression, Linear Discriminant Analysis, and Extra Trees classifiers.

The optimized logistic regression model was well calibrated with Brier score = 0.104 and expected-to-observed ratio = 1.01, and SHAP and LIME analyses confirmed robustness; high academic achievement was found to be protective while bullying, sexual violence exposure, and other substance use were predictive.
sources:
- peer_reviewed | Public Health | https://doi.org/10.1016/j.puhe.2026.106453 | 2026-08-22
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
81e844a457f98faaa11c5443e35201d7d87c5abdbf1f4de1c4ad0abfb2f23db4
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0864 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.