Section 7 of 10
Analyses
Maaike Verhagen, Desi Beckers, Nina van den Broek, Kirsten J. M. van Hooijdonk, Suhaavi Kochhar, Laila Qodariah, Milagros Rubio, Eveline Sarintohe, and Jacqueline M. Vink · about 2 minutes
All analyses were performed in R (R Core Team, 2023; R codes are available via OSF: https://osf.io/4sk5h/). For all samples, descriptive statistics of all study measures were explored.
Adolescent Samples (Samples 1, 2, and 3)
To examine whether loneliness predicts smoking initiation, we conducted discrete-time survival analyses using logistic regression models. The datasets were transformed into person-period files, where each participant had a separate row representing a time period. Once a participant initiated smoking, the following time points of data for that participant were removed (i.e. censoring). The models were performed with the glm function in R, using the complementary log–log link function. These models contained smoking initiation as the dependent variable, and the following independent variables: time (categorical; 5 time points), loneliness (at t-1; continuous, time-varying), sex (categorical, time-invariant), educational level (categorical, time-invariant, not available in sample 1 (KLS)), and age at baseline (continuous; time-invariant). We used sum-to-zero contrasts for categorical variables (i.e. sex and educational level), with − 1 reflecting boys and the lowest category of educational level. The threshold for statistical significance was set to α = 0.05.
University Student Samples (Samples 4 and 5)
Structural equation modelling (SEM) was used to explore the longitudinal association between baseline (T1) loneliness and smoking status (non-smokers, occasional smokers, or regular smokers) after 6 months (T2—short term) and 18 months (T3—long term). The R package ‘lavaan’ was used (Rosseel, 2011). In the model, loneliness at T1 was added as the independent variable, smoking status at T2 and T3 were added as (ordered) dependent variables, and smoking status at T1, gender, age at T1, and educational level (al-RISCO only) were added as covariates. For HSL, dummy coding was used for gender, while for al-RISCO, the category gender other was disregarded due to small numbers. Before conducting the analyses, missing data were imputed using KNN imputation (Memon et al., 2023). In the model, the weighted least squares mean and variance adjusted (WLSMV) estimator was used to estimate the model parameters (Suh, 2015) and the dependent variables were defined as ordered (lavaan project, 2023; Rosseel, 2011). The threshold for statistical significance was set to α = 0.05. In both samples, sensitivity analyses (with complete cases/smoking quantity (only for the HSL sample)) and attrition analyses were performed (see Supplementary Materials IV and V).
Exploratory analyses (deviation from pre-registration): In the main analyses, smoking status (non-smoker, occasional smoker, regular smoker) was included as an ordered variable (resulting in one beta). To better understand possible contrasts between the different smoking status groups, we have complemented the main analyses in both student samples with two nominal regression analyses, one with the dependent variable smoking status at T2 and one with the dependent variable smoking status at T3. The independent variables were the same as in the main analyses (gender, educational level (al-RISCO only), smoking status at T1, and loneliness at T). With these exploratory analyses, two different betas for occasional and regular smokers (compared to the reference group of non-smokers) were estimated.