Work overview

Section 02 of 09

Methods

Testing Front-of-Package Labels on Packaged Food for Adult and Adolescent Population in Bangladesh

Abu Ahmed Shamim, Lindsey Smith Taillie, Oumma Halima, Nisarga Bahar, Md Hafizul Islam, Md Mokbul Hossain, Sneha Sarwar, Ahmed K Abrar, Ummay Afroza, Sohel Reza Choudhury, Lindsay Steele, and Nazma Shaheen · 2026

Contents

Section 02 of 09

  1. 01Introduction
  2. 02Methods
  3. 03Results
  4. 04Discussion
  5. 05Author contributions
  6. 06Data availability
  7. 07Funding
  8. 08Declaration of generative AI and AI-assisted technologies in the writing process
  9. 09Conflict of interest
Text size
Work overview

Section 2 of 9

Methods

Abu Ahmed Shamim, Lindsey Smith Taillie, Oumma Halima, Nisarga Bahar, Md Hafizul Islam, Md Mokbul Hossain, Sneha Sarwar, Ahmed K Abrar, Ummay Afroza, Sohel Reza Choudhury, Lindsay Steele, and Nazma Shaheen · about 13 minutes

Institutional review board approvals

This study was reviewed and approved by the Ethics Review Committee of the National Heart Foundation Hospital and Research Institute in Dhaka, Bangladesh (Ref: N.H.F.H. & amp; R.I/4/14/7/Ad-2783) on 29 December 2024. Before starting the interview, the participants were given a short brief of the study and its objectives. Written informed consent was obtained from all adult participants. For adolescents (aged 10–18 y), written assent was obtained from the adolescents themselves, and consent was simultaneously obtained from their legal guardians.

This study was preregistered at the Open Science Framework on February 5, 2025. Registration DOI: https://doi.org/10.17605/OSF.IO/8ZDK6

Study design

The present study employed a multiarm, individually randomized experimental design to evaluate the impact of 4 different FOPL formats on consumer understanding, perceptions, and purchase intentions of packaged foods. Participants were individually randomly assigned to one of 5 study arms: barcode (control), HSR, WL, GDA, or MTL.

Study population

The study population comprised adults and adolescents residing in rural and urban Bangladesh. Adults were eligible if they were aged 19 to 60 y and had purchased packaged foods at least once in the past month and were involved in ≥50% of the household grocery purchasing decisions. Eligible adolescents aged 10 to 18 y had also purchased packaged foods at least once in the past month.

Exclusion criteria for both groups included vision impairment. To ensure representativeness, the proportion of primary incomplete participants (0–4 y of education) was capped at 25% of the sample; the final sample included ∼17% such participants; of them, <1% adolescents and <8% adults did not complete any education year. We included them because they are buyers and consumers of processed foods and because it is important for any labeling policy that is developed to work for illiterate populations as well as literate populations. We hypothesized that interpretive FOPLs (WL and MTL) would be appropriate for such participants. This hypothesis was based on the pictorial and color-coded nature of these formats, which are easier to interpret without requiring advanced literacy skills, thereby making them more appropriate for individuals with limited formal education.

Sample size and sampling strategy

Sample size estimation drew on evidence from a similar Indian trial [16]. Using a 5% type I error rate, 20% type II error rate (80% power), and a design effect of 1.61 [5,19], the required sample size was 600 participants per arm, totaling 3000 per age group.

A 4-stage cluster sampling design was used to obtain a sample across all 8 administrative divisions of Bangladesh (Rajshahi, Rangpur, Barishal, Khulna, Chattogram, Sylhet, Dhaka, and Mymensingh) (Supplemental Figure 1). From each division, one district was randomly selected (Naogaon, Gaibandha, Barishal, Kushtia, Lakshmipur, Moulvibazar, Narayanganj, and Mymensingh). Study locations were selected based on the primary sampling units of the Bangladesh Household Income and Expenditure Survey, ensuring consistency with the national census framework and preserving representativeness across rural and urban settings, ultimately resulting in 45 study site locations (27 rural and 18 urban). Data collection took place between February 28, 2025 and June 2, 2025.

In each study cluster, 250 to 280 eligible individuals (50% adults and 50% adolescents) were listed through household visits. From this list, 68 adults and an equal number of adolescents were randomly selected per cluster, with equal representation of males and females in both age groups. All randomly assigned participants were included in the interview process. In case of nonresponses mainly due to unavailability of the respondent, the study team attempted at least 3 times to meet the participant. The nonresponse or refusal rate was nearly 5% (4–6 nonresponses per cluster). Whenever there was a refusal or nonresponse, the enumerator informed the statistician, and the statistician replaced this from the list of additional potential participants. To minimize allocation bias, participants were randomly assigned to one of 4 FOPL formats or a control condition, ensuring equal distribution across study arms.

Stimuli

The FOPL formats were displayed on 5 commonly consumed products in Bangladesh [4]: cake, salty biscuit, chips, chanachur (salty snack), and sweet biscuit. A graphic designer developed mock products, and a team of nutritionists created mock nutrient profiles (i.e., calories, sugar, saturated fat, and sodium) for each of the 5 products based on a top-selling commercial brand for each category (Figure 1) [16].

FIGURE 1: Mock products of 5 commonly consumed packaged foods selected based on prior surveys.

FIGURE 1: Mock products of 5 commonly consumed packaged foods selected based on prior surveys.

Labels

Supplemental Table 1 compares 4 FOPL systems (MTL, Summary Indicators, Reference Intake/GDA, and WL) in terms of their design and interpretability. It shows that WL provide the clearest, nutrient-specific information, whereas 3 other formats vary in clarity and may present mixed messages, rely on numerical data, or be more susceptible to industry manipulation.

The FOPL formats tested in this study were selected based on qualitative research conducted preceding this randomized controlled trial [15] and were finalized by the Technical Working Group (Figure 2). The study used Bangla language versions of the FOPLs, which are presented in Supplemental Figure 2. Four FOPLs—the WL, MTL, HSR, and the GDA [16]—were selected in addition to the barcode control. A barcode label was used as a control label because it serves as a visual information on the front of the food package while conveying neutral information about the product’s nutritional content. A professional designer developed 5 sets of each mock product containing WL, MTL, HSR, GDA, and a barcode.

FIGURE 2: Five types of front-of-package labels used in the study (Control, MTL, WL, HSR, GDA) (English translations of the Bangla versions used in the study). GDA, guideline daily allowance; HSR, health star rating; MTL, multiple traffic light; WL, warning label.

FIGURE 2: Five types of front-of-package labels used in the study (Control, MTL, WL, HSR, GDA) (English translations of the Bangla versions used in the study). GDA, guideline daily allowance; HSR, health star rating; MTL, multiple traffic light; WL, warning label.

Pilot testing and modification of questionnaire

A pilot test, conducted with 964 adult and adolescent participants across rural and urban sites, assessed the clarity of mock food products, FOPLs, study procedures, and questionnaire items. Insights from this phase led to refinements in the questionnaire, including adding an “I do not know” option for nutrient-identification questions to avoid forcing uninformed responses. Additionally, assets associated with financial status were revised after participants showed confusion with the term “bills,” prompting a shift to more contextually relevant descriptions of household economic conditions.

Procedure

Participants were randomly assigned to 1 of 5 arms: control label, HSR, WL, GDA, or MTL label using an allocation ratio of 1:1:1:1:1. Participants then viewed images of products in random order with an FOPL on the product according to the assigned arm (as described above). In the control condition, all products had barcode labels. In the HSR condition, all products displayed stars. In the WL condition, products are shown with the relevant warning(s) for sugar, sodium, and/or saturated fat. In the GDA condition, all products had a GDA with the appropriate nutritional information. In the MTL condition, products were shown with MTLs with the applicable color code (green, yellow, or red) for each sugar, saturated fat, and salt.

Interviewers showed participants images of products using an A5-size booklet in random order. They asked them to assess the product based on the nutrients of concern and their reactions to the label. The participants viewed pictures of FOPL formats and answered questions about which label they preferred. All data were entered into a smartphone app (KoboCollect Software) during the interview.

Outcome measures

Primary outcomes

The primary outcomes of the study were as follows: 1) participants’ ability to correctly identify HFSS products and 2) intentions to purchase HFSS products. Ability to identify HFSS products was assessed using product-specific questions asking, “Do you think this product has high (nutrient of concern)?” Response options included yes, no, and I do not know. A response was considered correct if it accurately reflected the predefined nutrient profile of the product based on the study’s nutrient thresholds. Responses of “no” or “I do not know” for products that were high in a given nutrient were classified as incorrect.

Intentions to purchase HFSS products were assessed using product-specific questions asking participants whether they would purchase the product in the next week if it were available. Those responding “yes” were further asked to indicate the likelihood of purchase, allowing assessment of both purchase intention and its strength.

Secondary outcomes

Secondary outcomes included perceptions of unhealthiness, PME, and cognitive elaboration. Participants were asked, “Is this product unhealthy?” (yes/no). If they answered yes, they were asked, “How unhealthy is it?,” with response options ranging from 1 to 3 (very much, somewhat, very little). For PME, participants were asked whether the label made them feel concerned about the health consequences of consuming the product, made the product seem unpleasant, and made them feel discouraged from wanting to consume the product. For all label assessment items, response options were yes/no, and respondents who answered yes were subsequently asked, “How much?,” with responses ranging from 1 (very much) to 3 (very little). The mean was calculated from responses to the 3 PME items to create a composite PME score when internal consistency was acceptable (Cronbach’s α > 0.70). Cognitive elaboration was assessed by asking whether the label prompted participants to think about health problems associated with consuming the product (yes/no), followed by the extent of that thinking.

Tertiary outcomes

Tertiary outcomes assessed participants’ reactions to different FOPL formats. These included perceived label understanding, perceived learning of new information, perceived truthfulness of the label, desire to have the label displayed on products, attention capture, perceived unpleasantness, and concern about health consequences. Participants were also asked to indicate which FOPL format most discouraged product consumption. These outcomes were measured using structured response options and reported as the proportion of participants endorsing each response. These measures were assessed separately among adolescents and adults following exposure to the assigned labeling condition.

Statistical analysis

In our study, there were no missing values because an enumerator asked participants each question, which prevents participants from skipping items. Descriptive statistics were used to summarize participant characteristics. For all items measured on a Likert scale, responses from the agreement item (yes/no) and the corresponding strength-of-agreement item were combined for each participant to construct a 4-point Likert scale. The resulting scale was subsequently recoded from 1 (“not at all”) to 4 (“very much”) to facilitate more intuitive interpretation. In the main analyses, each FOP label condition was compared with the control arm. Pair-wise comparisons between label conditions were also conducted. Poisson regression was used for comparing proportion (percentage) of responses by study arm to the control condition for binary outcomes (correct identification compared with incorrect identification of nutrients of concern by study arms). This approach was selected because the prevalence of several study outcomes was relatively high (e.g., 73% correct identification among adolescents with WL labels) and prevalence ratios are more interpretable than odds ratios in cross-sectional analyses in such conditions [20], where logistic regression can substantially overestimate the effect size compared with risk ratios. Poisson regression models were also applied in a similar study conducted in South Africa [17] to compare the effects of different FOPL conditions.

For main analysis, results are presented as a percentage of participants with a positive outcome response. For pairwise comparisons, results are presented as relative risk (RR) estimates comparing pairs of label conditions, where an RR > 1 indicates a higher proportion of participants exposed to label X correctly identified products high in nutrients of concern or unhealthy products compared with those exposed to label Y. Moreover, pairwise comparisons of the outcome variable by the sociodemographic characteristics (gender, age, education, area of residence, and financial situation) are presented as percentages.

Although individual randomization was employed, some characteristics were not perfectly balanced, as indicated in Table 1. Therefore, the models were adjusted for potential confounders, including age, sex, level of education, and area of residence. However, the adjusted models yielded almost similar estimates in all cases. The covariate adjustment was not prespecified in the trial registration but was introduced during the analysis phase. Finally, the unadjusted models were used and described in the manuscript (Supplemental Figures 3 and 4). However, in line with CONSORT principles of transparency, the findings of the adjusted models were also presented.

Characteristics | ControlN = 1230 | HSRN = 1230 | WLN = 1244 | GDAN = 1233 | MTLN = 1237
Age (adolescents)
10–14 y | 310 (50.7) | 306 (50.4) | 316 (50.6) | 321 (51.3) | 323 (52.3)
15–18 y | 301 (49.3) | 301 (49.6) | 309 (49.4) | 305 (48.7) | 295 (47.7)
Sex (adolescents)
Girl | 294 (48.1) | 302 (49.8) | 323 (51.7) | 337 (53.8) | 285 (46.1)
Boy | 317 (51.9) | 305 (50.2) | 302 (48.3) | 289 (46.2) | 333 (53.9)
Education (adolescents)
Grade 0–4 | 174 (28.5) | 159 (26.2) | 160 (25.6) | 168 (26.8) | 153 (24.8)
Grade 5 | 101 (16.5) | 81 (13.3) | 105 (16.8) | 104 (16.6) | 102 (16.5)
Grade 6–9 | 250 (40.9) | 293 (48.3) | 283 (45.3) | 270 (43.1) | 289 (46.8)
Grade ≥10 | 86 (14.1) | 74 (12.2) | 77 (12.3) | 84 (13.4) | 74 (12.0)
Age (adults)
19–30 | 234 (37.8) | 230 (36.9) | 218 (35.2) | 194 (32.0) | 244 (39.4)
31–40 | 210 (33.9) | 202 (32.4) | 195 (31.5) | 234 (38.6) | 195 (31.5)
41–50 | 123 (19.9) | 133 (21.3) | 137 (22.1) | 124 (20.4) | 114 (18.4)
51–60 | 52 (8.4) | 58 (9.3) | 69 (11.1) | 55 (9.1) | 66 (10.7)
Sex (adults)
Male | 311 (50.2) | 286 (45.9) | 335 (54.1) | 297 (48.9) | 312 (50.4)
Female | 308 (49.8) | 337 (54.1) | 284 (45.9) | 310 (51.1) | 307 (49.6)
Education (adults)
Grade 0–4 | 88 (14.2) | 104 (16.7) | 128 (20.7) | 103 (17.0) | 109 (17.6)
Grade 5 | 176 (28.4) | 181 (29.1) | 161 (26.0) | 200 (32.9) | 154 (24.9)
Grade 6–9 | 172 (27.8) | 175 (28.1) | 161 (26.0) | 164 (27.0) | 194 (31.3)
Grade ≥10 | 183 (29.6) | 163 (26.2) | 169 (27.3) | 140 (23.1) | 162 (26.2)
Area of residence
Rural | 738 (60.0) | 740 (60.2) | 744 (59.8) | 740 (60.0) | 740 (59.8)
Urban | 492 (40.0) | 490 (39.8) | 500 (40.2) | 493 (40.0) | 497 (40.2)
Household financial situation
I can meet necessities and buy some luxury items with my income | 287 (23.3) | 305 (24.8) | 312 (25.1) | 283 (23.0) | 320 (25.9)
I can only meet my necessities with my income | 728 (59.2) | 717 (58.3) | 732 (58.8) | 724 (58.7) | 706 (57.1)
To meet my necessities, I have to take loans, I cannot buy the necessary stuff | 215 (17.5) | 208 (16.9) | 200 (16.1) | 226 (18.3) | 211 (17.1)

Consistent with prior studies, design weights were not required to be incorporated into the analysis, as the primary objective was to estimate the effect of FOPLs compared with a control label, rather than to produce nationally representative estimates of population-level indicators [21,22]. In addition, clusters of approximately equal size were employed, and an equal number of participants was randomly selected for each arm from the designated clusters. However, we replicated the analyses both with and without adjustment for clustering effects. The variation in risk ratio estimates was <5% in most cases, and the direction of association as well as the statistical significance levels remained unchanged. This approach enhanced the likelihood of obtaining unbiased estimates.

Although stratified analyses of the outcomes by sociodemographic characteristics were not planned during the preregistration process of the trial, we conducted exploratory stratified analyses and presented them as supplementary materials to understand whether the impact of the labels differed according to key sociodemographic variables, including gender, age, urban compared with rural, and household financial situation. Education was originally specified as a 4-level variable (with stratified analyses for each level): grade 0 to 4, grade 5, grade 6 to 9, and grade ≥10. We were also interested in assessing whether effects were moderated by literacy; however, we did not directly measure literacy. Thus, we used an education level of grade 0 to 4 as a proxy for illiteracy and conducted an additional stratified analyses for grade 0 to 4 compared with all higher grades.

All the statistical analysis was conducted using the STATA software version 18.