On this page
How to audit packaging labels at retail
A reproducible protocol for measuring which sustainability marks and claims are present on packages, how they are expressed and which legal or scheme status they occupy.
The research object
A shelf audit is a content analysis of packages available in a defined retail environment during a defined period. It does not measure all packages placed on a national market, consumer exposure weighted by sales, legal compliance, or the environmental performance of the products. Those outcomes need different sampling frames and evidence. A defensible audit begins by stating which of them it does and does not estimate.
The unit of observation should ordinarily be the stock-keeping unit as presented to the consumer, with multipacks and size variants governed by a rule fixed in advance. Where identical artwork appears across sizes, the protocol may keep one representative observation and record the deduplication key. Where the amount or position of a claim changes, the packages are distinct observations even if the product is otherwise the same.
Sampling frame and site selection
Specify the market, retail channels, store formats, geography and field dates before collection. A convenience sample of one supermarket can show that a mark exists; it cannot estimate national prevalence. A prevalence study should stratify by channel and retailer, include urban and non-urban locations where distribution differs, and state whether private-label and imported products are oversampled. Treat online retail as a separate channel: the image set often omits back and side panels, so the observation process is different.
Category selection should follow the research question. A whole-store audit maximises breadth but is costly and may under-sample uncommon packaging formats. A category-stratified audit allows comparison of rigid, flexible, fibre, glass, metal and compostable formats. The protocol should publish its inclusion and exclusion rules, replacement procedure for unavailable stores, and treatment of seasonal or promotional packaging.
Image capture and chain of evidence
Every visible panel should be photographed at sufficient resolution to read qualification text, licence identifiers and material codes. The minimum image set is front, back, both sides, base, closure and any peel-back or removable label accessible without damaging stock. A colour and size reference should be included where the study measures prominence. File names should use a non-semantic observation identifier rather than a coder's interpretation of the mark.
The record should include retailer, store identifier, location, date and time, category, brand, product name, package size, barcode or GTIN where visible, packaging format and image completeness. Images should be retained unchanged; crops and enhanced copies used for coding should be derived files. The audit log should record substitutions, inaccessible panels, photography restrictions and uncertainty over whether a sleeve, label, cap or multipack is part of the observed packaging unit.
Codebook design
The codebook should separate observation from interpretation. Observational fields record the exact words, symbol family, position, size, colour, presence of qualification, data carrier and component to which the statement appears to refer. Interpretive fields then classify function: disposal instruction, material identification, recyclability claim, compostability claim, content claim, system-participation mark, sourcing mark, generic environmental benefit, future claim or digital disclosure.
Legal status is a third layer and should not be coded from appearance. It requires a jurisdiction and verification date: mandatory, adopted but not yet applicable, permitted conditional, voluntary licensed, self-declared, restricted, prohibited, proposed or not assessed. Scheme ownership, certificate number and territorial validity should be stored separately. “Unknown” and “not legible” are valid values; forcing a substantive code where the image cannot support it understates measurement error.
Pilot testing and inter-coder reliability
The codebook should be piloted on packages that represent difficult boundary cases, not only obvious marks. Two or more coders should independently code a shared subset after training. Agreement should be reported field by field, using raw agreement for transparent description and an appropriate chance-corrected statistic for categorical variables. A single overall reliability percentage can conceal poor performance in the legally important fields.
Disagreement should first improve definitions and examples, then trigger a second pilot. Adjudication creates the final dataset but does not replace reporting of pre-adjudication reliability. Changes to the codebook after fieldwork begins should be versioned, with affected observations re-coded. Automated image or text classification may assist retrieval, but a model's output should be validated against a human-coded sample and its errors reported by mark family and package format.
Legal and scheme assessment
A shelf audit can identify apparent non-conformity. It rarely proves a violation by observation alone. Scope may depend on manufacture date, producer size, packaging role, destination, evidence held off-pack, or a licence not visible to the auditor. Distinguish “observed representation inconsistent with the stated rule on the available facts” from “unlawful”. Check high-risk observations against the operative text, transition provisions, regulator guidance and scheme registry before publication.
Analysis and reporting
Prevalence should use the sampled SKU as denominator unless sales weights are available and their source is disclosed. Results should be reported with category, retailer and format distributions so that sample composition is visible. Separate counts should be given for packages, claims and marks: one package may carry several representations. Missing panels and illegible images should remain in the denominator where they could conceal a label, with sensitivity analysis showing the effect of alternative treatment.
Comparisons across time require a stable sampling frame, category definitions and coding rules. If a law changes the classification between waves, the original observations should be preserved and reclassified under both the old and new rule where the research question requires trend analysis. The publication should include the codebook, sampling protocol, field dates, flow of included observations, reliability results and a de-identified data dictionary.
Minimum protocol checklist
| Stage | Required record | Failure prevented |
|---|---|---|
| Scope | Market, channel, period, category and estimand. | Generalising a convenience sample to the whole market. |
| Sampling | Store frame, selection method, substitutions and SKU rule. | Hidden retailer or category bias. |
| Capture | All package panels, metadata, identifier and completeness flag. | Missing qualifications or misattributing a claim to the product rather than package. |
| Coding | Exact text and graphic before functional and legal classification. | Reading meaning into a familiar symbol. |
| Reliability | Independent duplicate coding and field-level statistics. | Treating coder judgement as objective observation. |
| Verification | Instrument, status, scope, transition and scheme registry. | Calling a lawful transitional package non-compliant. |
| Reporting | Denominators, missingness, sample composition and materials. | Publishing an uninterpretable prevalence percentage. |
References
INFORMAS, Food Labelling Protocol and module materials. Available at: Open source (Accessed: 20 August 2026). The protocol is an adjacent retail-labeling model; this Atlas adapts its principles to packaging sustainability claims.
Krippendorff, K. (2018) Content Analysis: An Introduction to Its Methodology, 4th edn. Thousand Oaks, CA: Sage.
Neuendorf, K.A. (2017) The Content Analysis Guidebook, 2nd edn. Thousand Oaks, CA: Sage.
National Academies of Sciences, Engineering, and Medicine (2022) Recycling and Waste Reduction: A Research Agenda. Washington, DC: National Academies Press.
Note on sources and verification
This is a proposed Atlas protocol, not a formally validated measurement instrument. Its content-analysis requirements follow established methodological literature and borrow the retail-labeling frame from INFORMAS, but the specific sustainability-label taxonomy has not yet undergone an external reliability study. Any project adopting the protocol should publish its modifications and validation results rather than cite this page as evidence that the resulting data are representative.
Last verified: 20 August 2026.