Wellness programs are a standard feature of American employer benefits, and most report their results in terms of engagement. The choice of metric explains a great deal about the programs themselves.

Why participation became the metric

Participation is recorded automatically — logins, biometric screenings completed, challenges joined. It requires no additional data collection and produces a number every quarter.

Health outcomes require years, clinical data and a comparison group. Few employers have the population size or the tenure stability to detect a change even if one occurred.

A program is therefore evaluated on the thing it can measure, and designed to move that thing. Features that generate logins outcompete features that might change health.

What the incentive structure does

Programs commonly attach financial incentives to participation, often as premium differentials. Federal rules constrain how large those differentials can be and how they may be structured.

The rules exist because tying too much money to a health measure begins to resemble charging sick employees more, which other statutes prohibit.

The compromise is that incentives generally attach to activity rather than results — completing a screening rather than reaching a target value.

Where the selection problem enters

Participants are volunteers. People who join fitness challenges tend to differ from those who do not in ways related to health before the program starts.

Comparing participants to non-participants therefore measures who signed up rather than what the program did. This is the single largest flaw in reported savings figures.

Randomized evaluations that avoid the problem have generally found smaller effects on medical spending than the observational reports the industry cites.

What the programs reliably deliver

Screening uptake does rise, and screenings do identify conditions that would otherwise go undetected — which shifts spending forward rather than reducing it.

Programs also serve as a recruitment and retention signal, communicating something about the employer regardless of health effects.

These are real returns, but they are not the medical cost reductions that justified the category commercially, and the distinction is often blurred in vendor material.

Why the design is shifting

Attention has moved toward mental health support, caregiving assistance and financial wellbeing, categories where employees report unmet need rather than programs the employer finds easy to run.

These are harder to gamify and harder to count, which is precisely why they were underweighted while participation dominated reporting.

Employees evaluating a program should read what data is collected and who receives it, since privacy terms vary between vendors and are set by contract rather than by a single national standard.