RESEARCH METHODS; On the Use – and Abuse – of Power Analyses in Biomedical and Clinical Research

“Not powered to detect” has become a way to end the conversation, not a statistical finding.

by James Lyons-Weiler, PhD, Popular Rationalism, ©2026

(Oct. 1, 2026) — Imagine that researchers test whether a treatment relieves symptoms. They also watch for a serious complication, but their study includes too few people to reliably detect an important increase in that complication. When the results arrive, the treatment improves symptoms, and the investigators report “no statistically significant increase” in the complication. That description may be accurate. The problem begins when the summary turns an unresolved safety question into reassurance that the treatment does not cause the complication.

This is the distinction that the phrase “not powered to detect” can illuminate—or obscure. Used honestly, it identifies a limitation and helps define the next investigation. Used lazily, it replaces examination of the actual results with a statistical label. Used strategically, it dismisses an inconvenient signal while leaving an equally uncertain claim of safety undisturbed. A design’s limitations do not change according to which conclusion its audience prefers.

Safety must be a prospective research objective, not an inference borrowed from a successful efficacy result. A trial that supports a specific safety claim needs the participants, observation time, measurements, and analysis required to support that claim. Statistical power helps investigators plan that work. It does not turn a failure to detect harm into evidence that the harm cannot occur. Goodman and Berlin explained the separation between planning and interpretation in 1994; the distinction remains essential to reading clinical research. [1]

What power means—and what it does not

Statistical power describes how likely a planned analysis is to detect a specified effect, assuming that effect truly exists at the size used in the calculation. For a safety question, the effect might be an increase from one complication per 1,000 people to two per 1,000. For an efficacy question, it might be a particular improvement in pain or a reduction in hospitalizations. “Eighty percent power” has no useful meaning until someone explains power to detect what, in whom, and over what period. [2]

Think of repeating the same study under the same assumptions. A design with 80 percent power would detect the specified effect in about 80 out of 100 repetitions, on average. Approximately 20 would miss it. That does not mean the treatment is 80 percent safe, that 80 percent of participants experience the effect, or that the conclusion of a completed study has an 80 percent probability of being correct. Figure 1 illustrates this distinction using entire hypothetical trials, not individual patients.


Read the rest here.


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.