The probability of seeing a difference at least this large if the treatment actually did nothing. Below 0.05 (a 1 in 20 chance) is the conventional threshold for calling a result 'statistically significant'. It measures surprise, not the size of the benefit: a trivial gain in a huge trial can have a tiny p-value.
A small p-value means the result would be unlikely under pure chance, but it does not measure how big or how clinically meaningful the effect is; a trivial benefit in a huge trial can have a tiny p-value while a large benefit in a small trial can miss significance. Trials with several endpoints or several interim looks must spend their 0.05 across them (hierarchical testing, alpha spending), which is why an endpoint can be 'nominally significant' yet not count formally. The confidence interval conveys the same information about significance while also showing the size of the effect, and is generally more useful.
Shares Confidence interval, Statistical significance (P values, alpha, multiplicity), Group sequential design, stopping rules and alpha spending, Hazard ratio (HR).
Shares Reading a hazard ratio, Confidence interval, Hazard ratio (HR).
Shares Pre-specified vs post-hoc analysis, Statistical significance (P values, alpha, multiplicity), Group sequential design, stopping rules and alpha spending, Endpoint.
Shares Reading a hazard ratio, Confidence interval, Hazard ratio (HR).
Shares Reading a hazard ratio, Confidence interval, Randomised trial, Hazard ratio (HR).
Shares Pre-specified vs post-hoc analysis, Statistical significance (P values, alpha, multiplicity).
Shares Endpoint, Randomised trial.
Shares Pre-specified vs post-hoc analysis, Randomised trial.