A result is 'statistically significant' when it would be unlikely (usually under 5%) to arise by chance if the treatment did nothing. It is a threshold, not a measure of benefit size: a tiny gain can pass it in a huge trial and a large one can miss it in a small trial, so ESMO and ASCO grade benefit separately.
Trials are sized ('powered') to detect a pre-specified effect with 80-90% probability at a 5% two-sided alpha; testing several endpoints or several interim looks would inflate false positives unless alpha is split (allocated) or spent sequentially, so hierarchical testing means a secondary endpoint cannot be claimed if a higher one failed. 'Numerically better', 'trend' and 'nominal P' are phrases for results that did not meet the pre-set bar. Clinical meaningfulness is separate: ASCO and ESMO scales (ESMO-MCBS) grade the size of benefit, and a significant 6-week PFS gain may not matter to patients. Confidence intervals convey both size and uncertainty and are preferred to bare P values.
Shares Statistical power, sample size and re-estimation, Primary, secondary and co-primary endpoints, Interim analysis, readout and data cut-off, Hazard ratio (HR).
Shares Futility analysis (stopped for futility), Statistical power, sample size and re-estimation, Why trials fail: underpowered, wrong endpoint, control arm drift, subgroup fishing, crossover.
Shares P-value, Confidence interval, Hazard ratio (HR).
Shares Subgroup analysis (forest plots), Pre-specified vs post-hoc analysis, Futility analysis (stopped for futility), Group sequential design, stopping rules and alpha spending.
Shares P-value, Confidence interval, Primary, secondary and co-primary endpoints, Hazard ratio (HR).
Shares Futility analysis (stopped for futility), Group sequential design, stopping rules and alpha spending, Interim analysis, readout and data cut-off.
Shares Subgroup analysis (forest plots), Pre-specified vs post-hoc analysis, Why trials fail: underpowered, wrong endpoint, control arm drift, subgroup fishing, crossover.
Shares Pre-specified vs post-hoc analysis, Group sequential design, stopping rules and alpha spending, Primary, secondary and co-primary endpoints, Why trials fail: underpowered, wrong endpoint, control arm drift, subgroup fishing, crossover.