Chi-Square Test

← Back to Index (🔢 Statistics)

Overview Of The Chi-Square Test

The Chi-square test is a fundamental non-parametric statistical procedure utilized extensively in medical research and epidemiology. It is designed specifically for the analysis of unordered categorical or qualitative data.

Types Of Chi-Square Tests

Depending on the specific research question and study design, the basic principles of the test can be applied in three distinct ways.

Goodness Of Fit Test

Test Of Independence Or Association

Test For Trend

Core Mathematical Principles And Formulas

The execution of the test requires meticulous calculation of expected counts and the subsequent test statistic.

Calculation Of Expected Frequency

The Basic Chi-Square Statistic

Degrees Of Freedom

The 2x2 Fourfold Contingency Table

The simplest and most frequent application of the test compares two binary variables.

Treatment / Exposure Group Disease Present Disease Absent Total
Group A (Exposed) a b a + b
Group B (Unexposed) c d c + d
Total a + c b + d a + b + c + d

Yates Continuity Correction

Core Assumptions And Limitations

Violating the underlying assumptions compromises the validity of the resulting p-value.

Data Requirements

The Rule Of 5

Limitation Regarding Effect Size

Alternative Tests For Categorical Data

When the standard assumptions are not met, alternative statistical procedures must be deployed.

Clinical Scenario Suggested Statistical Test Rationale And Characteristic
Independent samples, expected cell counts < 5 Fisher's Exact Test Used when the "Rule of 5" is violated or sample size is very small (<40). It calculates the exact probability using the hypergeometric distribution rather than an approximation.
Paired or matched observations McNemar's Test Used for dependent data (e.g., pre-test/post-test crossover designs on the same subjects). It evaluates subjects who changed their status.
Evaluating specific covariates Logistic Regression While Chi-square tests association, logistic regression allows for predictive modeling and adjustment for multiple confounding variables simultaneously.

Handling Larger Contingency Tables (R x C)