Tutorial on Sample Size Calculations for the precision and the testing of Cohen’s Kappa: a methodological review and new proposals

Sample Size for qualitative agreement

Authors

DOI:

https://doi.org/10.54103/2282-0930/29771

Keywords:

Raters' agreement, Cohen's kappa statistic, Sample size calculation

Abstract

Introduction

We have revised the steps of the sample sizes calculation for the biomedical research.

Then, focusing on the agreement studies on qualitative variables we have shown the derivation of the Cohen’s kappa together with the factors influencing its value, leading to have low kappa values even in presence of a relevant observed agreement.

We have revised the sample size calculation approaches for the Cohen’s kappa proposed in the statistical literature.

Methods

We have calculated the sample sizes for agreement studies based on the Cohen’s kappa statistics according to the main relevant literature proposals, such as those of Flack et al.[72] and of Donner et al. [90,92,96]. Then, we have proposed a partial extension of the common correlation model (PCCM) for 2x2 contingency table to cxc square contingency tables with a common correlation coefficient of the cells on the principal diagonal, considered pairwise. Finally, we have conceived a full generalization with a common correlation coefficient model (FCCM) of all cells of the square contingency table.

Results

From our PCCM, we have obtained very similar maximum and minimum values of the kappa variance, leading to have sample sizes slightly different. Otherwise, from FCCM, it is obtained a unique contingency table and, consequently, a unique value of the kappa variance to be used under the null and the alternative hypothesis leading to have only one value of the sample size. More relevantly, this latter sample size is within or equal to the sample sizes calculated by using the maximum and minimum value of the kappa variance under PCCM.

In the case of 2x2 contingency tables, all our sample sizes calculation proposed methods (SS-A&C-max, SS-A&C-min, and SS-A&C-full) gave equal sample sizes which are, in addition, equal to those calculated according to Flack et al.[72], being the sample sizes from Donner et al. [90,92] generally greater.

For 3x3 contingency tables, the sample sizes calculated according to our proposed models are lower than or equal to those calculated according to Flack et al.[72] and, also almost always, lower than those from Donner et al. [90,92]. The same occurs for the 4x4 tables.

Discussion / conclusions

Sample sizes from our proposed models can be generally recommended since they are lower than or at maximum equal to those obtained by Flack et al.’s approach[72]. In addition, our proposed procedures give very similar maximum and minimum sample size values under PCCM and only one sample size value under FCCM. Furthermore, sample size proposed by Donner et al [90,92,96] can be considered only in the case in which they are lower and the primary objective of the study is the “agreement or not” instead of a diversified level of agreement assessed through the weighted kappa and, finally, there are more than two raters.

Finally, the fact that there is a unique kappa variance value under the FCCM, leading to sample sizes always within those calculated from the maximum and minimum values of the kappa variance under PCCM, allows to argue that this model actually refers to the population probability table, and consequently, it has to be warmly recommended

Downloads

Download data is not yet available.

Author Biography

Paolo Antonelli, Vita-Salute San Raffaele University

Retired Professor of Calculus of Probabilities, Statistics and Operative Research, at the State Industrial Technical Institute (ITIS) Benedetto Castelli, Brescia, Italy.

 

Contract Professor of Medical Statistics. Section of Medical Statistics and Biometry, Department of Molecular and Translational Medicine, Faculty of Medicine and Surgery, University of Brescia. V.le Europa 11, 25123 Brescia, Italy.

References

https://www.ncss.com/software/pass/pass-documentation/

PASS®13 Power Analysis and Sample Size Software 2014. NCSS, LLC. Kaysville, Utah, USA, ncss.com/software/pass. PASS®16 Power Analysis and Sample Size Software 2018. NCSS, LLC. Kaysville, Utah, USA, ncss.com/software/pass.

Flack VFA, Afifi A, Lachenbruch PA, Schouten HJ. A. Sample Size Determinations for the two Rater Kappa Statistic. Psychometrika, 1988; 53(3): 321-325.

Fleiss JL, Cohen J, Everitt BS. Large sample standard errors of kappa and weighted kappa Psychol Bull. 1969, 72,5: 323-327.

Cohen JA. A coefficient of agreement for nominal scales. Educ Psychol Meas. 1960; 20: 213–220.

Bujang MA, Baharum N. Guidelines of the minimum sample size requirements for Cohen’s Kappa. Epidemiol Biostat Public Health 2017, 14 (2): e12267-1-10.

Donner A, Eliasziw M. A Goodness-Of-Fit Approach to Inference Procedures for the Kappa Statistic: Confidence and Sample Size Estimation Interval Construction, Significance-Testing. Statist Med. 1992; 11: 1511-1519.

Donner A, Eliasziw M. Statistical Implications of the Choice Between a Dichotomous or Continuous Trait in Studies of Interobserver Agreement. Biometrics. 1994; 50: 550-555.

Donner A, Eliasziw M, Klar N., Testing the Homogeneity of Kappa Statistics. Biometrics. 1996; 52: 176-183

Donner A. Sample Size Requirements for the Comparison of two or more Coefficients of Inter-Observer Agreement. Statist. Med. 1998;17: 1157-1168.

Altaye M, Donner A, Klar N. Inference procedures for assessing interobserver agreement among multiple raters. Biometrics. 2001; 57: 584-588.

Altaye M., Donner A., Eliasziw M. A general goodness-of-fit approach for inference procedures concerning the kappa statistic. Statist. Med. 2001; 20:2479–2488 doi: 10.1002/sim.911.

Feinstein AR, Cicchetti DV. High agreement but low kappa: I. Resolving the paradoxes. J Clin Epidemiol. 1990;43, 6: 543-549.

Thompson WD, Walter SD. A reappraisal of the kappa coefficient. J Clin Epidemiol. 1988;41:949–958.

Shoukri MM. Measures of Interobserver Agreement. 2004 Boca Raton, Fla: Chapman & Hall/CRC.

Brennan RL, Prediger DJ. Coefficient kappa: Some uses, misuses, and alternatives. Educ Psychol Meas. 1981; 41: 687–699.

Gwet K. Computing inter-rater reliability and its variance in presence of high agreement. Br. J. Math. Stat. Psychol. 2008; 61: 29-48.

Gwet KL. Handbook of Inter-Rater Reliability: The Definitive Guide to Measuring the Extent of Agreement Among Raters. 4th ed. Gaithersburg, MD: Advanced Analytics. 2014.

Klein D. Implementing a general framework for assessing interrater agreement in Stata. SJ 2018;18: 871–901. doi: 10.1177/1536867x1801800408. 2018.

Falcaro M, Newson RB. tabagree: Nonparametric measures of agreement and disagreement in paired ordinal data. SJ 2025; 25(3):627–645 doi: 10.1177/1536867X251365495.

Svensson E. Analysis of Systematic and Random Differences Between Paired Ordinal Categorical Data. Stockholm: Almqvist and Wiksell 1993.

Svensson E. A coefficient of agreement adjusted for bias in paired ordered categorical data. Biom J. 1997;39:643–657. doi:10.1002/bimj.4710390602.

Cohen J. Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit. Psychol Bull. 1968;70:213–220.

Fleiss JL, Nee JCM, Landis JR. Large sample variance of kappa in the case of different sets of raters. Psychol Bull. 1979;86:974–977.

Kraemer HC, Periyakoil VS, Nod A. Chapter 1.3 Agreement Statistics; Tutorial In Biostatistics; Kappa coefficients in medical research. Tutorials in Biostatistics Volume 1: Statistical Methods in Clinical Studies Edited by RB. D’Agostino 2004 John Wiley & Sons, Ltd.

Donner A. Sample size requirements for interval estimation of the intraclass kappa statistic. Commun Stat-Simul C Journal 1999; 28(2): 415-429.

Donner A, Rotondi MA. Sample Size Requirements for Interval Estimation of the Kappa Statistic for Interobserver Agreement Studies with a Binary Outcome and Multiple Raters. Int J Biostat. 2010; 6(1) Article 31.

Rotondi MA, Donner A. A Confidence Interval Approach to Sample Size Estimation for Interobserver Agreement Studies with Multiple Raters and Outcomes. J. Clin. Epidemiol. 2012;65:778-784.

Choudhary PK, Nagaraja HN. Measuring Agreement: Models, Methods, and Applications. 2017 John Wiley & Sons, Inc. Hoboken, USA.

Cantor AB. Sample-Size Calculations for Cohen's Kappa. Psychol. Methods. 1996; 1(2):150-153.

Rotondi MA. Package ‘kappaSize’, Title: Sample Size Estimation Functions for Studies of Interobserver Agreement Version 1.1 released on February 20, 2015 and Version 1.2 released on November 26, 2018, Author: Michael A. Rotondi, Maintainer: Michael A. Rotondi

Von Eye A, Mun EY. Analyzing Rater Agreement: Manifest Variable Methods. Lawrence Erlbaum Associates; 2005.

https://cran.r-project.org/web/packages/lpSolve/index.html Package ‘lpSolve’, January 24, 2020, Version 5.6.15 Title Interface to 'Lp_solve' v. 5.5 to Solve Linear/Integer Programs URL https://github.com/gaborcsardi/lpSolve.

Elashoff JD. nQuery Advisor ® 2007 Version 7.0 User’s Guide. Statistical Solutions Ltd., Republic of Ireland. http://www.statsolusa.com info@statsolusa.com.

Vanbelle S, Hernandez Engelhart C, Blix E. Measuring Agreement in Diagnostics: A Practical Guide for Researchers (Tutorial in Biostatistics). Stat Med, 2025;44, e70299

https://cran.r-project.org/web/packages/magree/index.html (via r-universe) July 23, 2025.

Fleiss JL. Measuring Nominal Scale Agreement among Many Raters. Psychol. Bull. 1971; 76(5): 378-382.

Fleiss J. Statistical Methods for Rates and Proportions. Wiley, New York. 1981.

Schouten HJA. Measuring pairwise agreement among many observers Biom J. 1880;22(6): 497-504].

Moss J. Measures of Agreement with Multiple Raters: Fréchet Variances and Inference Psychometrika 2024,89(2): 517–541. https://doi.org/10.1007/s11336-023-09945-2.

Published

2026-07-28

How to Cite

1.
Cesana BM, Antonelli P. Tutorial on Sample Size Calculations for the precision and the testing of Cohen’s Kappa: a methodological review and new proposals: Sample Size for qualitative agreement. ebph [Internet]. 2026 [cited 2026 Aug. 5];21(2). Available from: https://riviste.unimi.it/index.php/ebph/article/view/29771

Issue

Section

Biostatistics
Received 2025-09-04
Accepted 2026-04-02
Published 2026-07-28