Intercoder reliability thresholds
| Statistic | Description | Interpretative thresholds | ||
|---|---|---|---|---|
| Weak | Moderate | Strong | ||
| Krippendorff α | agreement across all Study IDs × Codes; provides a single measure of reliability | <0.67 | 0.67–0.80 | ≥0.80 |
| Fleiss' K | agreement among coders across the entire set of codes, averaged over all individual per-code κ values | <0.40 | 0.40–0.75 | >0.75 |
| Cohen's K | overall agreement between each pair of coders; used to identify coder-to-coder alignment patterns | <0.59 | 0.60–0.79 | ≥0.80 |
| Statistic | Description | Interpretative thresholds | ||
|---|---|---|---|---|
| Weak | Moderate | Strong | ||
| Krippendorff α | agreement across all Study IDs × Codes; provides a single measure of reliability | <0.67 | 0.67–0.80 | ≥0.80 |
| Fleiss' | agreement among coders | <0.40 | 0.40–0.75 | >0.75 |
| Cohen's | overall agreement between each pair of coders; used to identify coder-to-coder alignment patterns | <0.59 | 0.60–0.79 | ≥0.80 |
Note(s): Codes that were never assigned to a transcript in each Round were excluded from the Fleiss' K calculation
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.