Qualitative assessment categories for AI-generated subject terms across the four tools: The metadata set. For each AI tool, the left column shows the number of generated terms assigned to each evaluation category, while the right column shows the percentage relative to the total number of terms generated by that tool
| Claude (n = 156) | ChatGPT (n = 96) | Deepseek (n = 156) | Gemini (n = 110) | Average | |||||
|---|---|---|---|---|---|---|---|---|---|
| Fully correct | 39 | 25.00% | 35 | 36.46% | 36 | 23.08% | 39 | 35.45% | 30.00% |
| Broader | 20 | 12.82% | 13 | 13.54% | 18 | 11.54% | 15 | 13.64% | 12.88% |
| Narrower | 2 | 1.28% | 1 | 1.04% | 1 | 0.64% | 2 | 1.82% | 1.20% |
| Related | 5 | 3.21% | 0 | 0.00% | 6 | 3.85% | 4 | 3.64% | 2.67% |
| Fully incorrect | 65 | 41.67% | 34 | 35.42% | 67 | 42.95% | 33 | 30.00% | 37.51% |
| Incorrect by QLIT definition or subject indexing policy | 14 | 8.97% | 10 | 10.42% | 16 | 10.26% | 13 | 11.82% | 10.37% |
| Not assigned by Queerlit but partially correct | 11 | 7.05% | 3 | 3.13% | 12 | 7.69% | 4 | 3.64% | 5.38% |
| Not assigned by Queerlit but fully correct | 0 | 0.00% | 0 | 0.00% | 0 | 0.00% | 0 | 0.00% | 0.00% |
| Claude ( | ChatGPT ( | Deepseek ( | Gemini ( | Average | |||||
|---|---|---|---|---|---|---|---|---|---|
| Fully correct | 39 | 35 | 36 | 39 | 30.00% | ||||
| Broader | 20 | 13 | 18 | 15 | 12.88% | ||||
| Narrower | 2 | 1 | 1 | 2 | 1.20% | ||||
| Related | 5 | 0 | 6 | 4 | 2.67% | ||||
| Fully incorrect | 65 | 34 | 67 | 33 | 37.51% | ||||
| Incorrect by QLIT definition or subject indexing policy | 14 | 10 | 16 | 13 | 10.37% | ||||
| Not assigned by Queerlit but partially correct | 11 | 3 | 12 | 4 | 5.38% | ||||
| Not assigned by Queerlit but fully correct | 0 | 0 | 0 | 0 | 0.00% | ||||
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.