Table 1

Qualitative assessment categories for AI-generated subject terms across the four tools: The metadata set. For each AI tool, the left column shows the number of generated terms assigned to each evaluation category, while the right column shows the percentage relative to the total number of terms generated by that tool

Claude (n = 156)ChatGPT (n = 96)Deepseek (n = 156)Gemini (n = 110)Average
Fully correct3925.00%3536.46%3623.08%3935.45%30.00%
Broader2012.82%1313.54%1811.54%1513.64%12.88%
Narrower21.28%11.04%10.64%21.82%1.20%
Related53.21%00.00%63.85%43.64%2.67%
Fully incorrect6541.67%3435.42%6742.95%3330.00%37.51%
Incorrect by QLIT definition or subject indexing policy148.97%1010.42%1610.26%1311.82%10.37%
Not assigned by Queerlit but partially correct117.05%33.13%127.69%43.64%5.38%
Not assigned by Queerlit but fully correct00.00%00.00%00.00%00.00%0.00%
Source(s): Author’s own work

or Create an Account

Close subscription notice
Close access options