Table 2

Qualitative assessment categories for AI-generated subject terms across the four tools: The full-text set. For each AI tool, the left column shows the number of generated terms assigned to each evaluation category, while the right column shows the percentage relative to the total number of terms generated by that tool

Claude (n = 178)ChatGPT (n = 123)Deepseek (n = 178)Gemini (n = 111)Average
Fully correct1810.11%86.50%158.43%98.11%8.29%
Broader2312.92%97.32%168.99%1412.61%10.46%
Narrower10.56%00.00%21.12%10.90%0.65%
Related84.49%21.63%42.25%54.50%3.22%
Fully incorrect8849.44%7964.23%10156.74%6054.05%56.12%
Incorrect by QLIT definition or subject indexing policy2715.17%1814.63%2011.24%1614.41%13.86%
Not assigned by Queerlit but partially correct137.30%75.69%2011.24%65.41%7.41%
Not assigned by Queerlit but fully correct00.00%00.00%00.00%00.00%0.00%
Source(s): Author’s own work

or Create an Account

Close subscription notice
Close access options