Table A1

Glossary of terms and terminologies used in this paper based on Budhwar et al. (2023), Guo et al. (2022) and Chase (2023) 

TermExplanation
Generative pre-trained transformers (GPT)LLMs based on the transformer architecture, capable of generating coherent text responses by learning from vast amounts of unlabelled data (Budhwar et al., 2023)
Word embeddingsDense vector representations of words in multi-dimensional space, capturing semantic similarity and relationships among words (Guo et al., 2022)
Vector databasesEfficient databases that store and search text embeddings, capturing the semantic meaning of textual data (Guo et al., 2022)
Vectors in vector databasesVectors representing data characteristics (beyond usual attributes like colours or shapes) for efficient retrieval and grouping (Guo et al., 2022)
LangChainA popular Python LLM framework that simplifies building applications on top of LLMs like ChatGPT (Chase, 2023)
Chunk (as in chunks of text)In LangChain a “chunk” refers to a segment of text that has been divided from a larger body of text for easier processing and analysis by NLP models (Chase, 2023)
word2vecA widely used word embedding technique that represents words as dense vectors, enabling semantic similarity calculations (Guo et al., 2022)

or Create an Account

Close subscription notice
Close access options