Glossary of terms and terminologies used in this paper based on Budhwar et al. (2023), Guo et al. (2022) and Chase (2023)
| Term | Explanation |
|---|---|
| Generative pre-trained transformers (GPT) | LLMs based on the transformer architecture, capable of generating coherent text responses by learning from vast amounts of unlabelled data (Budhwar et al., 2023) |
| Word embeddings | Dense vector representations of words in multi-dimensional space, capturing semantic similarity and relationships among words (Guo et al., 2022) |
| Vector databases | Efficient databases that store and search text embeddings, capturing the semantic meaning of textual data (Guo et al., 2022) |
| Vectors in vector databases | Vectors representing data characteristics (beyond usual attributes like colours or shapes) for efficient retrieval and grouping (Guo et al., 2022) |
| LangChain | A popular Python LLM framework that simplifies building applications on top of LLMs like ChatGPT (Chase, 2023) |
| Chunk (as in chunks of text) | In LangChain a “chunk” refers to a segment of text that has been divided from a larger body of text for easier processing and analysis by NLP models (Chase, 2023) |
| word2vec | A widely used word embedding technique that represents words as dense vectors, enabling semantic similarity calculations (Guo et al., 2022) |
| Term | Explanation |
|---|---|
| Generative pre-trained transformers (GPT) | LLMs based on the transformer architecture, capable of generating coherent text responses by learning from vast amounts of unlabelled data ( |
| Word embeddings | Dense vector representations of words in multi-dimensional space, capturing semantic similarity and relationships among words ( |
| Vector databases | Efficient databases that store and search text embeddings, capturing the semantic meaning of textual data ( |
| Vectors in vector databases | Vectors representing data characteristics (beyond usual attributes like colours or shapes) for efficient retrieval and grouping ( |
| LangChain | A popular Python LLM framework that simplifies building applications on top of LLMs like ChatGPT ( |
| Chunk (as in chunks of text) | In LangChain a “chunk” refers to a segment of text that has been divided from a larger body of text for easier processing and analysis by NLP models ( |
| word2vec | A widely used word embedding technique that represents words as dense vectors, enabling semantic similarity calculations ( |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.