The pipeline summary
| Step | Method | Purpose | Output |
|---|---|---|---|
| Preprocessing | Tokenization, lemmatization (spaCy) | Clean and normalize text | Standardized corpus |
| Topic modelling | LDA (Gensim) | Identify latent themes | Topic distributions |
| Clustering | TF-IDF + PCA + k-means | Group similar discourse segments | Semantic clusters |
| Key-Phrase extraction | RAKE | Extract salient ESG terms | Keyword sets |
| Sentiment analysis | VADER | Measure evaluative tone | Sentiment scores |
| Digital framing | Keyword detection | Identify digital mediation | Binary/cluster-level tagging |
| Validation | Coherence, silhouette, expert review | Ensure robustness | Validated outputs |
| Step | Method | Purpose | Output |
|---|---|---|---|
| Preprocessing | Tokenization, lemmatization (spaCy) | Clean and normalize text | Standardized corpus |
| Topic modelling | LDA (Gensim) | Identify latent themes | Topic distributions |
| Clustering | TF-IDF + PCA + k-means | Group similar discourse segments | Semantic clusters |
| Key-Phrase extraction | RAKE | Extract salient ESG terms | Keyword sets |
| Sentiment analysis | VADER | Measure evaluative tone | Sentiment scores |
| Digital framing | Keyword detection | Identify digital mediation | Binary/cluster-level tagging |
| Validation | Coherence, silhouette, expert review | Ensure robustness | Validated outputs |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.