Table 5

The pipeline summary

StepMethodPurposeOutput
PreprocessingTokenization, lemmatization (spaCy)Clean and normalize textStandardized corpus
Topic modellingLDA (Gensim)Identify latent themesTopic distributions
ClusteringTF-IDF + PCA + k-meansGroup similar discourse segmentsSemantic clusters
Key-Phrase extractionRAKEExtract salient ESG termsKeyword sets
Sentiment analysisVADERMeasure evaluative toneSentiment scores
Digital framingKeyword detectionIdentify digital mediationBinary/cluster-level tagging
ValidationCoherence, silhouette, expert reviewEnsure robustnessValidated outputs

or Create an Account

Close Modal
Close Modal