Multilingual Political Ideology Embeddings Dataset: Metadata and Sentence Embeddings from Multiple Embedding Models
收藏资源简介:
This dataset contains multilingual sentence embedding files and the corresponding metadata used for cross-lingual political ideology analysis. The archive includes NumPy embedding matrices (.npy) generated with multiple embedding models, including BGE-M3, LaBSE, mDES, paraphrase-mpnet, and Qwen3-Embedding, together with a metadata.csv file. Rows in the embedding matrices correspond to the rows in metadata.csv. The dataset is intended to support research on multilingual representation learning, political ideology classification, cross-lingual retrieval, and comparative embedding evaluation. The files are provided to facilitate reproducibility of downstream analyses, including classification, retrieval, within-class similarity comparison, gender-matched evaluation, and cross-lingual embedding experiments. File contents:- metadata.csv: metadata associated with the observations used in the embedding analyses.- embeddings_*.npy: sentence embedding matrices generated by different multilingual embedding models. The dataset does not include model weights. Embeddings are provided as derived numerical representations for research and reproducibility purposes.



