
This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Indonesian Pre-trained Language Models with OCEAN-aware Query Expansion Ranking for Personality Prediction from Social Media Text
Corresponding Author(s) : Gede Aditra Pradnyana
Kinetik: Game Technology, Information System, Computer Network, Computing, Electronics, and Control,
Vol. 11, No. 3, August 2026 (Article in Progress)
Abstract
Personality prediction from social media text has become a critical topic in computational social science, as online posts frequently reflect individual behavioral and psychological characteristics. However, predicting personality from Indonesian social media material remains challenging due to informal language, slang, abbreviations, noisy expressions, and implicit personality-related cues. This study proposed a new OCEAN personality prediction framework that integrates pre-trained language models with an OCEAN-aware Query Expansion Ranking feature learning mechanism. Unlike conventional approaches that rely mainly on contextual embeddings, the proposed framework introduced trait-specific class-discriminative lexical features to strengthen personality-related representation. The prediction task was structured as five separate binary classification problems, with each personality attribute divided into High and Low groups.. The experiments were conducted in two stages. First, numerous pre-trained language models, including IndoBERT, IndoBERTweet, Indonesian RoBERTa, and multilingual BERT, were fine-tuned and compared to identify the most suitable contextual representation model. The best baseline model, multilingual BERT, achieved an average accuracy of 75.48% and an average F1-score of 73.58%. Second, multilingual BERT was integrated with class-discriminative lexical features generated by the proposed OCEAN-aware Query Expansion Ranking mechanism. The trait-specific configuration improved the average accuracy to 78.76% and the average F1-score to 76.48%. These results demonstrated that integrating contextual semantic representations with OCEAN-aware lexical feature learning enhanced personality prediction from Indonesian social media material while also offering a more clear representation of trait-relevant language signals.
Keywords
Download Citation
Endnote/Zotero/Mendeley (RIS)BibTeX
- F. Celli et al., “Twenty Years of Personality Computing: Threats, Challenges and Future Directions,” ACM Comput. Surv., vol. 58, no. 11, pp. 1–37, Aug. 2026, doi: 10.1145/3806009.
- G. A. Pradnyana, W. Anggraeni, E. M. Yuniarno, and M. H. Purnomo, “An explainable ensemble model for revealing the level of depression in social media by considering personality traits and sentiment polarity pattern,” Online Soc. Netw. Media, vol. 46, May 2025, doi: 10.1016/j.osnem.2025.100307.
- G. Z. Nabiilah and D. Suhartono, “Personality Classification Based on Textual Data using Indonesian Pre-Trained Language Model and Ensemble Majority Voting,” Revue d’Intelligence Artificielle, vol. 37, no. 1, pp. 73–81, Feb. 2023, doi: 10.18280/ria.370110.
- Y. Mehta, N. Majumder, A. Gelbukh, and E. Cambria, “Recent trends in deep learning based personality detection,” Artif. Intell. Rev., vol. 53, no. 4, pp. 2313–2339, Apr. 2020, doi: 10.1007/s10462-019-09770-z.
- T. Thurairasa and L. Rupasinghe, “A Literature Review in Personality Predictions Based o n Twitter Text Modality,” in International Conference on Advances in Computing and Technology (ICACT–2020) Proceedings, 2020, pp. 172–174.
- A. Bruno and G. Singh, “Personality Traits Prediction from Text via Machine Learning,” 2022 IEEE World Conference on Applied Intelligence and Computing (AIC), 2022, doi: 10.1109/aic.2022.99.
- W. Kang, F. Steffens, S. Pineda, K. Widuch, and A. Malvaso, “Personality traits and dimensions of mental health,” Sci. Rep., vol. 13, no. 1, p. 7091, May 2023, doi: 10.1038/s41598-023-33996-1.
- G. A. Pradnyana, W. Anggraeni, E. M. Yuniarno, and M. H. Purnomo, “Fine-Tuning IndoBERT Model for Big Five Personality Prediction from Indonesian Social Media,” in 2023 International Seminar on Intelligent Technology and Its Applications (ISITIA), IEEE, Jul. 2023, pp. 93–98. doi: 10.1109/ISITIA59021.2023.10221074.
- M. L. Smith, D. Hamplová, J. Kelley, and M. D. R. Evans, “Concise survey measures for the Big Five personality traits,” Res. Soc. Stratif. Mobil., vol. 73, Jun. 2021, doi: 10.1016/j.rssm.2021.100595.
- J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Oct. 2018, [Online]. Available: http://arxiv.org/abs/1810.04805
- F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” Sep. 2021, [Online]. Available: http://arxiv.org/abs/2109.04607
- F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Online, 2020, pp. 757–770. [Online]. Available: https://huggingface.co/
- R. L. Vásquez and J. Ochoa-Luna, “Transformer-based Approaches for Personality Detection using the MBTI Model,” in Proceedings - 2021 47th Latin American Computing Conference, CLEI 2021, Institute of Electrical and Electronics Engineers Inc., 2021. doi: 10.1109/CLEI53233.2021.9640012.
- E. Kerz, Y. Qiao, S. Zanwar, and D. Wiechmann, “Pushing on Personality Detection from Verbal Behavior: A Transformer Meets Text Contours of Psycholinguistic Features,” in Proceedings of the 12th Workshop on Computational Approaches to Subjectivity, Sentiment & Social Media Analysis, Stroudsburg, PA, USA: Association for Computational Linguistics, 2022, pp. 182–194. doi: 10.18653/v1/2022.wassa-1.17.
- A. F. Hidayatullah, R. A. Apong, D. T. C. Lai, and A. Qazi, “Pre-trained language model for code-mixed text in Indonesian, Javanese, and English using transformer,” Soc. Netw. Anal. Min., vol. 15, no. 1, p. 30, Mar. 2025, doi: 10.1007/s13278-025-01444-9.
- L. Hu, H. He, D. Wang, Z. Zhao, Y. Shao, and L. Nie, “LLM vs Small Model? Large Language Model Based Text Augmentation Enhanced Personality Detection Model,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, pp. 18234–18242, Mar. 2024, doi: 10.1609/aaai.v38i16.29782.
- T. Parlar, S. A. Özel, and F. Song, “QER: a new feature selection method for sentiment analysis,” Human-centric Computing and Information Sciences, vol. 8, no. 1, Dec. 2018, doi: 10.1186/s13673-018-0135-8.
- T. Parlar and S. A. Ozel, “A new feature selection method for sentiment analysis of Turkish reviews,” in Proceedings of the 2016 International Symposium on Inovations in Intelligent Systems and Applications, INISTA 2016, Institute of Electrical and Electronics Engineers Inc., Sep. 2016. doi: 10.1109/INISTA.2016.7571833.
- P. H. Prastyo, R. Hidayat, and I. Ardiyanto, “Enhancing sentiment classification performance using hybrid Query Expansion Ranking and Binary Particle Swarm Optimization with Adaptive Inertia Weights,” ICT Express, vol. 8, no. 2, pp. 189–197, Jun. 2022, doi: 10.1016/j.icte.2021.04.009.
- G. A. Pradnyana, W. Anggraeni, E. M. Yuniarno, and M. H. Purnomo, “Enhancing MBTI Personality Trait Prediction from Imbalanced Social Media Data Using Hybrid Query Expansion Ranking and Glo Ve- BiLSTM,” in 2023 IEEE International Conference on Fuzzy Systems (FUZZ), IEEE, Aug. 2023, pp. 1–6. doi: 10.1109/FUZZ52849.2023.10309718.
- V. Ong, A. D. S. Rahmanto, W. Williem, N. H. Jeremy, D. Suhartono, and E. W. Andangsari, “Personality Modelling of Indonesian Twitter Users with XGBoost Based on the Five Factor Model,” International Journal of Intelligent Engineering and Systems, vol. 14, no. 2, pp. 248–261, 2021, doi: 10.22266/ijies2021.0430.22.
- H. Lucky, Roslynlia, and D. Suhartono, “Towards Classification of Personality Prediction Model: A Combination of BERT Word Embedding and MLSMOTE,” in Proceedings of 2021 1st International Conference on Computer Science and Artificial Intelligence, ICCSAI 2021, Institute of Electrical and Electronics Engineers Inc., 2021, pp. 346–350. doi: 10.1109/ICCSAI53272.2021.9609750.
- Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” Jul. 2019, Accessed: May 04, 2023. [Online]. Available: https://arxiv.org/abs/1907.11692v1
- S. Wu and M. Dredze, “Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 833–844. doi: 10.18653/v1/D19-1077.
- T. Pires, E. Schlinger, and D. Garrette, “How Multilingual is Multilingual BERT?,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4996–5001. doi: 10.18653/v1/P19-1493.
References
F. Celli et al., “Twenty Years of Personality Computing: Threats, Challenges and Future Directions,” ACM Comput. Surv., vol. 58, no. 11, pp. 1–37, Aug. 2026, doi: 10.1145/3806009.
G. A. Pradnyana, W. Anggraeni, E. M. Yuniarno, and M. H. Purnomo, “An explainable ensemble model for revealing the level of depression in social media by considering personality traits and sentiment polarity pattern,” Online Soc. Netw. Media, vol. 46, May 2025, doi: 10.1016/j.osnem.2025.100307.
G. Z. Nabiilah and D. Suhartono, “Personality Classification Based on Textual Data using Indonesian Pre-Trained Language Model and Ensemble Majority Voting,” Revue d’Intelligence Artificielle, vol. 37, no. 1, pp. 73–81, Feb. 2023, doi: 10.18280/ria.370110.
Y. Mehta, N. Majumder, A. Gelbukh, and E. Cambria, “Recent trends in deep learning based personality detection,” Artif. Intell. Rev., vol. 53, no. 4, pp. 2313–2339, Apr. 2020, doi: 10.1007/s10462-019-09770-z.
T. Thurairasa and L. Rupasinghe, “A Literature Review in Personality Predictions Based o n Twitter Text Modality,” in International Conference on Advances in Computing and Technology (ICACT–2020) Proceedings, 2020, pp. 172–174.
A. Bruno and G. Singh, “Personality Traits Prediction from Text via Machine Learning,” 2022 IEEE World Conference on Applied Intelligence and Computing (AIC), 2022, doi: 10.1109/aic.2022.99.
W. Kang, F. Steffens, S. Pineda, K. Widuch, and A. Malvaso, “Personality traits and dimensions of mental health,” Sci. Rep., vol. 13, no. 1, p. 7091, May 2023, doi: 10.1038/s41598-023-33996-1.
G. A. Pradnyana, W. Anggraeni, E. M. Yuniarno, and M. H. Purnomo, “Fine-Tuning IndoBERT Model for Big Five Personality Prediction from Indonesian Social Media,” in 2023 International Seminar on Intelligent Technology and Its Applications (ISITIA), IEEE, Jul. 2023, pp. 93–98. doi: 10.1109/ISITIA59021.2023.10221074.
M. L. Smith, D. Hamplová, J. Kelley, and M. D. R. Evans, “Concise survey measures for the Big Five personality traits,” Res. Soc. Stratif. Mobil., vol. 73, Jun. 2021, doi: 10.1016/j.rssm.2021.100595.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Oct. 2018, [Online]. Available: http://arxiv.org/abs/1810.04805
F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” Sep. 2021, [Online]. Available: http://arxiv.org/abs/2109.04607
F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Online, 2020, pp. 757–770. [Online]. Available: https://huggingface.co/
R. L. Vásquez and J. Ochoa-Luna, “Transformer-based Approaches for Personality Detection using the MBTI Model,” in Proceedings - 2021 47th Latin American Computing Conference, CLEI 2021, Institute of Electrical and Electronics Engineers Inc., 2021. doi: 10.1109/CLEI53233.2021.9640012.
E. Kerz, Y. Qiao, S. Zanwar, and D. Wiechmann, “Pushing on Personality Detection from Verbal Behavior: A Transformer Meets Text Contours of Psycholinguistic Features,” in Proceedings of the 12th Workshop on Computational Approaches to Subjectivity, Sentiment & Social Media Analysis, Stroudsburg, PA, USA: Association for Computational Linguistics, 2022, pp. 182–194. doi: 10.18653/v1/2022.wassa-1.17.
A. F. Hidayatullah, R. A. Apong, D. T. C. Lai, and A. Qazi, “Pre-trained language model for code-mixed text in Indonesian, Javanese, and English using transformer,” Soc. Netw. Anal. Min., vol. 15, no. 1, p. 30, Mar. 2025, doi: 10.1007/s13278-025-01444-9.
L. Hu, H. He, D. Wang, Z. Zhao, Y. Shao, and L. Nie, “LLM vs Small Model? Large Language Model Based Text Augmentation Enhanced Personality Detection Model,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, pp. 18234–18242, Mar. 2024, doi: 10.1609/aaai.v38i16.29782.
T. Parlar, S. A. Özel, and F. Song, “QER: a new feature selection method for sentiment analysis,” Human-centric Computing and Information Sciences, vol. 8, no. 1, Dec. 2018, doi: 10.1186/s13673-018-0135-8.
T. Parlar and S. A. Ozel, “A new feature selection method for sentiment analysis of Turkish reviews,” in Proceedings of the 2016 International Symposium on Inovations in Intelligent Systems and Applications, INISTA 2016, Institute of Electrical and Electronics Engineers Inc., Sep. 2016. doi: 10.1109/INISTA.2016.7571833.
P. H. Prastyo, R. Hidayat, and I. Ardiyanto, “Enhancing sentiment classification performance using hybrid Query Expansion Ranking and Binary Particle Swarm Optimization with Adaptive Inertia Weights,” ICT Express, vol. 8, no. 2, pp. 189–197, Jun. 2022, doi: 10.1016/j.icte.2021.04.009.
G. A. Pradnyana, W. Anggraeni, E. M. Yuniarno, and M. H. Purnomo, “Enhancing MBTI Personality Trait Prediction from Imbalanced Social Media Data Using Hybrid Query Expansion Ranking and Glo Ve- BiLSTM,” in 2023 IEEE International Conference on Fuzzy Systems (FUZZ), IEEE, Aug. 2023, pp. 1–6. doi: 10.1109/FUZZ52849.2023.10309718.
V. Ong, A. D. S. Rahmanto, W. Williem, N. H. Jeremy, D. Suhartono, and E. W. Andangsari, “Personality Modelling of Indonesian Twitter Users with XGBoost Based on the Five Factor Model,” International Journal of Intelligent Engineering and Systems, vol. 14, no. 2, pp. 248–261, 2021, doi: 10.22266/ijies2021.0430.22.
H. Lucky, Roslynlia, and D. Suhartono, “Towards Classification of Personality Prediction Model: A Combination of BERT Word Embedding and MLSMOTE,” in Proceedings of 2021 1st International Conference on Computer Science and Artificial Intelligence, ICCSAI 2021, Institute of Electrical and Electronics Engineers Inc., 2021, pp. 346–350. doi: 10.1109/ICCSAI53272.2021.9609750.
Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” Jul. 2019, Accessed: May 04, 2023. [Online]. Available: https://arxiv.org/abs/1907.11692v1
S. Wu and M. Dredze, “Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 833–844. doi: 10.18653/v1/D19-1077.
T. Pires, E. Schlinger, and D. Garrette, “How Multilingual is Multilingual BERT?,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4996–5001. doi: 10.18653/v1/P19-1493.