Research / Publications

Publications

This section includes a) a list of all CRC publications in which project S was involved, and b) bibliographies for DH-related research in the areas of the supported subprojects.

Project S: Publications Download BibTeX

Jurczyk, Thomas, Roman Seidel, Adrian Bernhard, Tatjana Scheffler, and Johann Buessow. 2025. “Text Mining Tafsir: Compilation and Preliminary Explorations of a Curated Corpus of 80 Qurʾanic Commentaries.” Journal of Digital Islamicate Research (Leiden, The Netherlands) 3 (1): 97–167. https://doi.org/10.1163/27732363-bja00010.

A01: DH / Islamic Studies Bibliography Download BibTeX

Al-Kabi, Mohammed, Heider Wahsheh, Izzat Alsmadi, and Abdallah Al-Akhras. 2015. “Extended Topical Classification of Hadith Arabic Text.” International Journal on Islamic Applications in Computer Science And Technology 3 (September): 13–24.

Alraddadi, Rawan Abdullah, and Moulay Ibrahim El-Khalil Ghembaza. 2021. “Anti-Islamic Arabic Text Categorization Using Text Mining and Sentiment Analysis Techniques.” International Journal of Advanced Computer Science and Applications 12 (8). https://doi.org/10.14569/IJACSA.2021.0120889.

Antoun, Wissam, Fady Baly, and Hazem Hajj. 2020. “AraBERT: Transformer-Based Model for Arabic Language Understanding.” In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, edited by Hend Al-Khalifa, Walid Magdy, Kareem Darwish, Tamer Elsayed, and Hamdy Mubarak. European Language Resource Association. https://aclanthology.org/2020.osact-1.2/.

Ayish, Mohammad. 2025. “Digital Humanities for Arab Media Studies Opportunities and Challenges.” International Journal of Digital Humanities 7 (2): 249–66. https://doi.org/10.1007/s42803-025-00105-9.

Ayu, Media Anugerah, Edi Irawan, and Teddy Mantoro. 2022. Text Mining Approaches for Analyzing an Indonesian Tafseer and Translation of the Holy Quran. No. 3. 25 (3): 3. https://doi.org/10.11591/ijeecs.v25.i3.pp1469-1480.

Badawy, Amro Ali. 2025. “Topic Discovery in the Digital Quran: A Text Mining Approach.” Journal of Information Systems Engineering and Management 10 (18s): 642–49. https://doi.org/10.52783/jisem.v10i18s.2976.

Bednarkiewicz, Maroussia, Aslisho Qurboniev, and Gowaart Van Den Bossche. 2023. “Studying Hadith Commentaries in the Digital Age.” In Hadith Commentary, edited by Joel Blecher and Stefanie Brinkmann. Edinburgh University Press. https://doi.org/doi:10.1515/9781474461061-014.

Bernhard, Adrian. 2024. Development of the Quranic WAY Metaphor in Tafsir: A Corpus Analysis Approach. January 1.

Elaziz, Mohamed Abd, Mohammed A. A. Al-qaness, Ahmed A. Ewees, and Abdelghani Dagou, eds. 2020. Recent Advances in NLP: The Case of Arabic Language. Studies in Computational Intelligence, volume 874. Springer.

Guellil, Imane, Houda Saâdane, Faical Azouaou, Billel Gueni, and Damien Nouvel. 2021. “Arabic Natural Language Processing: An Overview.” Journal of King Saud University - Computer and Information Sciences 33 (5): 497–507. https://doi.org/10.1016/j.jksuci.2019.02.006.

Jurczyk, Thomas, Roman Seidel, Adrian Bernhard, Tatjana Scheffler, and Johann Buessow. 2025. “Text Mining Tafsir: Compilation and Preliminary Explorations of a Curated Corpus of 80 Qurʾanic Commentaries.” Journal of Digital Islamicate Research (Leiden, The Netherlands) 3 (1): 97–167. https://doi.org/10.1163/27732363-bja00010.

Karam, Rimane. 2026. Is Arabic Really a Well-Resourced Language? Digital Exploration of a Corpus in Middle Arabic, a Family of Varieties Omitted by Arabic Natural Language Processing. Text/html,application/pdf. https://doi.org/10.60693/67A7-WB88.

Kelly, Tynan. 2025. “Detecting Text Reuse in Historical Arabic Texts: Challenges and Strategies.” Journal of Digital Islamicate Research 3 (2): 362–95. https://doi.org/10.1163/27732363-bja00013.

Khondaker, Md Tawkat Islam, Abdul Waheed, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023. “GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP.” In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, edited by Houda Bouamor, Juan Pino, and Kalika Bali. Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.16.

Kirmizialtin, Suphan, and David Joseph Wrisley. 2024. “Exploring Gulf Manumission Documents with Word Vectors.” Journal of Digital Islamicate Research 2 (1–2): 1–29. https://doi.org/10.1163/27732363-bja00005.

“KITAB: Text Reuse.” n.d. https://kitab-project.org/methods/text-reuse.

Lachkar, Abdelmonaime, Karim Bouzoubaa, Azzedine Mazroui, Abdelfettah Hamdani, and Abdelhak Lekhouaja, ed. 2018. Arabic Language Processing: From Theory to Practice. Vol. 782. Communications in Computer and Information Science. Springer International Publishing. https://doi.org/10.1007/978-3-319-73500-9.

Lange, Christian, Maksim Abdul Latif, Yusuf Çelik, A. Melle Lyklema, Dafne E. van Kuppevelt, and Janneke van der Zwaan. 2021. “Text Mining Islamic Law.” Islamic Law and Society 28 (3): 234–81. https://doi.org/10.1163/15685195-bja10009.

Mosa, Mohamed Atef. 2025. “Synergizing Structure and Semantics: A Knowledge Graph-Transformer Framework for Narrator Disambiguation in Hadith Networks.” Digital Scholarship in the Humanities 40 (4): 1085–100. https://doi.org/10.1093/llc/fqaf088.

Nigst, Lorenz, Maxim Romanov, Sarah Bowen Savant, Masoumeh Seydi, Peter Verkinderen, and Hamidreza Hakimi. 2023. “OpenITI: A Machine-Readable Corpus of Islamicate Texts.” Version 2023.1.8. Zenodo, October 17. https://doi.org/10.5281/ZENODO.10007820.

Obeid, Ossama, Nasser Zalmout, Salam Khalifa, et al. 2020. “CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing.” Proceedings of the Twelfth Language Resources and Evaluation Conference, May, 7022–32. https://aclanthology.org/2020.lrec-1.868.

Sabbeh, Sahar F., and Heba A. Fasihuddin. 2023. “A Comparative Analysis of Word Embedding and Deep Learning for Arabic Sentiment Classification.” Electronics 12 (6): 1425. https://doi.org/10.3390/electronics12061425.

Sara, Zanotta. 2026. Persian-Arab Life Trajectories in the Red Sea (19th-20th Centuries): Tracing Mixedness through Digital Humanities. Text/html,application/pdf. https://doi.org/10.60693/9RCJ-RJ41.

Shahid, Usama, Muhammad Zunnurain Hussain, and William Sayers. 2025. “Computational Analysis of Quran Text Using Machine Learning and Large Language Models.” 2025 8th International Conference on Data Science and Machine Learning Applications (CDMA), February, 18–24. https://doi.org/10.1109/CDMA61895.2025.00009.

Soliman, Abu Bakr, Kareem Eissa, and Samhaa R. El-Beltagy. 2017. “AraVec: A Set of Arabic Word Embedding Models for Use in Arabic NLP.” Procedia Computer Science 117: 256–65. https://doi.org/10.1016/j.procs.2017.10.117.

Wardini, Elie. 2022. “Arabic Computational Linguistics: Potential, Pitfalls and Challenges.” In Natural Language Processing in Artificial Intelligence — NLPinAI 2021, edited by Roussanka Loukanova, vol. 999. Studies in Computational Intelligence. Springer International Publishing. https://doi.org/10.1007/978-3-030-90138-7_4.

A03: DH / Tibetan Studies Bibliography Download BibTeX

Bingenheimer, Marcus, Jen-Jou Hung, and Cheng-en Hsieh. 2017. “Stylometric Analysis of Chinese Buddhist Texts - Do Different Chinese Translations of the Gaṇḍavyūha Reflect Stylistic Features That Are Typical for Their Age?” Journal of the Japanese Association for Digital Humanities 2 (1): 1–30. https://doi.org/10.17928/jjadh.2.1_1.

Erhard, Franz Xaver. 2025. “Text and Information Extraction for Modern (Post-1950) Tibetan Newspapers and Print Publications with Transkribus.” Revue d’Etudes Tibétaines, no. 74 (February): 128–71.

Faggionato, Christian. 2024. “A Universal Dependency Treebank for Classical Tibetan.” Revue d’Etudes Tibétaines, no. 72 (July): 52–69.

Griffiths, Rachael. 2024. “Handwritten Text Recognition (HTR) for Tibetan Manuscripts in Cursive Script.” Revue d’Etudes Tibétaines, no. 72 (July): 43–51.

Huang, Cheng, Fan Gao, Nyima Tashi, et al. 2026. “TFD: A Comprehensive Structured Tibetan Foundation Dataset for Low-Resource Language Processing and Large-Scale Modeling.” arXiv:2503.18288. Preprint, arXiv, February 14. https://doi.org/10.48550/arXiv.2503.18288.

Huang, Cheng, Fan Gao, Nyima Tashi, Yutong Liu, et al. 2024. “Sun-Shine: A Large Language Model for Tibetan Culture.” arXiv Preprint arXiv:2407.10671.

Huang, Cheng, Nyima Tashi, Fan Gao, et al. 2025. “Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges.” Version 1. Preprint, arXiv. https://doi.org/10.48550/ARXIV.2510.19144.

Kinadeter, Michael. 2026. Analyzing, Comparing and Mapping Zen Buddhist Genealogies – Facilitating Research in Buddhist Studies by Creating a Database. Text/html,application/pdf. https://doi.org/10.60693/TV2D-AJ54.

Krishna, Ravi, Norman Mu, and Kurt Keutzer. 2021. “Applying Text Analytics to the Mind-Section Literature of the Tibetan Tradition of the Great Perfection.” ACM Transactions on Asian and Low-Resource Language Information Processing 20 (2): 1–32. https://doi.org/10.1145/3392047.

Kyogoku, Yuki, Franz Xaver Erhard, Robert Barnett, and Nathan Hill. 2024. “Basic Modern Tibetan SpaCy Model.” Version 0.1.2. Preprint, Zenodo, December 15. https://doi.org/10.5281/ZENODO.14494472.

Li, Yan, Xiaomin Li, Yiru Wang, Hui Lv, Fenfang Li, and La Duo. 2022. “Character-Based Joint Word Segmentation and Part-of-Speech Tagging for Tibetan Based on Deep Learning.” ACM Transactions on Asian and Low-Resource Language Information Processing 21 (5): 1–15. https://doi.org/10.1145/3511600.

Liu, Fei-Fei, and Zhi-Juan Wang. 2018. “Active Learning for Tibetan Named Entity Recognition Based on CRF.” In Proceedings of the LREC 2018 Workshop MLP–MoMent, edited by Jinhua Du and Mihael Arcan.

Luo, Queenie, and Leonard W. J. van der Kuijp. 2024. “Norbu Ketaka: Auto-Correcting BDRC’s E-Text Corpora Using Natural Language Processing and Computer Vision Methods.” Revue d’Etudes Tibétaines, no. 72 (July): 26–42.

Meelen, Marieke, and Nathan Hill. 2018. “Segmenting and POS Tagging Classical Tibetan Using a Memory-Based Tagger.” Himalayan Linguistics 16 (2). https://doi.org/10.5070/H916234501.

Meelen, Marieke, Sebastian Nehrdich, and Kurt Keutzer. 2024. “Breakthroughs in Tibetan NLP & Digital Humanities.” Revue d’Etudes Tibétaines, no. 72 (July): 5–25.

Meelen, Marieke, and Élie Roux. 2020. “The Annotated Corpus of Classical Tibetan (ACTib) - Version 2.0 (Segmented & POS-Tagged).” Zenodo, May 4. https://doi.org/10.5281/ZENODO.3951503.

Meelen, Marieke, Élie Roux, and Nathan Hill. 2021. “Optimisation of the Largest Annotated Tibetan Corpus Combining Rule-Based, Memory-Based, and Deep-Learning Methods.” ACM Trans. Asian Low-Resour. Lang. Inf. Process. (New York, NY, USA) 20 (1). https://doi.org/10.1145/3409488.

Nehrdich, Sebastian. 2023. “Observations on the Intertextuality of Selected Abhidharma Texts Preserved in Chinese Translation.” Religions 14 (7): 911. https://doi.org/10.3390/rel14070911.

Nehrdich, Sebastian, Oliver Hellwig, and Kurt Keutzer. 2024. “One Model Is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks.” In Findings of the Association for Computational Linguistics: EMNLP 2024, edited by Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.805.

Nehrdich, Sebastian, and Kurt Keutzer. 2026. “MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan.” arXiv:2601.06400. Preprint, arXiv, January 10. https://doi.org/10.48550/arXiv.2601.06400.

Pan, Leiyu, Bojian Xiong, Lei Yang, et al. 2025. “Advancing Large Language Models for Tibetan with Curated Data and Continual Pre-Training.” arXiv:2507.09205. Preprint, arXiv, July 28. https://doi.org/10.48550/arXiv.2507.09205.

Schmidt, Dirk. 2024. “NLP for Readability, Graded Literature, & Materials Development in Tibetan.” Revue d’Etudes Tibétaines, no. 72 (July): 70–85.

Zhang, Jin, Ziyue Zhang, Lobsang Yeshi, et al. 2025. “Tibetan Medical Named Entity Recognition Based on Syllable‐Word‐Sentence Embedding Transformer.” CAAI Transactions on Intelligence Technology 10 (4): 1148–58. https://doi.org/10.1049/cit2.70029.

A06: DH / Film & Images Bibliography Download BibTeX

Arnold, Taylor, and Lauren Tilton. 2019. “Distant Viewing: Analyzing Large Visual Corpora.” Digital Scholarship in the Humanities 34 (Supplement_1): i3–16. https://doi.org/10.1093/llc/fqz013.

Arnold, Taylor, and Lauren Tilton. 2023. Distant Viewing: Computational Exploration of Digital Images. The MIT Press. https://doi.org/10.7551/mitpress/14046.001.0001.

Arnold, Taylor, Lauren Tilton, and Annie Berke. 2019. “Visual Style in Two Network Era Sitcoms.” Journal of Cultural Analytics 4 (2): 1194. https://doi.org/10.22148/16.043.

Ayyadevara, V. Kishore. 2024. Modern Computer Vision with PyTorch: A Practical Roadmap from Deep Learning Fundamentals to Advanced Applications and Generative AI. 1st ed. With Yeshwanth Reddy. Packt Publishing Limited.

Bamman, David, Rachael Samberg, Richard Jean So, and Naitian Zhou. 2024. “Measuring Diversity in Hollywood through the Large-Scale Computational Analysis of Film.” Proceedings of the National Academy of Sciences 121 (46): e2409770121. https://doi.org/10.1073/pnas.2409770121.

Burges, Joel, Nora Dimmock, and Joshua Romphf. 2016. “Collective Reading: Shot Analysis and Data Visualization in the Digital Humanities.” Cinema Journal Teaching Dossier 3 (3). https://teachingmedia.org/collective-reading-shot-analysis-and-data-visualization-in-the-digital-humanities/.

Burghardt, Manuel, Michael Kao, and Christian Wolff. 2016. “Beyond Shot Lengths – Using Language Data and Color Information as Additional Parameters for Quantitative Movie Analysis.” Digital Humanities 2016: Conference Abstracts (Kraków), 753–55. https://www.researchgate.net/publication/305277308_Beyond_Shot_Lengths_-_Using_Language_Data_and_Color_Information_as_Additional_Parameters_for_Quantitative_Movie_Analysis.

Butler, Jeremy. 2014. “Statistical Analysis of Television Style: What Can Numbers Tell Us about TV Editing?” Cinema Journal 54 (1): 25–44. https://doi.org/10.1353/cj.2014.0066.

Chen, Daoling, and Pengpeng Cheng. 2025. “Digital Extraction and Segmentation of Intangible Cultural Heritage Paper-Cut Patterns.” Digital Scholarship in the Humanities, December 16, fqaf006. https://doi.org/10.1093/llc/fqaf006.

Cutting, James E., Kaitlin L. Brunick, Jordan E. DeLong, Catalina Iricinschi, and Ayse Candan. 2011. “Quicker, Faster, Darker: Changes in Hollywood Film over 75 Years.” I-Perception 2 (6): 569–76. https://doi.org/10.1068/i0441aap.

Dancygier, Barbara. 2016. “Multimodality and Theatre: Material Objects, Bodies and Language.” In Theatre, Performance, and Cognition: Languages, Bodies and Ecologies, edited by Rhonda Blair and Amy Cook. Bloomsbury Methuen.

Ferguson, Kevin L. 2016. “Digital Surrealism: Visualizing Walt Disney Animation Studios.” Digital Humanities Quarterly 11 (1). https://doi.org/10.63744/xtpq8gu3cfct.

Fischer, Norbert, Dominik Kimmel, and Frank Puppe. 2025. “Semiautomatische Erschließung von Fotografien auf beschrifteten Bildkarten im Archiv. Dokumentenerkennung mit Deep Learning sowie Large-Language-Modellen.” Zeitschrift für digitale Geisteswissenschaften 10. https://doi.org/10.17175/2025_009.

Gonzalez, Rafael C., and Richard E. Woods. 2018. Digital Image Processing. Fourth, Global edition. Pearson Education.

Löffler, Beate. 2026. “Irgendwas in Japan. Erfahrungswerte zur Bildsuche und Bilderkennung in der architekturhistorischen Forschung.” Bildähnlichkeit und Bildsuche: Geistes- und informationswissenschaftliche Zugänge zu historischem Material Sonderband 8. https://doi.org/10.17175/SB008_007.

Nantke, Julia, Johannes Leitgeb, and Christian Reul. 2026. “Quantitative Analyse visueller Charakteristika von Briefen.” Zeitschrift für digitale Geisteswissenschaften 11. https://doi.org/10.17175/2026_011.

Otto, Ulf. 2026. “Bodies in Relation. Theater Photography and the Distant Viewing of Performance.” Bildähnlichkeit Und Bildsuche: Geistes- Und Informationswissenschaftliche Zugänge Zu Historischem Material Sonderband 8. https://doi.org/10.17175/SB008_006.

Salt, Barry. 1974. “Statistical Style Analysis of Motion Pictures.” Film Quarterly 28 (1): 13–22. https://doi.org/10.2307/1211438.

Schmidt, Thomas, Alina El-Keilany, Johannes Eger, and Sarah Kurek. 2021. Exploring Computer Vision for Film Analysis: A Case Study for Five Canonical Movies. https://doi.org/10.5283/EPUB.50867.

Smits, Thomas, and Melvin Wevers. 2023. “A Multimodal Turn in Digital Humanities. Using Contrastive Machine Learning Models to Explore, Enrich, and Analyze Digital Visual Historical Collections.” Digital Scholarship in the Humanities 38 (3): 1267–80. https://doi.org/10.1093/llc/fqad008.

Somandepalli, Krishna, Tanaya Guha, Victor R. Martinez, Naveen Kumar, Hartwig Adam, and Shrikanth Narayanan. 2021. “Computational Media Intelligence: Human-Centered Machine Analysis of Media.” Proceedings of the IEEE 109 (5): 891–910. https://doi.org/10.1109/JPROC.2020.3047978.

Soriano-Gonzalez, Laura, and Jose Belda-Medina. 2025. “Exploring Image–Text Combinations in Visual Humour through Large Language Models (LLMs).” Digital Scholarship in the Humanities 40 (1): 280–94. https://doi.org/10.1093/llc/fqae068.

B02: DH / Korean Studies Bibliography tba