Research
Selected papers, preprints, and the code, data, and slides behind them. Filter by topic below.
Journal articles
-
Herring, S. C., Lee, S., Wang, H., & Yan, Y. (forthcoming). Can LLMs do computer-mediated discourse analysis? Language@Internet.
-
Wang, H., Clark, J., Yan, Y., Bradley, S., Chen, R., Zhang, Y., Fu, H., & Tian, Z. (2026). Fairness evaluation of large language models in academic library reference services. Humanities and Social Sciences Communications.
-
Wang, H., Juola, P., & Riddell, A. (2026). Reproduction and replication of an adversarial stylometry experiment. Replication Research, 2.
-
Wang, H., Lee, J., Walsh, J. A., Flanders, J., & Lee, B. C. G. (2025). Building and deploying the Digital Humanities Quarterly recommender system. Code4Lib Journal, (61).
-
Wang, H., Clark, J., McKelvey, H., Sterman, L., Gao, Z., Tian, Z., Kübler, S., & Liu, X. (2025). Science out of its Ivory Tower: Improving accessibility with reinforcement learning. Scientometrics, 130(8), 4519–4543.
-
Wang, H., Clark, J., McKelvey, H., Sterman, L., Gao, Z., Tian, Z., & Liu, X. (2025). Improving scholarship accessibility with reinforcement learning. Information Research, 30(iConf), 203–218.
-
Xie, X., Li, J., & Wang, H. (2025). The many voices of Duying: Revisiting the disputed essays between Lu Xun and Zhou Zuoren. Digital Scholarship in the Humanities, 40(1), 338–353.
Conference papers
-
Wang, H., & Clark, J. (2024). Simplifying scholarly abstracts for accessible digital libraries using language models. Proceedings of the 2024 ACM/IEEE Joint Conference on Digital Libraries (JCDL '24).
-
Wang, H. (2024). A content-based novelty measure for scholarly publications: A proof of concept. iConference 2024 (LNCS), 409–420.
-
Wang, H., & Riddell, A. (2022). CCTAA: A reproducible corpus for Chinese authorship attribution research. Proceedings of the 13th Language Resources and Evaluation Conference (LREC 2022), 5889–5893.
-
Tian, Z., Dong, X., Gao, F., Wang, H., & Lin, C. (2022). Mandarin tone sandhi realization: Evidence from large speech corpora. INTERSPEECH 2022, 5273–5277.
-
Wang, H., Xie, X., & Riddell, A. (2021). The challenge of vernacular and classical Chinese cross-register authorship attribution. Proceedings of the Computational Humanities Research Conference (CHR 2021), 299–309.
-
Riddell, A., Wang, H., & Juola, P. (2021). A call for clarity in contemporary authorship attribution evaluation. Proceedings of Recent Advances in Natural Language Processing (RANLP 2021), 1178–1183.
-
Wang, H., Riddell, A., & Juola, P. (2021). Mode effects' challenge to authorship attribution. Proceedings of the 16th Conference of the European Chapter of the ACL (EACL 2021), 1146–1155.
Book chapters
-
Wang, H., Clark, J., & Peña, A. (forthcoming). Responsible intelligence in practice: A fairness audit of open large language models for library reference services. Artificial Intelligence and Social Justice Intersections in Library and Information Studies (Emerald).
-
Wang, H. (forthcoming). Social determinants of COVID-19 hospitalization in the All of Us Research Program: A propensity-matched case-control study. Social Vulnerability to COVID-19, 2nd ed. (Springer Nature).
Preprints
-
Zhang, J., Fu, H., & Wang, H. (2026). The shape of NSF openness to newcomers: Same share, fewer doors. SocArXiv.
-
Xie, F., Zhang, J., & Wang, H. (2026). Still funded, no longer counted: How NIH's 2025 award reviews changed what the government counts as minority health research. arXiv:2610.03443.
-
Wang, H. (2026). Authorship identification under domain shift: A survey of stylistic measures and learned author representations. arXiv:2310.00436.
-
Wang, H. (2026). Funding the runners-up beats a golden ticket. arXiv:2609.19552.
-
Wang, H. (2026). Ending the NIH embargo accelerated public access to funded research, but substantial delays remain. arXiv:2609.14842.
-
Wang, H. (2023). Defending against authorship identification attacks. arXiv:2310.01568.