Research

Selected papers, preprints, and the code, data, and slides behind them. Filter by topic below.

Journal articles

  1. Herring, S. C., Lee, S., Wang, H., & Yan, Y. (forthcoming). Can LLMs do computer-mediated discourse analysis? Language@Internet.
    Code NLP
  2. Wang, H., Clark, J., Yan, Y., Bradley, S., Chen, R., Zhang, Y., Fu, H., & Tian, Z. (2026). Fairness evaluation of large language models in academic library reference services. Humanities and Social Sciences Communications.
  3. Wang, H., Juola, P., & Riddell, A. (2026). Reproduction and replication of an adversarial stylometry experiment. Replication Research, 2.
    Code Data Preprint NLPOpen code, data, and materials; reproducibility certificate
  4. Wang, H., Lee, J., Walsh, J. A., Flanders, J., & Lee, B. C. G. (2025). Building and deploying the Digital Humanities Quarterly recommender system. Code4Lib Journal, (61).
    Code NLP
  5. Wang, H., Clark, J., McKelvey, H., Sterman, L., Gao, Z., Tian, Z., Kübler, S., & Liu, X. (2025). Science out of its Ivory Tower: Improving accessibility with reinforcement learning. Scientometrics, 130(8), 4519–4543.
  6. Wang, H., Clark, J., McKelvey, H., Sterman, L., Gao, Z., Tian, Z., & Liu, X. (2025). Improving scholarship accessibility with reinforcement learning. Information Research, 30(iConf), 203–218.
    PDF Code NLPiConference 2025 Best Paper Finalist
  7. Xie, X., Li, J., & Wang, H. (2025). The many voices of Duying: Revisiting the disputed essays between Lu Xun and Zhou Zuoren. Digital Scholarship in the Humanities, 40(1), 338–353.
    Code NLP

Conference papers

  1. Wang, H., & Clark, J. (2024). Simplifying scholarly abstracts for accessible digital libraries using language models. Proceedings of the 2024 ACM/IEEE Joint Conference on Digital Libraries (JCDL '24).
  2. Wang, H. (2024). A content-based novelty measure for scholarly publications: A proof of concept. iConference 2024 (LNCS), 409–420.
    Code Demo NLP Metascience
  3. Wang, H., & Riddell, A. (2022). CCTAA: A reproducible corpus for Chinese authorship attribution research. Proceedings of the 13th Language Resources and Evaluation Conference (LREC 2022), 5889–5893.
    PDF Code NLP
  4. Tian, Z., Dong, X., Gao, F., Wang, H., & Lin, C. (2022). Mandarin tone sandhi realization: Evidence from large speech corpora. INTERSPEECH 2022, 5273–5277.
    PDF NLP
  5. Wang, H., Xie, X., & Riddell, A. (2021). The challenge of vernacular and classical Chinese cross-register authorship attribution. Proceedings of the Computational Humanities Research Conference (CHR 2021), 299–309.
    PDF Data NLP
  6. Riddell, A., Wang, H., & Juola, P. (2021). A call for clarity in contemporary authorship attribution evaluation. Proceedings of Recent Advances in Natural Language Processing (RANLP 2021), 1178–1183.
    Data NLP
  7. Wang, H., Riddell, A., & Juola, P. (2021). Mode effects' challenge to authorship attribution. Proceedings of the 16th Conference of the European Chapter of the ACL (EACL 2021), 1146–1155.
    PDF Poster NLP

Book chapters

  1. Wang, H., Clark, J., & Peña, A. (forthcoming). Responsible intelligence in practice: A fairness audit of open large language models for library reference services. Artificial Intelligence and Social Justice Intersections in Library and Information Studies (Emerald).
  2. Wang, H. (forthcoming). Social determinants of COVID-19 hospitalization in the All of Us Research Program: A propensity-matched case-control study. Social Vulnerability to COVID-19, 2nd ed. (Springer Nature).
    Health

Preprints

  1. Zhang, J., Fu, H., & Wang, H. (2026). The shape of NSF openness to newcomers: Same share, fewer doors. SocArXiv.
    Code Metascience
  2. PDF Metascience Health
  3. PDF NLP
  4. Wang, H. (2026). Funding the runners-up beats a golden ticket. arXiv:2609.19552.
    PDF Code Metascience
  5. PDF Code Metascience Health
  6. Wang, H. (2023). Defending against authorship identification attacks. arXiv:2310.01568.
    PDF NLP