Hi 👋, I'm a postdoctoral fellow in the Department of Biostatistics and Health Data Science at Indiana University School of Medicine. I do NLP with real-world data. My research explores cultural and sociotechnical dynamics to help build better AI for everyone 🤗.

My representative work includes:


What's New

  • (Sep 13, 2026) My new preprint "Ending the NIH Embargo Accelerated Public Access to Funded Research, But Substantial Delays Remain" is out: the 2025 mandate roughly doubled one-week PubMed Central availability for NIH-funded articles in non-DOAJ journals (11% to 22%), but most of them still weren't there a week after publication.

  • (Aug 14, 2026) Two papers are out! 🎉

  • (Jul 21, 2026) I gave an invited lightning talk, "Can AI Serve Everyone Equally? Auditing Fairness in LLM-Powered Reference," at the AI Literacy & Learning for Libraries, Archives, and Museums (LAMs) Summit, Library of Congress, Washington, DC.

  • (May 27, 2026) Our half-day short course, "Practical Large Language Models: Foundations and Applications," was successfully delivered at MBSW 2026 in Carmel, Indiana, together with Drs. Jiang Bian, Xing He, and Yuhang Jiang.

  • (Mar 20, 2026) DualR is live! 💊 Our clinical risk assessment tool is now available at dualr.hainingwang.org. DualR asks a simple question: given that a patient took this drug, how likely are they to have this disease? and turns the answer into a portable risk score from any medication list. In our study across the All of Us Research Program and Indiana's largest clinical data network (N > 1.3M), DualR flagged future alcohol use disorder 300+ days before diagnosis — all with equity across race, gender, and social determinants of health.

  • (Nov 17, 2025) Our poster, "Thinking, Fast and Slow: DualReasoning Enhances Clinical Knowledge Extraction from Large Language Models," was nominated as a distinguished poster at AMIA 2025!

  • (Nov 8, 2025) Our short course, "Practical Large Language Models: Foundations and Applications," has been accepted by MBSW 2026!

  • (Oct 21, 2025) Our paper "Building and Deploying the Digital Humanities Quarterly Recommender System" is out in Code4Lib Journal! 🎉 As the Data Analytics team at DHQ, we open-sourced DHQ's recommender system. With DHQ's corpus (also fully open) and self-explanatory code, this serves a good starting point for CS/DS/LIS students interested in building their first retrieval system.

  • (Aug 18, 2025) Excited to share that our ACCESS Explore request, "Practice-Informed, Explainable Multi-Agent Decision Support for Immune-Related Acute Kidney Injury (irAKI) Phenotyping" (MED250055), has been approved!

  • (Aug 3, 2025) I gave an invited talk, "Thinking, Fast and Slow: DualReasoning Enhances Clinical Knowledge Extraction from Large Language Models", at the International Conference on Intelligent Biology and Medicine (ICIBM 2025).

  • (Jul 19, 2025) Our paper "Science Out of Its Ivory Tower: Improving Accessibility with Reinforcement Learning" is out in Scientometrics! 🎉

  • (Jul 8, 2025) Our new preprint "Fairness Evaluation of Large Language Models in Academic Library Reference Services" is out! TL;DR: Current LLMs show a promising degree of readiness to support equitable and contextually appropriate communication in academic library reference services.

  • (May 20, 2025) I presented "Thinking, Fast and Slow: Knowledge Extraction to Facilitate Phenotyping Using Drug Records in Real-World Data" at the Midwest Biopharmaceutical Statistics Workshop (MBSW 2025). Check out our DualReasoning, which enables effective and secure use of medication records through simple chat with LLMs.

  • (May 9, 2025) I successfully defended my PhD dissertation, Defending Against Authorship Attribution Attacks with Large Language Models. I am deeply grateful to my committee members: Dr. Allen Riddell (Chair), Dr. John Walsh, Dr. Kahyun Choi, Dr. Staša Milojević, and Dr. Xiaozhong Liu. I also extend my sincere thanks to Dr. Sandra Küber for her guidance. Check out the slides and manuscript.

  • (Mar 20, 2025) I presented "Improving Scholarship Accessibility With Reinforcement Learning" at iConference 2025, hosted at Indiana University Bloomington. Our jargon-busting AI paper became a Best Paper finalist—give it a read!

  • (Feb 12, 2025) I presented a poster titled "Thinking, Fast and Slow: DualReasoning Enhances Clinical Knowledge Extraction from Large Language Models" at the Regenstrief Healthcare AI Conference*.
    DualReasoning shows great promise in incorporating messy medication information for chronic disease phenotyping (e.g., diabetes and hypertension) while remaining fully privacy-preserving.

  • (Dec 27, 2024) Our DSH paper on solving authorship disputes between Lu Xun and Zhou Zuoren's early works is out—give it a read!

  • (Dec 18, 2024) I presented "Simplifying Scholarly Abstracts for Accessible Digital Libraries Using Language Models" at JCDL 2024.

  • (Oct 24, 2024) Our new preprint is out! "Science Out of Its Ivory Tower: Improving Accessibility with Reinforcement Learning" simplifies scholarly abstracts from postgraduate 🧑‍🎓 to high-school 🧑‍🏫 reading levels without sacrificing faithfulness or quality. All of this is done by RL-tuning a Gemma-2B.

  • (Oct 15, 2024) I presented "Language Modeling: From Teaching to Incentivizing" at the IU Computational Linguistics Discussion Group.

  • (Sep 28, 2024) I presented "Science Out of Its Ivory Tower: Improving Accessibility with Reinforcement Learning" at the ILS Doctoral Research Forum 2024 and won 🥇.

  • (Aug 7, 2024) Our new preprint "Simplifying Scholarly Abstracts for Accessible Digital Libraries" is now online. We introduced a novel corpus designed to simplify scholarly abstracts and demonstrated that mainstream LLMs perform well with straightforward supervised fine-tuning. Although the improvement doesn't yet make the content fully understandable for a middle-school audience, these models provide a strong baseline for further enhancement.

  • (May 1, 2024) I joined Dr. Jing Su's lab as a research intern on the GAIPA project (Graph Artificial Intelligence for Precision Identification of Alcohol Use Disorder) in the Department of Biostatistics and Health Data Science at Indiana University School of Medicine.

  • (Apr 30, 2024) I passed my proposal defense for my PhD dissertation.

  • (Apr 15, 2024) I presented the manuscript titled "A Content-Based Novelty Measure for Scholarly Publications: A Proof of Concept" at iConference 2024.

  • (Mar 13, 2024) I delivered a presentation titled "AI4Library" at Nankai University, covering the fundamentals of LLMs, their capabilities, surrounding hype, and their applications within librarianship.

  • (Jan 31, 2024) I joined the Center for Rare Book Conservation and Restoration Research at Wuhan University as a Research Affiliate.

  • (Jan 29, 2024) I joined Digital Humanities Quarterly as Data Analytics Editor.

  • (Jan 18, 2024) My invited column "Defending Against Authorship Identification Attacks" was published by the Montreal AI Ethics Institute.

  • (Jan 8, 2024) The manuscript for NovEval, "A Content-Based Novelty Measure for Scholarly Publications: A Proof of Concept", is up. I will be presenting it at iConference 2024.

  • (Oct 7, 2023) NovEval (pronounced as "Nawv-Ee-val") demo is now online! This GPT-2-based model evaluates scientific novelty automatically, and its assessments align well with human evaluation. It's still in beta, and I would love to hear your feedback😉!

  • (Oct 5, 2023) I successfully completed my qualifying defense🎉. The committee consisted of Dr. Allen Riddell (Chair), Dr. Xiaozhong Liu, and Dr. Staša Milojević (Minor Advisor).

  • (Oct 4, 2023) Our new papers are available on arXiv. Check them out!

  • (May 19, 2023) I delivered a lightning talk introducing our jargon-busting AI at LEADING Forum 2022. You can access our poster titled "Science Out of the Ivory Tower: Scientific Abstract Simplification for Everyone" here.

  • (May 10, 2023) I uploaded a LaTeX poster template to Overleaf. The template, based on Gemini, is minimal and modern, and features Indiana University's official color palette. You can find the template here, or simply search for "IU poster" in the Overleaf gallery.

  • (Apr 14, 2023) I gave a lightning talk and presented a poster titled "The Many Voices of the Detached: Revisiting the Disputed Writings of Lu Xun and Zhou Zuoren" at the IDAH HASTAC Symposium 2023.

  • (Apr 13, 2023) I gave a guest lecture titled "Authorship Attribution: An Introduction" for Allen's Digital Humanities course.

  • (Jan 9, 2023) I began working with Allen on the IARPA HIATUS (Human Interpretable Attribution of Text Using Underlying Structure) Task Three.

  • (Dec 12, 2022) Blog-1K was published on Zenodo. It is a redistributable authorship identification testbed for contemporary English prose. As a midterm output of HASTAC, we will use it to benchmark a deep learning-based authorship verification model against a corresponding attribution model.

  • (Nov 19, 2022) We launched our AI on Hugging Face. It rewrites complex scientific abstracts into simpler yet accurate versions, helping lay audiences enjoy the fruits of open science.

  • (Nov 14, 2022) I gave a guest lecture titled "Modern Stylometry: Theory & Practice" at Wuhan University.

  • (Nov 4, 2022) I presented "Protecting Author Identity with Bible-Reading Artificial Intelligence" and won 🥈 in the ILS Doctoral Research Forum 2022.

  • (Aug 22, 2022) I worked with Dr. Kahyun Choi as an Associate Instructor in her Music Data Mining course that fall.

  • (Aug 15, 2022) Our new paper "Reproduction and Replication of an Adversarial Stylometry Experiment" was released on arXiv.

  • (May 4, 2022) I was accepted into the 2022–2023 Institute for Digital Arts & Humanities (IDAH) Humanities, Arts, Science, and Technology Alliance and Collaboratory (HASTAC) Scholarship program and began working on Leveraging Small Humanities Datasets With Few-Shot Learning: A Case Study in Stylometry.

  • (May 4, 2022) The package writeprints-static was released on PyPI. Check it out!

  • (May 1, 2022) Scripts for generating the Chinese Cross-Topic Authorship Attribution (CCTAA) Corpus and running baselines were hosted on Codeberg.

  • (Apr 14, 2022) I gave a guest lecture titled "Authorship Attribution: A Case Study with Classical Chinese" at ILS Z-764/657 Digital Humanities, Indiana University Bloomington.

  • (Apr 13, 2022) I was accepted as a LIS Education and Data Science Integrated Network Group (LEADING) Fellow and worked with Montana State University Library on the project "TL;DR it": Automating Article Synopses for Search Engine Optimization and Citizen Science.

  • (Apr 4, 2022) Our paper "CCTAA: A Reproducible Corpus for Chinese Authorship Attribution Research" was accepted by the 13th European Language Resources Association's Language Resources and Evaluation Conference (LREC 2022).

  • (Jun 28, 2022) I presented our paper The Many Voices of Du Ying: Revisiting the Disputed Writings of Lu Xun and Zhou Zuoren at DH 2022. The Book of Abstracts is available here. Watch the presentation if registered.

  • (Jun 14, 2022) Our paper "Mandarin Tone Sandhi Realization: Evidence from Large Speech Corpora" was accepted by the 23rd INTERSPEECH Conference (INTERSPEECH 2022).

  • (Mar 4, 2022) Our papers "Minimum Text Length for Chinese Authorship Attribution" and "The Many Voices of Du Ying: Revisiting the Disputed Writings of Lu Xun and Zhou Zuoren" were accepted by DH Unbound 2022 and DH 2022, respectively.

  • (Dec 22, 2021) I gave a guest lecture titled "Authorship Attribution: Theory & Practice" at Shanghai Normal University.

  • (Nov 19, 2021) I presented our study "The Challenge of Vernacular and Classical Chinese Cross-Register Authorship Attribution" at CHR 2021. Check out the Proceedings of the Conference on Computational Humanities Research 2021.

  • (Nov 12, 2021) Katherine Morrison, Meredith Dedema, and I organized the ILS Doctoral Research Forum 2021. I won 🥉 in the contest with my presentation "The Challenge of Vernacular and Classical Chinese Cross-Register Authorship Attribution."

  • (Nov 1, 2021) I gave a guest lecture on modern stylometry, authorship attribution, and adversarial stylometry at Soochow University.

  • (Oct 13, 2021) I gave a guest lecture on adversarial stylometry in Allen's Social Media Mining course.

  • (Sep 16, 2021) The Cross-Register Authorship Attribution Corpus (v1.0) was released on Zenodo.

  • (Sep 3, 2021) Our paper "The Challenge of Vernacular and Classical Chinese Cross-Register Authorship Attribution" was accepted by the Second Conference on Computational Humanities Research (CHR 2021).

  • (Sep 1, 2021) I presented our paper "A Call for Clarity in Contemporary Authorship Attribution Evaluation" at RANLP 2021 via Zoom.

  • (Aug 17, 2021) The Reproducible Authorship Attribution Benchmark Tasks (RAABT v1.0) was released on Zenodo.

  • (Jun 3, 2021) We presented our paper "Cross-Register Authorship Attribution Using Vernacular and Classical Chinese Texts" at DH Benelux 2021.

  • (May 17–18, 2021) I gave talks at Nankai University and Shanghai Normal University on the topic of modern stylometry and authorship attribution.

  • (Apr 23, 2021) I presented our paper at EACL 2021. We also made a poster.

  • (Apr 16, 2021) Our long abstract "Cross-Register Authorship Attribution Using Vernacular and Classical Chinese Texts" was accepted by the Digital Humanities Benelux 2021 (DH Benelux 2021).

  • (Jan 11, 2021) Our long paper "Mode Effects' Challenge to Authorship Attribution" was accepted by the 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2021).

  • (Nov 16, 2020) I gave a talk at Clingding@IU about the progress we made on the paper "Mode Effects' Challenge to Authorship Attribution."