@ehudreiter.bsky.social

New blog: AI for Healthcare: Great benchmarks but minimal impact AI in healthcare benchmarks show soaring perf, but in the real world AI is not actually helping people much. Reasons include messy real world, weak eval, deployment, and lack of utility ehudreiter.com/2026/08/05/a...

AI for Healthcare: Great benchmarks but minimal impact

The paradox of AI in healthcare is that while benchmarks show soaring (and sometimes superhuman) performance, in the real world AI is not actually helping people much. This is partially because rea…

ehudreiter.com

Somebody emailed me to ask about being a reviewer at TACL, and unfortunately this ended up in my spam folder and got deleted. Apologies! If you emailed me about this and did not get a response, please email me again

Congrats to my student Mengxuan Sun for submitting her PhD thesis! Mengxuan is joint CS/Medicine and is looking at using NLP to help cancer patients.

Somewhat scary paper showing dubious Kaggle medical data sets being used in both research papers and real-world clinical practice. Blind use of dubious data is not just a problem in NLP and AI... doi.org/10.1186/s129...

Evidence of unreliable data and poor data provenance in clinical prediction model research and clinical practice - BMC Medicine

Background Clinical prediction models are often created using large routinely collected datasets. It is essential that prediction models are developed with appropriate data and methods and transparently reported to ensure that decisions are based on reliable predictions. Kaggle is a popular competition and data repository website where users learn and apply analysis skills on a range of datasets. Methods We identified two large, publicly available Kaggle datasets, on stroke and diabetes, that lack clear data provenance, but are widely used in clinical prediction models in peer reviewed publications. We used exploratory analyses to examine the quality of data and reporting of information using nine items from the TRIPOD+AI statement checklist. Results Data provenance assessment using nine TRIPOD+AI items revealed major deficiencies, with minimal details for either dataset including no information on when, where, why or how the data were collected. The authenticity of both datasets could not be verified and have no reliable provenance of authenticity and should not be used for informing research or practice. From these two datasets, we found 125 clinical prediction model studies. Three prediction models had evidence of use in clinical practice, one model was cited in a medical device patent, and the models were cited in 86 review articles. Conclusions We recommend that journals and data repositories mandate data provenance reporting to safeguard published research. Prediction models based solely on inauthentic or unreliable datasets should never be used to directly inform decisions on patient care.

doi.org

New blog: Memories of being an NL researcher in 1990 I “reminisce ” about being an NL researcher in 1990, when I got my PhD. Community was much smaller than 2026, but in many ways it was nicer, including less pressure and a more open research culture. ehudreiter.com/2026/07/20/m...

Memories of being an NL researcher in 1990

Since I am about to retire, I decided to “reminisce ” about what it was like to be an NL researcher in 1990, when I got my PhD. The community was much smaller than 2026, but in many way…

ehudreiter.com

AI safety news UK: AISI eliminates its societal resilience team (which covered things like suicide risk from bots) China: new government regulations aim to reduce emotional dependence on AI Looks like China takes emotional risks far more seriously than UK (or US)...

As reviewer, I pointed out a fundamental misconception in paper. Authors reponded that they had seen many published ACL (etc) papers with same problem, its also embedded in benchmarks. If they cannot trust what they see published in ACL, how can they build on other peoples work? I dont have answer..

New blog: What is the purpose of ACL conferences? What is the main pupose of ACL conferencess: meeting people, enhancing CVs, identifying good papers, or providing a home for exciting science? The best reviewing system depends on the goal of our conf ehudreiter.com/2026/07/09/w...

What is the purpose of ACL conferences?

The reviewing system for ACL conferences is struggling. In order to fix it, we should be clear about what the main pupose of the conferences is: meeting people, enhancing CVs, identifying good pape…

ehudreiter.com

Congratulations to Zeerak and Dirk for their well-deserved Test of Time award on hate speech detection! Its also striking to see a Test of Time award go to a paper presented at a Student Research Workshop. Maybe SRW are more open to crazy new ideas than xACL main conferences?

I'm giving a talk to the local branch of the Royal Statistical Society next week. Looking forward to it! Statisticians have been building models for a long time, we can learn from them. They also take data issues very seriously unlike most machine learning people I know.

Are there any papers which won both Best Paper award when presented and a Test of Time award later? I cannot find any in xACL. If so, suggests BP awards do not signify long-term research impact

I am becoming an editor (EiC) at TACL (Transactions of ACL) journal. Not what I expected to be doing in retirement, but I am a strong believer in journals as the best way to present research, and I want to help TACL become an even better venue for publishing NLP research.

INLG (@inlg.bsky.social) submissions are now open! Please submit all of your work on Natural Language Generation. Submit systems, demos, experiments, position papers, squibs, linguistic analyses, as long as it is NLG-related. For more information, see: 2026.inlgmeeting.org #NLProc

INLG2026

The 19th International Natural Language Generation Conference is scheduled to be held in Utrecht, the Netherlands from October 17 to 21, 2026.

2026.inlgmeeting.org

New blog: Future of NLG evaluation In a recent position paper, I argued that NLG evaluation in the future needs to be become more rigorous. It also needs to move beyond benchmarks, and focus more on impact, qualitative, and safety evaluation. ehudreiter.com/2026/06/26/f...

Future of NLG evaluation

In a recent position paper, I argued that NLG evaluation in the future needs to be become more rigorous. It also needs to move beyond benchmarks, and focus more on impact, qualitative, and safety e…

ehudreiter.com

Really enjoyed helping my colleague Jakub Zbrzezny from our Divinity dept look at how well LLMs can translate biblical materials in a local Arabic dialect. Not surprisingly, LLMs good at translating out of dialect, but struggle to translate into dialect. aclanthology.org/2026.retroev...

The Arabic Bible as an Evaluation Tool: The Case Study of the Khalīlī Arabic Dialect

Jakub Zbrzeżny, Ehud Reiter, Wei Zhao. Proceedings of the 1st Symposium on Natural Language Generation Evaluations. 2026.

aclanthology.org

Radical suggestion. Why not *lower* prestige of xACL, for example by including Findings in main conf. Then people chasing N papers in "top" venues will submit elsewhere, making xACL more manageable. Let Neurips deal with AI slop...

Reviewed a paper which extensively used a dataset I helped to create, but showed zero awareness of data issues and the domain. I guess people just want data to throw into LLMs, dont care about data issues even when these are carefully explained in dataset paper.

[AI leads to] an expansion of individual scientists’ impact but a contraction in collective science’s reach, as AI-augmented work moves collectively towards areas richest in data... AI tools seem to automate established fields rather than explore new ones www.nature.com/articles/s41...

Artificial intelligence tools expand scientists’ impact but contract science’s focus - Nature

Artificial intelligence boosts individual scientists’ output, citations and career progression, but collectively narrows research diversity and reduces collaboration, concentrating work in data-rich a...

nature.com

New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...

I am worried by NLP research culture

In most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have…

ehudreiter.com