Marzena Karpinska

@markar.bsky.social

#nlp researcher interested in evaluation including: multilingual models, long-form input/output, processing/generation of creative texts previous: postdoc @ umass_nlp phd from utokyo https://marzenakrp.github.io/

Regular reminder that we have an alt-ARR slack workspace where ACs and SACs can support each other through the sometimes confusing process of the ARR cycle! Post or DM me a good email address for a Slack invitation and I will add you. #EMNLP2026

this checkbox at #arr really seems like a get-out-of-jail-free card; if we really need to allow for lang edits, why not require the reviewers to submit their pre-GPTed draft along with the edited one? 😭😭😭

Bild

Peer review was one of the most-discussed topics at #ACL2026 . Many folks were concerned about the incredible growth in the number of ARR submissions (17K for the May26 ARR cycle 😱), and even more shocked that ~40% didn't have any authors qualified to review. What is going on?? I did some digging...

hot take: don’t even use it for polishing your reviews. as an AC and SAC, the last thing i care about is your grammar, spelling, or formatting. i get the presentation issues with papers, where you might face reviewers’ language biases, but you really don’t need to worry about this for reviews.

Marzena Karpinska@markar.bsky.social · 4w ago

I think I will be posting this after each #ARR cycle: Please 🥺🙏 let's prohibit AI review writing. Otherwise, we will get lazy reviewers claiming they wrote bullet points and used AI only to put that in prose.

I think I will be posting this after each #ARR cycle: Please 🥺🙏 let's prohibit AI review writing. Otherwise, we will get lazy reviewers claiming they wrote bullet points and used AI only to put that in prose.

5 years later MTurk is on its way out. A lot has improved in open-ended gen eval, but it still suffers from underreporting, lack of statistical analysis, and sometimes sloppy design. Perhaps authors should always do their own tasks to understand the implications of how they designed evals.

Bild
Ethan Mollick@emollick.bsky.social · 4w ago

Unsurprising but still big: MTurk is on its way out, killed by AI. Mechanical Turk was a mainstay of social & survey research through the 2010s, as it allowed you to quickly buy access to many representative humans. It was pretty good at it, until LLMs came along and everyone started using AI

In the spirit of #paperoverflow looking for emergency reviewers: - aspect-based summarization - hallucination detection - evaluation invariance - evaluation, evaluation bias - fine-grained evaluation for some, the reviewers filled in delay but became unresponsive later HELP, this house is on fire 🔥

It's invaluable in good literature to have someone who has actual, lived experience and a deep cultural and historical understanding of the literature performing the translation @jricole.bsky.social translation of the Rubaiyat is a great example. Omar's words ring out with truth and subtle humor.

Marzena Karpinska@markar.bsky.social · last mo.

A lot of people focus on AI-written fiction, but what about AI literary translation? 📚 We find that AI translation can be readable.. ‼️BUT it also flattens characters' voices 🧙and is less immersive 🫣 than published human translations. Below is my favorite quote from a reader 📖 lait.cs.sfu.ca

A lot of people focus on AI-written fiction, but what about AI literary translation? 📚 We find that AI translation can be readable.. ‼️BUT it also flattens characters' voices 🧙and is less immersive 🫣 than published human translations. Below is my favorite quote from a reader 📖 lait.cs.sfu.ca

Bild
yvesfrtl.bsky.social@yvesfrtl.bsky.social · last mo.

📚 AI-written stories get attention, but AI-translated literature is quietly shaping how readers experience the author So what gets lost ⁉️ 🤖 AI translation into English is readable and often ‘fine’ ✍️ But human translation is valued more ‼️Both vary in quality BUT AI more, even within a single book

First page of the paper "AI translation of literary texts is ‘fine’, but readers still prefer human translations".

📚 AI-written stories get attention, but AI-translated literature is quietly shaping how readers experience the author So what gets lost ⁉️ 🤖 AI translation into English is readable and often ‘fine’ ✍️ But human translation is valued more ‼️Both vary in quality BUT AI more, even within a single book

First page of the paper "AI translation of literary texts is ‘fine’, but readers still prefer human translations".

One more thing @tuhinchakr.bsky.social 's post reminded me of... people tend to rationalize and see things not there. We saw it already in GPT-2 stories - we *expect* things to *mean* something, so we tend to see things that are not there... (link to this old paper: aclanthology.org/2021.emnlp-m...)

BildBild
Marzena Karpinska@markar.bsky.social · 2mo ago

this is how massive illusion of 'creativity' gets crushed... please read it to understand why models may appear to produce coherent text but are in fact Frankenstein factory ...

"Load bearing," "I keep coming back to," "Not just X, but Y" A curse of using AI a lot is that you realize how much of the writing around you is just AI, now People who don't use AI have historically been unable to identify AI prose on sight, but those who use it a lot can spot the tells easily

Bild