Steffen Herbold

@sherbold.bsky.social

https://www.fim.uni-passau.de/ai-engineering/

That LLMs can be dangerous is not new. We demonstrate how dangerous they can be, when used to impersonate people. They are so authentic, that humans actually often rather belief the LLM is authentic, not what the person said. The tech is there. We need methods to deal with its misuse.

Uni Passau Research Magazine@unipassauresearch.bsky.social · 2mo ago

Can #AI respond more "human" than real people in political debates? 💬🤖 A new study from our researchers (a.o. @sherbold.bsky.social) found that #ChatGPT impersonated answers from #BBCQT were rated as more authentic than the human responses. @plosone.org 📄 journals.plos.org/plosone/arti...

The next two weeks will be busy for my chair. And we will mostly not be in Passau but in Rio. We will present our work at the CAIN, MSR, ICSE, and ICLR. I look forward to meeting many people I know and getting to know even more people. See you at the beach ... I mean conferences!

Bild

While the Grok-caused CSAM scandal is happening over on X, our work discussing the possible criminal liability of X (and others when publishing generative models) has been accepted the International Conference on AI Engineering. The preprint is already online: arxiv.org/abs/2601.03788

Criminal Liability of Generative Artificial Intelligence Providers for User-Generated Child Sexual Abuse Material

The development of more powerful Generative Artificial Intelligence (GenAI) has expanded its capabilities and the variety of outputs. This has introduced significant legal challenges, including gray a...

arxiv.org

“The whos, whats, and whys of issues related to personal data and data protection in open-source projects on GitHub” by Anne Hennig, Lukas Schulte, Steffen Herbold, Oksana Kulyk, and Peter Mayer will be published in #EMSE! It examines discussions on personal data and data protection on #GitHub. 1/2

The whos, whats, and whys of issues related to personal data and data protection in open-source projects on GitHub - Empirical Software Engineering

Data protection regulations such as the General Data Protection Regulation (GDPR) in the European Union and the California Consumer Privacy Act (CCPA) in the US affect how software may handle the pers...

link.springer.com

Sometimes reviewers still manage to surprise me. A reviewer suggested, we should do a field study for something were we argue that is criminal. We are now planning do address this and wondering if we the reviewer rather wants us to commit crimes or to become criminal investigators 🤨

Just accepted at TMLR: We found evidence of copyright violations by LLMs even when we ask questions that were not part of the training. Indeed, we found that the amount of memorized content was independent from the questions being part of the training or not. openreview.net/forum?id=ddo...

Studying memorization of large language models using answers to...

Large Language Models (LLMs) are capable of answering many software related questions and supporting developers by generating code snippets. These capabilities originate from training on massive...

openreview.net

Success, a luxury problem, and its solution: 🎉 Our quiz is a huge success and incredibly popular on YouTube with now over 100,000 views. 😐 We cannot answer all the feedback and comments individually anymore. 😀 We write a follow up article to answer the most important questions.

Uni Passau Research Magazine@unipassauresearch.bsky.social · last yr.

The article on the reactions on YouTube is now available in English: www.digital.uni-passau.de/en/beitraege...