Marcel Böhme

@mboehme.bsky.social

Software Security @ MPI for Security and Privacy PhD @NUS, Dipl.-Inf. @TUDresden Research Group: http://mpi-softsec.github.io

Thrilled to give a keynote at ACM India's ISEC'26 in Jaipur this Friday! How do we know whether our program has no bugs if we have never seen it have any, and if we don't even have anything (i.e., an oracle) that can tell us whether a behavior is a bug or feature? Stick around till Friday!

Bild

📢 Call for Papers for ISSTA 2026 We invite high-quality submissions on software testing and analysis from industry and academia, incl. * research papers * experience papers, and * replicability studies. 📆 29th January 2026 🖊️ issta2026.hotcrp.com 🌐 conf.researchr.org/track/issta-...

ISSTA 2026 - Research papers - ISSTA 2026

Welcome to the website of the ISSTA 2026 conference. The ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) is the leading research symposium on software testing and analysis...

conf.researchr.org

⏱️ 9 days until submission deadline (Dec 11, 23:59 AoE). Organized by: @yannicnoller.bsky.social, @rohan.padhye.org, @ruijiemeng.bsky.social, and Laszlo (@lszekeres.bsky.social) Szekeres.

Yannic Noller@yannicnoller.bsky.social · 11mo ago

#FUZZING'26 CALL FOR PAPERS ────── ✨ After 5 years, we will be again co-located with NDSS! 🔗 fuzzing-workshop.github.io 📅 11. Dec (Submission) //cc @mboehme.bsky.social (MPI-SP), @ruijiemeng.bsky.social (CISPA), @rohan.padhye.org (CMU), László Szekeres (Google)

🧵 A human review of our AI review at #AAAI26. 📝: arxiv.org/abs/2507.00057 🦋 : bsky.app/profile/did:... We are off to a good start. While the synopsis misses the motivation (*why* this is interesting), it offers the most important points. Good abstract-length summary. 1/

Bild
Marcel Böhme@mboehme.bsky.social · 10mo ago

Just accepted at #AAAI26 in Singapore: Our paper on estimating the *correctness* of LLM-generated code in the absence of oracles (e.g., a ground-truth implementation). 📝 arxiv.org/abs/2507.00057 with Thomas Valentin (ENS Paris-Saclay), Ardi Madadi, and Gaetano Sapia (#MPI_SP).

AAAI'26 has been adopting AI reviews in two stages of the review process. I can see that we have to handle the reviewer overload, but I don't think AI reviews are beneficial, at all, for our scientific progress. If our paper gets accepted at #AAAI26, I will review our AI-generated review here 🤠

Bild

After 5 years, we are back at NDSS in San Diego! Looking forward to submissions from the Security and the Software Engineering community!

Yannic Noller@yannicnoller.bsky.social · 11mo ago

#FUZZING'26 CALL FOR PAPERS ────── ✨ After 5 years, we will be again co-located with NDSS! 🔗 fuzzing-workshop.github.io 📅 11. Dec (Submission) //cc @mboehme.bsky.social (MPI-SP), @ruijiemeng.bsky.social (CISPA), @rohan.padhye.org (CMU), László Szekeres (Google)

Can we statistically estimate how likely an LLM-generated program is correct w/o knowing what is a correct program for that task? Sounds impossible-but it's actually really simple. In fact, our measure of "correctness" called incoherence can be estimated (PAC guarantees). arxiv.org/abs/2507.00057

Estimating Correctness Without Oracles in LLM-Based Code Generation

Generating code from natural language specifications is one of the most successful applications of Large Language Models (LLMs). Yet, they hallucinate: LLMs produce outputs that may be grammatically c...

arxiv.org