Johannes Hoffart

@hoffart.ai

CTO, AI at SAP - #foundationmodel #linkedbusinessdata #knowledgegraph #nlp #ai - www.hoffart.ai

A new player enters the arena of Foundation Models on Tabular Data: www.limix.ai - novel methods for pre-training and data generation that look highly relevant. Their evaluation on selected datasets is showing strong performance. Exciting times, looking forward to further in depth comparisons!

LimiX

limix.ai

Our team developing Foundation Models on Tables & Linked Business Data is looking for a new Senior Applied Research Scientist! Excited about pushing the frontier in foundation models on tabular data? Want to have business impact and academic visibility? Look no further: jobs.sap.com/job/Walldorf...

Senior/Principal Applied Research Scientist (f/m/d): Foundation Models on Linked Business Data

Senior/Principal Applied Research Scientist (f/m/d): Foundation Models on Linked Business Data

jobs.sap.com

For the past 3 years, I've taught a course on Machine Learning for Climate Change to undergrads. At times, people have asked if the course lectures could be made available online. While I can't offer that, I have decided to start making "5 Minute Papers on AI for the Planet" videos. Hope its useful!

5 Minute Papers on AI for the Planet

AI is more than just chatbots! Learn about how AI can be used to protect biodiversity, fight climate change, and just better understand our planet through 5-minute explainers covering academic papers ...

youtube.com

Can you train a performant language model using only openly licensed text? We are thrilled to announce the Common Pile v0.1, an 8TB dataset of openly licensed and public domain text. We train 7B models for 1T and 2T tokens and match the performance similar models like LLaMA 1 & 2

Bild

This was helpful. Also worth noting that Bluesky remains a very fraught place for AI discussions for a variety of reasons, good & bad, but with the impact of keeping a lot of the most relevant AI news, paper discussions & biggest names on X That might change, but it hasn’t yet. Still posting, tho.

Naomi Saphra@nsaphra.bsky.social · last yr.

I wrote something up for AI people who want to get into bluesky and either couldn't assemble an exciting feed or gave up doomscrolling when their Following feed switched to talking politics 24/7.

When using LLM-as-a-judge, practitioners often use greedy decoding to get the most likely judgment. But we found that deriving a score from the judgment distribution (like taking the mean) works better! ❌LLM-as-a-judge with greedy decoding 😎Using the distribution of the judge’s labels

Victor Wang@victorwang37.bsky.social · last yr.

LLM judges have become ubiquitous, but valuable signal is often ignored at inference. We analyze design decisions for leveraging judgment distributions from LLM-as-a-judge: 🧵 (w/ Michael J.Q. Zhang, @eunsol.bsky.social)

Discover European cities ✈️ while building your career! Check out the ELLIS PhD/Postdoc Program's 2025 Winter & Summer School Schedule! Dive deep into cutting-edge #AI research, learn from top researchers & connect with peers across Europe. Learn more: bit.ly/42iow66 #PhD #machinelearning

Bild

Our first release of 2025: 𝙨𝙢𝙤𝙡𝙖𝙜𝙚𝙣𝙩𝙨, 𝘁𝗵𝗲 𝘀𝗶𝗺𝗽𝗹𝗲𝘀𝘁 𝗹𝗶𝗯𝗿𝗮𝗿𝘆 𝘁𝗼 𝗯𝘂𝗶𝗹𝗱 𝗮𝗴𝗲𝗻𝘁𝗶𝗰 𝘀𝘆𝘀𝘁𝗲𝗺𝘀! 💥 Main logic in ~1000 LoC 🧑‍💻 Agent writes its actions in code! LLMs are much better at writing code than current standard of writing JSON => higher perf 🌍 Any LLM support (h/t LiteLLM) 🛡️ Secure code exec (h/t E2B)

Bild

Have a look at our work on foundation models on tabular data, published today at #TRL @ #NeurIPS2024: 📜 PORTAL, an open weight and code foundation model trained on tabular data, and 📜 SALT, a real business data set containing millions of sales orders across multiple tables. Further details 👇

Tired of saturated benchmarks? Want scope for a significant leap in capabilities? 🔥 Introducing BALROG: a Benchmark for Agentic LLM and VLM Reasoning On Games! BALROG is a challenging benchmark for LLM agentic capabilities, designed to stay relevant for years to come. 1/🧵

Bild

How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️