Jack Hessel

@jmhessel.bsky.social

jmhessel.com Seattle bike lane enjoyer. Opinions my own.

m̶e̶n̶ Americans will literally l̶e̶a̶r̶n̶ ̶e̶v̶e̶r̶y̶t̶h̶i̶n̶g̶ ̶a̶b̶o̶u̶t̶ ̶a̶n̶c̶i̶e̶n̶t̶ ̶R̶o̶m̶e̶ invest billions into self driving cars instead of g̶o̶i̶n̶g̶ ̶t̶o̶ ̶t̶h̶e̶r̶a̶p̶y̶ building transit

in llm-land, what is a tool, a function, an agent, and (most elusive of all): a "multi-agent system"? (This had been bothering me recently; are all these the same?) @yoavgo.bsky.social's blog is a clarifying read on the topic -- I plan to adopt his terminology :-) gist.github.com/yoavg/9142e5...

What makes multi-agent LLM systems multi-agent?

What makes multi-agent LLM systems multi-agent? GitHub Gist: instantly share code, notes, and snippets.

gist.github.com

I've spent the last two years trying to understand how LLMs might improve middle-school math education. I just published an article in the Journal of Educational Data Mining describing some of that work: "Designing Safe and Relevant Generative Chats for Math Learning in Intelligent Tutoring Systems"

Journal of Educational Data Mining

Large language models (LLMs) are flexible, personalizable, and available, which makes their use within Intelligent Tutoring Systems (ITSs) appealing. However, their flexibility creates risks: inaccura...

jedm.educationaldatamining.org

I'm not an """ AGI """ person or anything, but, I do think process reward model RL/scaling inference compute is quite promising for problems with easily verified solutions like (some) math/coding/ARC problems.

Post nicht verfügbar.

“They said it could not be done”. We’re releasing Pleias 1.0, the first suite of models trained on open data (either permissibly licensed or uncopyrighted): Pleias-3b, Pleias-1b and Pleias-350m, all based on the two trillion tokens set from Common Corpus.

Bild

Blue skies 🦋 , hot (?) takes 🔥 Constrained output for LLMs, e.g., outlines library for vllm which forces models to output json/pydantic schemas, is cool! But, because output tokens cost much more latency than input tokens, if speed matters: bespoke, low-token output formats are often better.

Information retrieval systems usually operate as a model "cascade" -- fast vector search over billions of documents followed by a more expressive LLM "re-ranking" the resulting top-K. But beware 👻 ! Despite expressivity, top-K re-rankers generalize poorly as K increases. arxiv.org/pdf/2411.11767

Figure 1 from the linked paper, which illustrates the performance of a re-ranker dropping as the number of re-ranked documents increases.

I'm recruiting 1-2 PhD students to work with me at the University of Colorado Boulder! Looking for creative students with interests in #NLP and #CulturalAnalytics. Boulder is a lovely college town 30 minutes from Denver and 1 hour from Rocky Mountain National Park 😎 Apply by December 15th!

A photo of Boulder, Colorado, shot from above the university campus and looking toward the Flatirons.

📣 I am recruiting 1-2 PhD students for Fall 2025 at the University of Maryland College of Information. Consider applying if you're interested in language, society/politics, and computers! Deadline Dec 3: ischool.umd.edu/academics/ph... And pls share with anyone who may be interested!

Doctor of Philosophy in Information Studies (PhD) - College of Information (INFO)

This doctoral program prepares students to address the hardest social and technical problems of today and tomorrow.

ischool.umd.edu