Chris Potts

@cgpotts.bsky.social

Stanford Professor of Linguistics and, by courtesy, of Computer Science, and member of @stanfordnlp.bsky.social and The Stanford AI Lab. He/Him/His. https://web.stanford.edu/~cgpotts/

I am confident OpenAI will become profitable. They are smart, creative, highly incentivized, and well-funded. On the other hand, any app/company that depends on capturing most of the value from OpenAI's models has an uncertain future, like the Twitter apps of old.

Bill Labov died this morning. I'm not coherent enough to talk about how important and influential and brilliant he was. I am very sad. I was so lucky to know him, and I am grateful every day that he (and Gillian, and Walt, etc) built an academic field where kindness is expected.

ReFT: Representation Finetuning for Language Models Spotlight Poster Zhengxuan Wu · Aryaman Arora · Zheng Wang · Atticus Geiger · Dan Jurafsky · Christopher D Manning · Christopher Potts Fri 13 Dec 07:00 PM UTC [West Ballroom A-D]

Papers (partly) from @stanfordnlp at #NeurIPS 2024: Oral: Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making Manling Li · Shiyu Zhao · Qineng Wang · Kangrui Wang · … · Weiyu Liu · Percy Liang · Li Fei-Fei · Jiayuan Mao · Jiajun Wu Wed 11 Dec 11:50 PM UTC [East Ballroom A, B]

My primary role as a Department Chair at Stanford has become complaining about bureaucratic overreach at Stanford. I have send dozens of messages on this topic just this quarter. And yet I have still not mastered the spelling of "bureaucratic".

Natural Language Processing—artificial intelligence that uses human language—has been on a roll lately. You’ve probably noticed! So the Stanford NLP Group has been growing, and diversifying into lots of new topics, including agents, language model programs, and socially aware #NLP. nlp.stanford.edu

Group picture of people in the Stanford NLP Group gathered in front of the shores of Lake Tahoe.

Lena is a masterful Borgesian fiction imagining the first human brain to be captured on disk. The program enters a state of "terror and extreme panic" on boot-up. If you deny that a machine could be sentient, do you deny the story's premise or the potential reality of such terror? qntm.org/mmacevedo

Lena

You can now buy this story as part of my collection, Valuable Humans in Transit and Other Stories. This collection also includes a sequel story, titled "Driver". Russian translation French transla...

qntm.org

I've been going through the Wordle archives seeking to achieve a mind-meld with WordleBot, and I finally did it (in a puzzle from winter 2024). I appreciate that it even seems to have known this was my goal.

My solution to a Wordle puzzle alongside the solution from WordleBot. Both of us guessed CRANE / SCALE / PLACE (luck 76). The WordleBot says, "Sensational. We are as one."

For PhD recommendation systems: this year, as in every year in my experience, MIT EECS gets my highest recommendation: (1) one subjective multiple choice question that does not really try to hide its subjectivity behind made-up numbers and (2) letter upload. I wish my own institution's were as good.

There's a known bug in how we compute "word" probabilities with subword-based LMs that mark beginnings of words -- as pointed out by Byung-doh Oh and Will Schuler, & @tpimentel.bsky.social and Clara Meister I'm pleased to announce that minicons now includes a fix which runs batch-wise!

Code: from minicons import scorer

lm = scorer.IncrementalLMScorer("gpt2-xl", "cuda:0")

stimuli = ["I was a matron in France", "I was a mat in France"]

# old way, no correction
# P.S. gpt2 does not automatically add a bos token at the beginning...
lm.token_score(stimuli, bos_token=True, surprisal=True, base_two=True, bow_correction=False)

'''Rounded Output
[[('<|endoftext|>', 0.0),
  ('I', 5.85),
  ('was', 4.28),
  ('a', 4.67),
  ('mat', 16.34),
  ('ron', 1.74),
  ('in', 2.12),
  ('France', 11.43)],
 [('<|endoftext|>', 0.0),
  ('I', 5.85),
  ('was', 4.28),
  ('a', 4.67),
  ('mat', 16.34),
  ('in', 10.78),
  ('France', 10.71)]]
'''

# the new way! notice the surprisal of "mat" in both cases
lm.token_score(stimuli, bos_token=True, surprisal=True, base_two=True, bow_correction=True)

'''Rounded Output
[[('<|endoftext|>', 0.0),
  ('I', 6.30),
  ('was', 3.84),
  ('a', 4.68),
  ('mat', 16.34),
  ('ron', 2.11),
  ('in', 1.75),
  ('France', 11.42)],
 [('<|endoftext|>', 0.0),
  ('I', 6.30),
  ('was', 3.84),
  ('a', 4.68),
  ('mat', 21.34),
  ('in', 5.80),
  ('France', 10.69)]]
'''Screenshot from Oh and Schuler showing surprisal values for the partial sentences "I was a matron in" and "I was a mat in" using GPT-2 XL with leading whitespaces and trailing whitespaces.

I'm excited to kick off my Bluesky presence with wonderful news: Our paper "Reference-Based Metrics Are Biased Against Blind and Low-Vision Users' Image Description Preferences" won a Best Paper Award at the NLP for Positive Impact Workshop at EMNLP! Read it here: aclanthology.org/2024.nlp4pi-...

Reference-Based Metrics Are Biased Against Blind and Low-Vision Users’ Image Description Preferences

Rhea Kapur, Elisa Kreiss. Proceedings of the Third Workshop on NLP for Positive Impact. 2024.

aclanthology.org