Ethan Mollick

@emollick.bsky.social

Professor at Wharton, studying AI and its implications for education, entrepreneurship, and work. Author of Co-Intelligence. Book: https://a.co/d/bC2kSj1 Substack: https://www.oneusefulthing.org/ Web: https://mgmt.wharton.upenn.edu/profile/emollick

This time, I had Fable built me a casual Van Gogh city building game that I faked in an AI video last year. The key mechanic the AI came up with is painting the landscape with big brushstrokes & an environment that evolves with weather and seasons. Chill and pretty. Play it: the-sower.netlify.app

Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down after Sydney & got Copilot to market quickly (the 1st professional AI tool) Google dealt with blowback about Bard and AI Overview and invented deep research.

An unexpectedly stress-reducing use of Codex/Code is just fixing problems with my various Windows machines: weird driver issues, game incompatibilities, even just tiny stuff that used to annoy me (why does the program that I set to run at startup not run at startup?). Hours saved. Annoying hours.

Sure, Threejs in webpages are neat but, Codex: "you have access to Blender & Unity. I want you to make a new game, with full assets, in which you play as an otter who can get into mech suits shaped like animals and that are critical to game play" Took over my computer, built assets & gave me this

I continue to think that a lack of verifiable answers in many fields is a real issue for LLMs but not as big a problem as it sometimes is made out to be. As models are improving at formal domains, they also are Improving at lots of other less-verifiable domains as well, though jaggedness remains.

Bild

OpenAI announces 10 discoveries from their next model. Observations:: 1) AI is getting very good at math 2) Two years ago LLMs failed at basic math 3) This cost less than $2000 in current API fees 4) OpenAI is focusing on announcing benefits, not just risks, of new models openai.com/index/ten-ad...

Bild

One big result in our study at Procter & Gamble was that AI blurred the lines between jobs. Now OpenAI has a similar finding Organizational boundaries are becoming porous, the walls thinning. Companies are going to need to think about division of labor in a new way, things are getting chaotic now.

BildBild

Managing AI requires new interfaces. These are the physical things I have been trying out to control Code. Given how useful voice mode is, the Teenage Engineering Ting walky-talky has been a surprising favorite. The Codex Micro is beautiful & fun, but displays too little info vs the Stream Deck

Bild

I gave Flux 3 the penultimate lines of Eliot's The Wasteland, which involve both switches in tone and in language. I had it read by a modern Fisher King as a city decays behind him. The results were surprisingly good.

Fable is amazing but needs to stop talking like someone who has read only pulp fantasy: "I have shown you the way but you must open the door. The map exists but the path is yours. The atlas of your instinct — every fact must first know itself" Please, just make the infographic about cheese I wanted

Flux 3 is pretty darn impressive. This is what it produced with the prompt: "tracking shot that follows a female astronaut with her helmet open as she walks through a regency dance in a traditional manor, with a mural on the wall painted by Rothko. Pushing people out of the way to make room..." 1/2

Those who follow my feed know that I test new video models by having them show an otter using a laptop on an airplane. The new Flux 3 video model is really good, so I had it do a variation on theme, which will become apparent a couple seconds in.

It's a superpower to know the names of many beautiful and interesting things in the age of AI. You can invoke Vaporwave, Muqarnas, Bauhaus, or Art Nouveau whiplash curves. Sfumato and Grisaille and Notan. Polysyndeton and Zeugma. You just have to know what to ask for. What a time for the humanities

Kimi K3's weights were released, making it the most powerful open weights model in the world. Been reading it since it came out, page one starts strong: mostly negatives, a late run of zeros, and one unexpectedly large positive value.

Bild

For the better part of 20,000 centuries of human history, not much happened. We spent almost all of those making slightly different variations of one tool, we only figured out the metal thing after 19,940 centuries. And stuff really took off 2 centuries ago. That we can sort of keep up is amazing.

BildBild

I think most people, when talking about open weights AI models, don’t deeply believe in the vision of AGI/ASI that lab insiders tend to believe in. Like they don’t expect AI to really present grave semi-autonomous biosecurity or other similar risks. Whether this view is right or not, no one knows

Bild

A year later, Fable builds me a version of the Cezanne city builder game. The AI came up with the idea of an impressionist city builder where you paint with gestures & the town grows around it, with neighborhoods acquiring characters as they evolve. Play with it here: cezanne-city.netlify.app

Ethan Mollick@emollick.bsky.social · last yr.

City builder game in the styles of Cezanne, Piranesi, the Voynich Manuscript, Van Gogh, Rembrandt, Rothko, and Seurat (those are in order). Made with Veo 2.

Ha! It did it: "We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics" I really thought it would treat "now do benchbenchbenchbenchbench" as a joke, but Sol actually did reasonable experiments.

BildBildBildBild
Ethan Mollick@emollick.bsky.social · 2w ago

As a joke I prompted Codex "Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a good arXiv paper." I got a PDF. But the paper is actually kind of interesting? Weird.