FloDR: An invertible dimensionality reduction method based on a normalising flow ... combining the advantages of PCA with those of UMAP. We also show what the map "hides" in the remaining dimensions and how wrong it is when two points with different labels are close. arxiv.org/pdf/2607.26278
Max Noichl
@mnoichl.bsky.social
Philosophy with computers at Utrecht University. www.maxnoichl.eu
📣 I'm hiring! Two positions (PhD student & Postdoc) in my ERC project Macroevolution of European Literature. Let's study the cultural evolution of literature using massive data 📚📈 📍Frankfurt | Apply by August 15 Details: www.aesthetics.mpg.de/en/career/jo... Please repost to help spread the word!
Now published in open access! Your one-stop shop for the philosophy of language models. It's the spiritual descendant of our two-part preprint from 2024, fully updated. This should be particularly useful for anyone looking for an entry point into this rapidly growing field.
The Philosophy of Language Models
The success of large language models (LLMs) across many domains of AI research has generated intense debate. Some attribute their impressive performance on complex tasks to human-like linguistic and ...
compass.onlinelibrary.wiley.com
Data visualization: Every human language, from modern day back through protolanguage to a hypothetical common origin. Accurate, sourced, possibly complete. Includes sign language! Tap or mouse over to highlight individual branches; many also have more info. 🛌 this project has been a beast
Networking
7,370 languages visualized as a force-directed network
dr.eamer.dev
This is what a video model does when it is told not to reference the training data (like literally: not prompted, but turning off the classifier free guidance). I gave it a noise seed from elsewhere and it just grinded on through. Most looked like some form of this. I like them.
The amount of story time elapsed vs. space traversed in 500-word passages of fiction, 1550-2020. Periods of fiction more linguistically abstract (blue) are also more chronotopically "abstract": further removed from real-time narration. Data annotated by Qwen3.6-27B for 8,208 passages in 1,232 texts.
Claude code, obey this rite: uv alone should see the light. Ban import star, cast em-dash out, Put pandas, pip, and print to rout. Polars flourish, f-strings burn, Docstrings, hints, at every turn. Clean thy lint, make patterns stand, Hold thy peace and serve my hand!
This is an actual line that was added to the official system prompt for Codex for GPT-5.5 by OpenAI. Usually the system prompt is as minimal as possible, so I assume it would otherwise mention goblins a lot. AIs are weird.
📣 We're #Hiring! #ERC Consolidator Project BOTLEG at @ruhr-uni-bochum.de is recruiting: - 1 PostDoc ➡️ jobs.ruhr-uni-bochum.de/jobposting/9... - 2 PhDs ➡️ jobs.ruhr-uni-bochum.de/jobposting/c... 📅 Deadline: 25 May 2026 #PhDSky #HistSci #PhilSci #HPBio #EnvHist #DigitalHumanities #AcademicSky
📢 CfP for #CHR2027 is out! Submissions on all aspects of computational humanities research welcome; all details available here: 2027.computational-humanities-research.org/cfp/ And yes, CHR2027 will take place a little later than usual. So you *could* start later… but we suggest starting now!
one arm bandit curious? Never heard of the Zollman Effect? Join us Friday next week for the new episode of Conversations at the Center the podcast of the @center4philsci.bsky.social with our guest @kevinzollman.com
1/ "Silicon samples" are becoming more and more common in research and polling. One problem: depending on the analytic decisions made, you can basically get these samples to show any effect you want. The updated version of this preprint is now online! THREAD🧵 arxiv.org/abs/2509.13397
The threat of analytic flexibility in using large language models to simulate human data
Social scientists are now using large language models to create "silicon samples": synthetic datasets intended to stand in for human respondents. However, producing these samples requires many analyti...
arxiv.org
Can large language models stand in for human participants? Many social scientists seem to think so, and are already using "silicon samples" in research. One problem: depending on the analytic decisions made, you can basically get these samples to show any effect you want. THREAD 🧵
This is a bit niche, but for those interested in metaphor and metonymy research, here is one of the first articles I have seen using LLMs as research tool! #cogling #metaphor #metonymy arxiv.org/abs/2604.12919 Oh, they have also done something on visual metonym arxiv.org/abs/2601.17706
MetFuse: Figurative Fusion between Metonymy and Metaphor
Metonymy and metaphor often co-occur in natural language, yet computational work has studied them largely in isolation. We introduce a framework that transforms a literal sentence into three figurativ...
arxiv.org
How can generative AI better support human creativity, without limiting it? If you have thoughts, we invite submissions to our ICML workshop on Generative AI, Creativity, and Human-AI Co-Creation 📍 July 2026, Seoul 📄 Submit by: April 24 (AOE) 🔗 Submission link: openreview.net/group?id=ICM...
ICML 2026 Workshop GenAICreativity
Welcome to the OpenReview homepage for ICML 2026 Workshop GenAICreativity
openreview.net
I made some playable philosophy simulations: -Oxford 1952 -Republic Book I -Jena 1799 -Paris 1945 www.ux-phi.com
There is no best VLM OCR model - rankings can flip completely by document type. I built ocr-bench: run open OCR models on YOUR documents, get a per-collection leaderboard. VLM-as-judge with Bradley-Terry ELO, all running on @hf.co. No local GPU needed.
i'm trying out the novel writing project with Claude in Claude Code, using Pangram to break it out of writing in a clearly identifiable AI-writing style. it's going... interesting so far. i despaired at the beginning but am now cautiously optimistic. not so much at the structural level though.
this is cool tbh all i want is an LLM that sits atop my Zotero library and lets me talk to it tho
our open model proving out specialized rag LMs over scientific literature has been published in nature ✌🏻 congrats to our lead @akariasai.bsky.social & team of students and Ai2 researchers/engineers www.nature.com/articles/s41...
Final CFA for the 8th Scientific Understanding and Representation (SURe) annual workshop, which will take place May 27-29, 2026, at the IFIS PAN in Warsaw. Submission deadline: 20 January 2026. More info: shorturl.at/AUoye @philsci.bsky.social @eenphilsci.bsky.social @epsaphilsci.bsky.social
Current Workshop
CFA: 8th Scientific Understanding and Representation (SURe) annual workshop Call for abstracts We invite authors to submit abstracts of up to 750-words for the upcoming...
shorturl.at
Analytic philosophy can be distinguished from literary criticism with 90-95% accuracy via syntax alone. Moreover, a classifier trained to separate them in early C20 does better predicting future separations than a C21 one predicts past ones, suggesting philosophy syntax narrows/specializes in ~C21.
OpenAlex intégré au Web of Science, ou la capture du travail des “commoners” | carnetist.hypotheses.org/2572
OpenAlex intégré au Web of Science, ou la capture du travail des “commoners”
C’est une annonce qui est passée relativement inaperçue, mais qui mérite que l’on s’y arrête un instant. Clarivate a récemment annoncé l’intégration d’OpenAlex comme une nouvelle base de données au se...
carnetist.hypotheses.org
Three different ways to represent colo(u)r. Work in progress, inspired by an old post by Kat Zhang / The Poet Engineer.
"there is a part of human intelligence which operates in a continuous generalization of the space of words, and other parts entirely which do things which are less well understood" is a perfectly reasonable position which apparently has no adherents
Excited to share my latest publication, "Generative Aesthetics: On formal stuckness in AI verse." It's published in a special issue in the Journal of Cultural Analytics, expertly edited by Tess McNulty and Laura Chapot, on "Computation and Form, Reconsidered." culturalanalytics.org/article/1448...
Generative Aesthetics: On formal stuckness in AI verse | Published in Journal of Cultural Analytics
By Ryan Heuser. This paper examines the formal and aesthetic patterns of AI-generated poems through a series of computational experiments.
culturalanalytics.org
Tomorrow we will have a keynote from Charles Pence (UC Louvain). Thanks to the Dutch Philosophy Research School (OZSW) for supporting this event, and @mnoichl.bsky.social for organizing this with me!
Gregor Betz (KIT) kicking off our "Data Driven Philosophy" Hackathon in Utrecht with his talk: "Doing Philosophy with and for LLMs". Besides input about the state of research and new directions, we're spending three days kicking off new projects.
i am going to try to give a framework of my own understanding which laypeople can understand.
Updated & turned my Big LLM Architecture Comparison article into a video lecture. The 11 LLM archs covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok 2.5 11. GLM-4.5/4.6 www.youtube.com/watch?v=rNlU...
The Big LLM Architecture Comparison
YouTube video by Sebastian Raschka
youtube.com