Srishti

@srishtiy.bsky.social

ELLIS PhD Fellow @belongielab.org | @aicentre.dk | University of Copenhagen | @amsterdamnlp.bsky.social | @ellis.eu Multi-modal ML | Alignment | Culture | Evaluations & Safety| AI & Society Web: https://www.srishti.dev/

Happy to share that our work on multi-modal framing analysis of news was accepted to #EMNLP2025! Understanding news output and embedded biases is especially important in today's environment and it's imperative to take a holistic look at it. Looking forward to presenting it in Suzhou!

Arnav Arora@rnv.bsky.social · last yr.

🚨New pre-print 🚨 News articles often convey different things in text vs. image. Recent work in computational framing analysis has analysed the article text but the corresponding images in those articles have been overlooked. We propose multi-modal framing analysis of news: arxiv.org/abs/2503.20960

🚀 Technical practitioners & grads — join to build an LLM evaluation hub! Infra Goals: 🔧 Share evaluation outputs & params 📊 Query results across experiments Perfect for 🧰 hands-on folks ready to build tools the whole community can use Join the EvalEval Coalition here 👇 forms.gle/6fEmrqJkxidy...

[EvalEval Infra] Better Infrastructure for LM Evals

Welcome to EvalEval Working Group Infrastructure! Please help us get set up by filling out this form - we are excited to get to know you! This is an interest form to contribute/collaborate on a research project, building standardized infrastructure for AI evaluation. Status Quo: The AI evaluation ecosystem currently lacks standardized methods for storing, sharing, and comparing evaluation results across different models and benchmarks. This fragmentation leads to unnecessary duplication of compute-intensive evaluations, challenges in reproducing results, and barriers to comprehensive cross-model analysis. What's the project? We plan to address these challenges by developing a comprehensive standardized format for capturing the complete evaluation lifecycle. This format will provide a clear and extensible structure for documenting evaluation inputs (hyperparameters, prompts, datasets), outputs, metrics, and metadata. This standardization enables efficient storage, retrieval, sharing, and comparison of evaluation results across the AI research community. Building on this foundation, we will create a centralized repository with both raw data access and API interfaces that allow researchers to contribute evaluation runs and access cached results. The project will integrate with popular evaluation frameworks (LM-eval, HELM, Unitxt) and provide SDKs to simplify adoption. Additionally, we will populate the repository with evaluation results from leading AI models across diverse benchmarks, creating a valuable resource that reduces computational redundancy and facilitates deeper comparative analysis. Tasks? As a collaborator, you would be expected to: Work towards merging/integrating popular evaluation frameworks (LM-eval, HELM, Unitxt) Group 1 - Extend to Any Task: Design universal metadata schemas that work for ANY NLP task, extending beyond current frameworks like lm-eval/DOVE to support specialized domains (e.g., machine translation) Group 2 - Save the Relevant: Develop efficient query/download systems for accessing only relevant data subsets from massive repositories (DOVE: 2TB, HELM: extensive metadata) The result will be open infrastructure for the AI research community, plus an academic publication. When? We're looking for researchers who can join ASAP and work with us for at least 5 to 7 months. We are hoping to find researchers who would take this on as an active project (8 hours+/week) in this period.

forms.gle

Can you train a performant language model using only openly licensed text? We are thrilled to announce the Common Pile v0.1, an 8TB dataset of openly licensed and public domain text. We train 7B models for 1T and 2T tokens and match the performance similar models like LLaMA 1 & 2

Bild

"I don’t want to just be entering text prompts for the rest of my life." I spoke to political cartoonists, including Pulitzer-winner Mark Fiore, about how they are using AI image generators in their work. My latest for @niemanlab.org. www.niemanlab.org/2025/06/i-do...

“I don’t want to outsource my brain”: How political cartoonists are bringing AI into their work

Pulitzer-winning cartoonists are experimenting with AI image generators.

niemanlab.org

Check out our new preprint 𝐓𝐞𝐧𝐬𝐨𝐫𝐆𝐑𝐚𝐃. We use a robust decomposition of the gradient tensors into low-rank + sparse parts to reduce optimizer memory for Neural Operators by up to 𝟕𝟓%, while matching the performance of Adam, even on turbulent Navier–Stokes (Re 10e5).

Bild

I am excited to announce our latest work 🎉 "Cultural Evaluations of Vision-Language Models Have a Lot to Learn from Cultural Theory". We review recent works on culture in VLMs and argue for deeper grounding in cultural theory to enable more inclusive evaluations. Paper 🔗: arxiv.org/pdf/2505.22793

Paper title "Cultural Evaluations of Vision-Language Models
Have a Lot to Learn from Cultural Theory"

This morning at P1 a handful of lucky of lab members got to see the telescope while centre secretary Björg had the dome open for a building tour 🔭 (1/7)

Bild

When you have a lot of work before the deadline push, you keep thinking of others things (distractions) you’d like to do. The day you get free, those things suddenly don’t seem important anymore. And kind of miss work! 🙄

I will present our #ICLR2025 Spotlight paper MM-FSS this week in Singapore! Curious how MULTIMODALITY can enhance FEW-SHOT 3D SEGMENTATION WITHOUT any additional cost? Come chat with us at the poster session — always happy to connect!🤝 🗓️ Fri 25 Apr, 3 - 5:30 pm 📍 Hall 3 + Hall 2B #504 More follow

Bild

Starting on a new social media is like moving to a new country and starting all over again. Find new friends, staying in touch with old ones, find what you would like (again!) and finding if you would fit in a new place (again!) 🙄