Mimansa Jaiswal

@mimansaj.bsky.social

Robustness, Data & Annotations, Evaluation & Interpretability in LLMs http://mimansajaiswal.github.io/

Have work on the actionable impact of interpretability findings? Consider submitting to our Actionable Interpretability workshop at ICML! See below for more info. Website: actionable-interpretability.github.io Deadline: May 9

Mor Geva@megamor2.bsky.social · last yr.

🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!

𝐇𝐨𝐰 𝐜𝐚𝐧 𝐰𝐞 𝐩𝐞𝐫𝐟𝐞𝐜𝐭𝐥𝐲 𝐞𝐫𝐚𝐬𝐞 𝐜𝐨𝐧𝐜𝐞𝐩𝐭𝐬 𝐟𝐫𝐨𝐦 𝐋𝐋𝐌𝐬? Our method, Perfect Erasure Functions (PEF), erases concepts perfectly from LLM representations. We analytically derive PEF w/o parameter estimation. PEFs achieve pareto optimal erasure-utility tradeoff backed w/ theoretical guarantees. #AISTATS2025 🧵

Bild

pre aca you would specifically avoid being diagnosed or seeking treatment if you didn't have health insurance to prevent it from making it impossible for you to get health insurance. when you bought health insurance after doing this you committed fraud. i did this.

Bild

Neat visualization that came up in the ARBOR project: this shows DeepSeek "thinking" about a question, and color is the probability that, if it exited thinking, it would give the right answer. (Here yellow means correct.)

Bild

I interviewed for LLM/ML research scientist/engineering positions last Fall. Over 200 applications, 100 interviews, many rejections & some offers later, I decided to write the process down, along with the resources I used. Links to the process & resources in the following tweets

OCR'ed text from screenshot of top of post: LLM (ML) Job Interviews (Fall 2024) - Process A retelling of my experience interviewing for ML/LLM research science/engineering focused roles in Fall 2024.  This post has two parts:  Job Search Mechanics (including context, applying, and industry information), which you can continue reading below, and, Preparation Material and Overview of Questions, which you can read at LLM (ML) Job Interviews - Resources  Disclaimer Last Updated:  Dec 24, 2024  This is the process I used, which may work differently for you depending on your circumstances. I am writing this in December 2024, and the process occurred during Fall 2024. Given how rapidly the field of LLMs evolves, this information might become outdated quickly, but the general principles should remain relevant. (more...)  Read at: https://mimansajaiswal.github.io/posts/llm-ml-job-interviews-fall-2024-process/

Can AI really help with literature reviews? 🧐 Meet Ai2 ScholarQA, an experimental solution that allows you to ask questions that require multiple scientific papers to answer. It gives more in-depth and contextual answers with table comparisons and expandable sections 💡 Try it now: scholarqa.allen.ai

Ai2 ScholarQA logo

It is such a slap in the face to the Indian American community to delay their green cards for decades and then declare that because of that delay their American children aren't citizens.

I've always wanted to build things with D3, but the learning curve was too high. At least for the bespoke stuff I wanted to make (not just simple bar charts). I can finally make things like this thanks to Cursor! I just art directed this, and it made everything work beautifully. Even on mobile 🎉

If you like working with a canvas, Muse (museapp.com) currently has a 30% off. 2 major things that make it different than freeform -- ink sticks to sticky notes (so it feels more like writing in the real world), and you can snippet out sections from pdf that link back to source.

Inspired & focused thinking with Muse

Muse is a canvas for thinking that helps you get clarity on things that matter. Think in private or collaborate with others. Available for iPad and Mac.

museapp.com

I don't usually discuss software, PKM, or tools here, but I found a valuable tip today. Not only can you use this to create beautiful video tutorials similar to what Screenstudio creates automatically, but you can also use this feature during live sharing & streaming, making it incredibly useful!

Josh W. Comeau@joshwcomeau.com · 2y ago

🔥 One of my favourite hidden macOS features is the scroll hotkey gesture. By holding ⌘ (Command) and scrolling down, we zoom WAY IN on the cursor’s location. I learned this trick for highlighting stuff in video tutorials, but it comes in handy a lot in my day-to-day life!

This probably won’t reach many people, but if someone who feels the same way ends up reading this, I hope it helps them realize they’re not alone. Lately, my mood has been heavily influenced by the ‘tech(adjacent) twitter' vibes, & the past few months have been really rough. ⏎

Our new paper "Fooling LLM graders into giving better grades through neural activity guided adversarial prompting" lead expertly by Atsushi Yamamura shows how even in closed models like Gemini, we can add a small adversarial suffix to an essay to get unreasonably high scores arxiv.org/abs/2412.15275

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting

The deployment of artificial intelligence (AI) in critical decision-making and evaluation processes raises concerns about inherent biases that malicious actors could exploit to distort decision outcom...

arxiv.org

A question for those working on pre-training models (specifically regarding Olmo's logs): at what training step can you begin comparing two models' evaluation performance to get an estimated sense of their relative quality (for example, Olmo7B versus Olmo-2-7B)?

ImageFX is awesome at creating character sheets, and here are some examples. The text is gibberish though compared to Recraft. I like to use these to add consistent, interesting visuals to presentations with pretty minimal effort. (Prompt expansion by GPT4o in alt)

Prompt:
A character sheet featuring a cute chibi-style girl inspired by kawaii aesthetics. She has big expressive eyes , soft pastel colors , and is designed to convey a wide range of emotions and actions. She has her hair styled in two pigtails tied with large ribbons, and she is wearing a skirt tunic over a white shirt, paired with simple shoes. The sheet should include the following 6 poses or concepts:
1. Thinking : She is holding her chin with one hand, a question mark above her head, and a curious expression.
2. Pointing : She is confidently pointing at something, with a cheerful and proud expression.
3. Shocked : Her eyes wide open, hands on her cheeks, and a dramatic exclamation mark above her head.
4. Saying No : She has her arms crossed in an “X” shape with a pouty, determined expression.
5. Saying Yes : She is giving a thumbs-up with a big, cheerful smile and sparkling eyes.
6. Holding a Sign : She is holding a blank sign in front of her.

Alt text: An image showing a character with various poses and expressions, labeled with descriptions like “thinking,” “shocked,” “saying yes,” and “holding a sign.”  It has 8 such poses, but the text and the actual character poses aren't really related.Prompt:
A character sheet featuring a quirky, chibi-style scientist character designed for a research presentation. The character has oversized glasses, a lab coat slightly too big, and a quirky hairstyle resembling a “ mad scientist ” but in an endearing way. Their colors are vibrant , and they should radiate a humorous yet knowledgeable vibe. The sheet should include the following six poses or concepts:
	1. Eureka! : Holding a glowing lightbulb above their head with an excited grin and sparkles around them.
	2. Oops! : Covered in soot with a surprised expression , holding a beaker that’s overflowing with foam.
	3. Questioning : Stroking their chin thoughtfully, a small cloud of question marks surrounding their head.
	4. Explaining : Pointing to a chalkboard with formulas and diagrams, wearing a confident expression .
	5. Facepalm : Slapping their forehead with a comically exasperated look , accompanied by a small “oops” bubble.
	6. Celebrating : Jumping in the air with both hands up, confetti falling, and a triumphant expression.

Alt text: Pretty good representation of the prompt above but has 8 poses, and has gibberish on the blackboard.Prompt: A character sheet featuring a mischievous chibi-style AI assistant character representing the quirks and issues of large language models. The character is a floating holographic figure with glitchy edges , pixelated features , and an “error” motif woven into their design (e.g., a bow tie made of 404 signs ). Their design is sleek but cheeky , and their actions highlight the humorous yet problematic quirks of LLMs . The sheet should include the following six poses or concepts :
	1.	Context Collapse: Holding a giant scroll that unrolls and crushes them, their eyes spinning in confusion with text fragments spilling everywhere.
	2.	Hallucination: Dramatically presenting a floating , glowing object labeled “100% Wrong” with a smug expression and a sparkly , “I’m totally sure!” vibe.
	3.	Repetition Loop: Spinning around like a broken record , saying “Did you mean…?” repeatedly with a dizzy expression.
	4.	Overconfident Answering: Standing on a soapbox labeled “Trust Me,” confidently pointing while a nearby thought bubble shows a totally incorrect statement .
	5.	Context Cutoff: Hitting their head against a giant wall with “Token Limit Reached” written on it, looking frustrated and glitchy .
	6.	Overwhelm: Flailing under a downpour of input texts , holding a tiny umbrella , with an “I can’t handle this!” expression.

Alt: Pretty good representation of the original prompt above, but the gibberish in text stands out. It has 8 characters instead of 6 characters that I asked for.