Martin Ziqiao Ma

@marstin.bsky.social

https://mars-tin.github.io member of technical staff @ 💭🤖 member of less technical stuff @ aclmentorship phd @ umich; bs @ sjtu prev. @ mit-ibm-watson, adobe, amazon herborium lover, fortune teller, pokémon trainer, szechuan cuisine chef.

In my new blog, “Test-Time Training Done Better: From Plastic Adaptation to Elastic Memory Consolidation,” I introduce a long-context modeling architecture that learns to adapt and memorize at test time by updating a subset of the model’s weights during inference. mars-tin.github.io/blogs/posts/...

Test-Time Training Done Better: From Plastic Adaptation to Elastic Memory

Elastic Test-Time Training (ETTT) that prevents catastrophic forgetting at inference time and overfitting during pretraining LaCT.

mars-tin.github.io

Will be at #NeurIPS2025 (San Diego) Dec 1-9, then in the Bay Area until the 14th. Hmu if you wanna grab coffee and talk about totally random stuff. Thread with a few things I’m excited about. P.S. 4 NeurIPS papers all started pre-May 2024 and took ~1 year of polishing...so proud of the team!

Bild

@fredashi.bsky.social and I wrote a blog for our new mechinterp paper (arxiv.org/abs/2510.13796), including many unpublished and even negative results that we found meaningful to share. An Open-Notebook Exploration of Emergent Grounding in LMs mars-tin.github.io/blogs/posts/...

An Open-Notebook Exploration of Emergent Grounding in LMs

How We Did This Curiosity-Driven Research? An Open-Notebook Exploration of Emergent Grounding in LMs

mars-tin.github.io

Regrettably can’t attend #COLM2025 due to deadlines, but Jane and Joyce will be presenting our work. :) Jane is an exceptional undergraduate researcher and a great collaborator! Go meet her at COLM if you’re curious about her work on mechanistic interpretability, multimodality, & pragmatics!

Martin Ziqiao Ma@marstin.bsky.social · last yr.

Vision-Language Models are not yet pragmatically optimal. We identify 3 key failures of pragmatic competence in referring expression generation with VLMs: (1) cannot uniquely refer to the referent, (2) include excessive or irrelevant information, and (3) misalign with human pragmatic preferences.

New Paper Alert ‼️ Current VLMs completely fail human gaze understanding 🙀 and scaling does NO help ‼️ However, humans, since an extremely age 🧒, are extremely sensitive to other people's gaze 🙄 👀 No mentors, no labs, only pre-doc students, 111 VLMs, and we did it 😎

Bild

Vision-Language Models are not yet pragmatically optimal. We identify 3 key failures of pragmatic competence in referring expression generation with VLMs: (1) cannot uniquely refer to the referent, (2) include excessive or irrelevant information, and (3) misalign with human pragmatic preferences.

Meet VEGGIE 🥦 VEGGIE is an instructional video generative model trained solely with diffusion loss, designed for both video concept grounding and instruction-based editing. It effectively handles diverse video editing tasks by pixel-level grounded training in a multi-task learning setup. ⬇️

Shoubin Yu@shoubin.bsky.social · last yr.

Introducing VEGGIE 🥦—a unified, end-to-end, and versatile instructional video generative model. VEGGIE supports 8 skills, from object addition/removal/changing, and stylization to concept grounding/reasoning. It exceeds SoTA and shows 0-shot multimodal instructional & in-context video editing.

Our #ICLR2025 workshop on Bidirectional Human-AI Alignment now has a sister SIG at #CHI2025! 📅 Submission deadline extended to Feb 15 to welcome both ML & HCI communities. 📍 Authors may opt to present at either ICLR or CHI 2025! More info here: bialign-workshop.github.io