Martin Ziqiao Ma
@marstin.bsky.social
https://mars-tin.github.io member of technical staff @ 💭🤖 member of less technical stuff @ aclmentorship phd @ umich; bs @ sjtu prev. @ mit-ibm-watson, adobe, amazon herborium lover, fortune teller, pokémon trainer, szechuan cuisine chef.
A glimpse of interaction models: early experiments with time-aligned micro-turns over native audio, video, and text streams, moving toward more seamless human-AI collaboration. thinkingmachines.ai/blog/interac...
Interaction Models: A Scalable Approach to Human-AI Collaboration
Interaction models move beyond turn-based AI interfaces by handling multimodal, real-time collaboration natively across audio, video, and text.
thinkingmachines.ai
Thanks for sharing, LaCET is now open-sourced :)
Fast Spatial Memory with Elastic Test-Time Training @marstin.bsky.social, Xueyang Yu, Haoyu Zhen, Yuncong Yang, Joyce Chai, Chuang Gan tl;dr: fast-weight module keeps anchor parameters and estimates their importance through an online Fisher-style statistic arxiv.org/abs/2604.07350
Thrilled to announce the 1st Workshop on Computational Developmental Linguistics (CDL) at ACL 2026 🎉 A new venue at the intersection of development linguistics × modern NLP, spearheaded by @fredashi.bsky.social @marstin.bsky.social, and and outstanding team of colleagues! A thread 🧵
NEPA: Next-Embedding Predictive Autoregression sihanxu.me/nepa/ Key ideas: - One self-supervised signal: cosine-style next-embedding prediction - Autoregression runs directly on the model's native embeddings - No pixel decoder (& loss), no contrastive pairs, no task-specific heads, no random masks
NEPA
sihanxu.me
In my new blog, “Test-Time Training Done Better: From Plastic Adaptation to Elastic Memory Consolidation,” I introduce a long-context modeling architecture that learns to adapt and memorize at test time by updating a subset of the model’s weights during inference. mars-tin.github.io/blogs/posts/...
Test-Time Training Done Better: From Plastic Adaptation to Elastic Memory
Elastic Test-Time Training (ETTT) that prevents catastrophic forgetting at inference time and overfitting during pretraining LaCT.
mars-tin.github.io
Gosh, I’m getting way too emotional writing my thesis acknowledgements...
Will be at #NeurIPS2025 (San Diego) Dec 1-9, then in the Bay Area until the 14th. Hmu if you wanna grab coffee and talk about totally random stuff. Thread with a few things I’m excited about. P.S. 4 NeurIPS papers all started pre-May 2024 and took ~1 year of polishing...so proud of the team!
Trying to decide what to do on the first day of #NeurIPS2025? Check out my, @marstin.bsky.social and @xiangyue96.bsky.social's tutorial, "The Science of Benchmarking: What's Measured, What's Missing, What's Next" on December 2 from 1:30 to 4:00pm. benchmarking.science What will we cover? 1/3
@fredashi.bsky.social and I wrote a blog for our new mechinterp paper (arxiv.org/abs/2510.13796), including many unpublished and even negative results that we found meaningful to share. An Open-Notebook Exploration of Emergent Grounding in LMs mars-tin.github.io/blogs/posts/...
An Open-Notebook Exploration of Emergent Grounding in LMs
How We Did This Curiosity-Driven Research? An Open-Notebook Exploration of Emergent Grounding in LMs
mars-tin.github.io
Regrettably can’t attend #COLM2025 due to deadlines, but Jane and Joyce will be presenting our work. :) Jane is an exceptional undergraduate researcher and a great collaborator! Go meet her at COLM if you’re curious about her work on mechanistic interpretability, multimodality, & pragmatics!
Vision-Language Models are not yet pragmatically optimal. We identify 3 key failures of pragmatic competence in referring expression generation with VLMs: (1) cannot uniquely refer to the referent, (2) include excessive or irrelevant information, and (3) misalign with human pragmatic preferences.
🚀 ACL ARR is looking for a Co-CTO to join me lead our amazing tech team and drive the future of our workflow. If you’re interested or know someone who might be, let’s connect! RTs & recommendations appreciated.
🚨 ARR is looking for a volunteer Co-CTO to help improve tech infrastructure! 🛠️ Preferred: • 5+ years in NLP research • Git, CLI tools, Python, and basic HTML • 2-year role, overlapping with current Co-CTO Interested? DM @fredashi.bsky.social or email fhs@uwaterloo.ca #ARR #ACL #NLProc
Unfortunately, I’ll be missing #ACL2025NLP this year — but here are a few things I’m excited about! 👇
📣 Excited to announce SpaVLE: #NeurIPS2025 Workshop on Space in Vision, Language, and Embodied AI! Join us in San Diego to push the frontiers of spatial understanding and reasoning across CV, NLP, and robotics! 👉 space-in-vision-language-embodied-ai.github.io
#CoreCognition #LLM #multimodal #GrowAI We spent 3 years to curate 1503 classic experiments spanning 12 core concepts in human cognitive development and evaluated on 230 MLLMs with 11 different prompts for 5 times to get over 3.8 millions inference data points. A thread (1/n) - #ICML2025 ✅
New Paper Alert ‼️ Current VLMs completely fail human gaze understanding 🙀 and scaling does NO help ‼️ However, humans, since an extremely age 🧒, are extremely sensitive to other people's gaze 🙄 👀 No mentors, no labs, only pre-doc students, 111 VLMs, and we did it 😎
& @tianminshu.bsky.social (+ @marstin.bsky.social, @zhitinghu.bsky.social, @lianhui.bsky.social & more) will present “SimWorld: A World Simulator for Scaling Photorealistic Multi-Agent Interactions,” an @unrealengine.bsky.social-based sim that generates unlimited/diverse urban environments: (13/14)
SimWorld
SimWorld: A World Simulator for Scaling Photorealistic Multi-Agent Interactions
simworld-cvpr2025.maitrix.org
See you at #NAACL2025! I will talk about grounded lexicon acquisition and scaling mechanistically grounded vision language models. Happy to chat if you are around :)
On my way to NAACL✈️! If you're also there and interested in grounding, don't miss our tutorial on "Learning Language through Grounding"! Mark your calendar: May 3rd, 14:00-17:30, Ballroom A. Another exciting collaboration with @marstin.bsky.social @kordjamshidi.bsky.social, Jiayuan, and Joyce!
Vision-Language Models are not yet pragmatically optimal. We identify 3 key failures of pragmatic competence in referring expression generation with VLMs: (1) cannot uniquely refer to the referent, (2) include excessive or irrelevant information, and (3) misalign with human pragmatic preferences.
I won’t be attending #ICLR2025 in person since #NAACL2025 follows right after, but here are a few things I’m excited about (all time in EDT) ⬇️
📢 Join us for the ACL Mentorship Session @naaclmeeting.bsky.social #NAACL2025 Mentors: • @amuuueller.bsky.social • @fredashi.bsky.social • Jiayuan Mao • @marstin.bsky.social • Oana Ignat • Weijia Shi • @zhijingjin.bsky.social
Meet VEGGIE 🥦 VEGGIE is an instructional video generative model trained solely with diffusion loss, designed for both video concept grounding and instruction-based editing. It effectively handles diverse video editing tasks by pixel-level grounded training in a multi-task learning setup. ⬇️
Introducing VEGGIE 🥦—a unified, end-to-end, and versatile instructional video generative model. VEGGIE supports 8 skills, from object addition/removal/changing, and stylization to concept grounding/reasoning. It exceeds SoTA and shows 0-shot multimodal instructional & in-context video editing.
I am looking for a postdoc. A serious-looking call coming soon, but this is to get it going. Topics include (but not limited to): LLMs (🫢!), multimodal LLMs, interaction+learning, RL, intersection with cogsci, ... see our work to get an idea: yoavartzi.com/pubs Plz RT 🙏
Publications
yoavartzi.com
[RT Appreciated] I'm excited to join #ACL @aclmentorship.bsky.social as a mentor! Your ideas can help make mentorship more impactful. Let’s plan together!🚀
We’re planning 12 monthly mentoring sessions for 2025, and we’d love your input! Have topic ideas or feedback? Want to nominate a panelist? Fill out this form and share your thoughts! 💡https://forms.gle/dURA4QUANH3pBxBG8
Submission #1 is in the bag, but we know your masterpiece is next, don’t keep us waiting. 😎 Also, consider joining the program committee and help shape the future of alignment research. Your reviews might just steer the field! forms.gle/MapXyCZhbcFr...
ICLR 2025 Bi-Align Workshop Reviewer Nomination Form
We strive to expand our reviewing pool by welcoming newer members of the community. We encourage nominations from senior community members as well as self-nominations from individuals who have either ...
forms.gle
Happy New Year, everyone!🌟If you’re interested in our ICLR’25 workshop of👫<>🤖Bidirectional Human-AI Alignment, check out our website for more details: bialign-workshop.github.io! Our submission portal is also live now. 😃 Excited to see you join and present your cool work! @iclr-conf.bsky.social 🚀