Spandana Gella

@spandanagella.bsky.social

Sr Mgr & Research Scientist @ServiceNowRSRCH, Montreal

🚨New Paper!🚨 How do reasoning LLMs handle inferences that have no deterministic answer? We find that they diverge from humans in some significant ways, and fail to reflect human uncertainty… 🧵(1/10)

Bild

🚀 New paper from our team at @servicenowresearch.bsky.social!⁣ ⁣ 💫𝐒𝐭𝐚𝐫𝐅𝐥𝐨𝐰: 𝐆𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐧𝐠 𝐒𝐭𝐫𝐮𝐜𝐭𝐮𝐫𝐞𝐝 𝐖𝐨𝐫𝐤𝐟𝐥𝐨𝐰 𝐎𝐮𝐭𝐩𝐮𝐭𝐬 𝐅𝐫𝐨𝐦 𝐒𝐤𝐞𝐭𝐜𝐡 𝐈𝐦𝐚𝐠𝐞𝐬⁣ We use VLMs to turn 𝘩𝘢𝘯𝘥-𝘥𝘳𝘢𝘸𝘯 𝘴𝘬𝘦𝘵𝘤𝘩𝘦𝘴 and diagrams into executable workflows 🖍️→⚙️⁣ ⁣ 🔗 arxiv.org/abs/2503.218... 📝 tinyurl.com/3utdbn97%E2%... #Sketch2Flow #AI #VLM

Very excited to announce our GUI benchmarking dataset UI-Vision : uivision.github.io Our evals reveal current GUI-models struggle with grounding small elements, dense UIs and has limited domain/spatial/motion understanding. Watch out this space for more exciting stuff from us!

UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction

UI-Vision

uivision.github.io

Xiangru (Edward) Jian@edwardjian.bsky.social · last yr.

🚀 Super excited to announce UI-Vision: the largest and most diverse desktop GUI benchmark for evaluating agents in real-world desktop GUIs in offline settings. 📄 Paper: arxiv.org/abs/2503.15661 🌐 Website: uivision.github.io 🧵 Key takeaways 👇

Web agents powered by LLMs can solve complex tasks, but our analysis shows that they can also be easily misused to automate harmful tasks. See the thread below for more details on our new web agent safety benchmark: SafeArena and Agent Risk Assessment framework (ARIA).

Xing Han Lu@xhluca.bsky.social · last yr.

Agents like OpenAI Operator can solve complex computer tasks, but what happens when users use them to cause harm, e.g. spread misinformation? To find out, we introduce SafeArena (safearena.github.io), a benchmark to assess the capabilities of web agents to complete harmful web tasks. A thread 👇

📢New Paper Alert!🚀 Human alignment balances social expectations, economic incentives, and legal frameworks. What if LLM alignment worked the same way?🤔 Our latest work explores how social, economic, and contractual alignment can address incomplete contracts in LLM alignment🧵

Bild

If you want to know all about the exciting stuff we do with web agents @servicenowresearch.bsky.social register here and interact with our team including the amazing @alex-lacoste.bsky.social and @adrouinenv.bsky.social

Alexandre Lacoste@alex-lacoste.bsky.social · 2y ago

Join us for a co-hosted Happy Hour NeurIPS 2024 with ServiceNow and IMean.ai as we explore the cutting edge of WebAgent development! 📅 Date: Dec 13th 6:00pm PST 📍 Location: 15min walk from Neurips see details after RSVP 🎉 RSVP Here: lu.ma/rw9x9vc6