"Everyone is an AI agent builder โ if you let them." At @arize.bsky.social's Observe 2026, CrewAI shared enterprise lessons on where agent ROI shows up and how to scale building across an org. Sketchnoted ๐ ๐ Video link in the comments.
Jim Bennett
@jimbobbennett.dev
World's most energetic dev rel Microsoft MVP. ๐ally. I โค๏ธ Star Wars Lego & ๐ปโโ๏ธ. Father, husband. He/him.
As agents become digital coworkers, access control built for humans starts to break. At @arize.bsky.social's Observe 2026, WorkOS's Michael Grinich explored the identity & security model for autonomous agents. Sketchnoted ๐ ๐ Video link in the comments.
What does agent adoption actually look like in production? At @arize.bsky.social's Observe 2026, Mastra shared patterns from thousands of teams โ what separates shipped agents from stuck prototypes. Sketchnoted ๐ ๐ Video link in the comments.
From experimentation to production, agents need whole-lifecycle platforms. At @arize.bsky.social's Observe 2026, Microsoft's Sebastian demoed building, deploying, evaluating & governing agents with Microsoft Foundry. Sketchnoted ๐ ๐ Video link in the comments.
Scaling agents from prototype to production is an infrastructure problem. At @arize.bsky.social's Observe 2026, Anyscale's Robert Nishihara explained how Ray scales RL, inference & multimodal AI โ and why RL is having a moment. Sketchnoted ๐ ๐ Video link in the comments.
The hardest problems in AI aren't model problems anymore โ they're evaluation problems. At @arize.bsky.social's Observe 2026, Hamel Husain argued agents are bringing the data scientist back to AI engineering. Sketchnoted ๐ ๐ Video link in the comments.
What are the smartest AI investors seeing before everyone else? At @arize.bsky.social's Observe 2026, Jaya Gupta of Foundation Capital shared where VC is flowing across the AI stack. Sketchnoted the fireside ๐ ๐ Video link in the comments.
When a hallucination is a regulatory + financial risk, responsible AI gets real. At @arize.bsky.social's Observe 2026, BlackRock shared how it deploys AI to support pros managing trillions โ with real guardrails & evaluation. Sketchnoted ๐ ๐ Video link in the comments.
With autonomous agents, observability shifts from "what happened" to "why did the agent do that." At @arize.bsky.social's Observe 2026, AWS's Nate Slater explored how agentic AI changes observability. Sketchnoted ๐ ๐ Video link in the comments.
Two of the fastest-growing open-source agent projects, one conversation. At @arize.bsky.social's Observe 2026, OpenClaw & Nous Research debated where agent frameworks go next โ memory, skill creation, long-term learning. Sketchnoted ๐ ๐ Video link in the comments.
"AI agents need specs, not prompts." At @arize.bsky.social's Observe 2026, George Zhang argued the engineer's real job is specifying the hill agents climb โ tests, evals, rubrics, constraints. Sketchnoted ๐ ๐ Video link in the comments.
What happens when a leading AI coding company turns its product inward? At @arize.bsky.social's Observe 2026, Cursor shared how it uses agents, evals & agent-powered workflows to build Cursor itself. Sketchnoted ๐ ๐ Video link in the comments.
"Kubernetes is not your sandbox." At @arize.bsky.social's Observe 2026, the Daytona team argued K8s wasn't built for agent workloads, and walked through what agent-native infrastructure actually needs. Sketchnoted ๐ ๐ Video link in the comments.
Building agents is harder than the demos make it look. At @arize.bsky.social's Observe 2026, Anthropic's Marius Buleandra shared why agent failures compound in production, how to design evals that catch them, and why human review still matters. Sketchnoted ๐ ๐ Video link in the comments.
How do you improve a product with hundreds of millions of users? At @arize.bsky.social's Observe 2026, OpenAI's Stuart Sy showed how ChatGPT turns fragmented 'vibes' into evidence + action. I sketchnoted the talk ๐ ๐ Video link in the comments.
Agents stopped being demos this year โ they're shipping code, fixing bugs, and running real workflows. In @arize.bsky.social's Observe 2026 keynote, the founders lay out what's next. I sketchnoted the whole keynote ๐ ๐ Video link in the comments.
A field guide to four of the latest from @jimbobbennett.dev, where each one leaks, and why no single pass rate was ever going to survive this. arize.com/blog/long-h...
Do you have an AI agent? Do you actually know what it is doing? Do you know if it works? Typically the answer to the first question is yes, and for the second it's we think so, based off 'vibes'. Which is a terrible way to build and run production software. 1/2
Apple paid Google ~$1B/yr to license memory for Siri. OpenAI rebuilt ChatGPT memory in place. Anthropic gave models an API to consolidate their own. All called "memory." None is what users mean. @jimbobbennett.dev wrote a field map: arize.com/blog/memory...
Memory is still a missing primitive: Cataloguing what the field is actually shipping
This week the field shipped four kinds of memory, and Apple paid Google a billion dollars a year for one of them. None of the four is what the demos imply. A field map of what's actually shipping, and the missing primitive that sits between the buckets.
arize.com
"I genuinely don't care. Pick one." That was my contribution to a meeting last week where the team was debating two tools. And it was the most useful thing I said all day. "Strong opinions, loosely held" is the "approved" take. I think it's mostly nonsense.
Strong opinions, strongly held - and why I don't care about your tooling debate
I was in a meeting last week where the team was debating which of two tools to use for a job. Both of them do the thing.
linkedin.com
Microsoft picked OpenInference. Twice. The open trust stack for AI agents announced at #MSBuild, ASSERT for evaluation, ACS for controls, both ride on the open tracing standard Arize built for agents. arize.com/blog/micros...
Microsoft's open trust stack runs on OpenInference
Microsoft's open trust stack for AI agents puts ASSERT and Agent Control Specification on top of OpenInference, connecting evaluation, runtime controls, and observability through a shared trace contract.
arize.com
At Microsoft Build? Our 2 must do things for today: 1. Catch Sarah Bird's session - Observe and control agents with OSS tools build.microsoft.com/en-US/sessi... 2. Head to the Microsoft AI expert booth to meet with @jimbobbennett.dev from our devrel team about AI observability and Evals #MSBuild
Observe and control agents across any framework with open source tools
As AI agents move into production, developers own safety, governance, and reliability across Microsoft Agent Framework and open-source stacks. This session shows how to govern agents end to end: turning your requirements into context-aware evaluations, stress-testing against adversarial risks, applying open controls that work across frameworks, and keeping humans in the loop on high-stakes actions. Leave with a blueprint for shipping agents at enterprise scale. Seating for this session is first-come, first-served. Add it to your schedule to plan your day and arrive early to secure a spot.
build.microsoft.com
Will you be at Microsoft Build this week, either in person in SF or virtually? Our very own @jimbobbennett.dev will be giving a demo session on understanding and fixing agents with open source observability and evals, Wednesday 3:30pm, Theater C. #MSBuild build.microsoft.com/en-US/sessi...
Understand and fix Agent Framework apps with observability and evals
Your AI apps are getting more complex, with multiple agents, tools, and different orchestration patterns. This makes them harder to understand, debug, and test. This session shows you how to visualize the decisions your LLMs are making in complex Microsoft Agent Framework applications, using open standards and open source tooling to provide you with instrumentation and observability. You'll also see how you can use another LLM as a judge to evaluate how well your AI app is working. Seating for this session is first-come, first-served. Add it to your schedule to plan your day and arrive early to secure a spot.
build.microsoft.com
Phoenix now lets you compose evaluation strategies in code. Most eval tooling hands you a fixed menu of judge templates. Real evaluation is rarely that tidy.
All Londoners who read this post will be reminded of it every day when they get on the tubeโฆ
Your AI agent disagrees with your human reviewers all day. Most teams treat that as noise. It's the most useful signal in the system. @jimbobbennett.dev wrote up how to mine the gap and feed it back to the agent. arize.com/blog/self-i...
Your AI agent disagrees with your human reviewers all day. Most teams treat that as noise. It's the most useful signal in the system. @jimbobbennett.dev wrote up how to mine the gap and feed it back to the agent. arize.com/blog/self-i...
Every AI agent deployed inside an enterprise sometimes quietly disagrees with the humans running the same process. The written policy says one thing. The institutional knowledge sitting in Slack threads, hallway conversations, and the heads of long-tenure employees says another. ๐งต 1/3
One AI Question with Cam Young We asked our Strategic AI Solutions Architect: What's a ๐ฅ take on evals? His answer: Stop guessing and start measuring. Use "LLM-as-a-judge" for nuance, but don't ignore code-based evals for speed and human annotators for ground truth. #AI #AIStrategy #AIEvals #LLM
Claude Code users - want to be notified when Claude wants your attention? If you have a RPi and a 3.5" screen, then here's a project that puts a happy character on the screen. Bored when Claude is busy, dances when Claude needs your attention. All the code is here: github.com/jimbobbennet...
GitHub - jimbobbennett/claude-notify: Raspberry Pi + 3.5" touchscreen Claude mascot that dances when Claude Code on your Mac needs your input
Raspberry Pi + 3.5" touchscreen Claude mascot that dances when Claude Code on your Mac needs your input - jimbobbennett/claude-notify
github.com
How many instructions can you give an LLM before it starts to forget about some of them? How long can your skill file be? How big can your prompt get? I did some actual *research*! https://www.linkedin.com/pulse/models-got-order-magnitude-better-following-one-year-laurie-voss-9fymc/