Jim Bennett

@jimbobbennett.dev

World's most energetic dev rel Microsoft MVP. ๐ŸŒˆally. I โค๏ธ Star Wars Lego & ๐Ÿปโ€โ„๏ธ. Father, husband. He/him.

As agents become digital coworkers, access control built for humans starts to break. At @arize.bsky.social's Observe 2026, WorkOS's Michael Grinich explored the identity & security model for autonomous agents. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

What does agent adoption actually look like in production? At @arize.bsky.social's Observe 2026, Mastra shared patterns from thousands of teams โ€” what separates shipped agents from stuck prototypes. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

From experimentation to production, agents need whole-lifecycle platforms. At @arize.bsky.social's Observe 2026, Microsoft's Sebastian demoed building, deploying, evaluating & governing agents with Microsoft Foundry. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

Scaling agents from prototype to production is an infrastructure problem. At @arize.bsky.social's Observe 2026, Anyscale's Robert Nishihara explained how Ray scales RL, inference & multimodal AI โ€” and why RL is having a moment. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

The hardest problems in AI aren't model problems anymore โ€” they're evaluation problems. At @arize.bsky.social's Observe 2026, Hamel Husain argued agents are bringing the data scientist back to AI engineering. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

When a hallucination is a regulatory + financial risk, responsible AI gets real. At @arize.bsky.social's Observe 2026, BlackRock shared how it deploys AI to support pros managing trillions โ€” with real guardrails & evaluation. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

With autonomous agents, observability shifts from "what happened" to "why did the agent do that." At @arize.bsky.social's Observe 2026, AWS's Nate Slater explored how agentic AI changes observability. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

Two of the fastest-growing open-source agent projects, one conversation. At @arize.bsky.social's Observe 2026, OpenClaw & Nous Research debated where agent frameworks go next โ€” memory, skill creation, long-term learning. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

"AI agents need specs, not prompts." At @arize.bsky.social's Observe 2026, George Zhang argued the engineer's real job is specifying the hill agents climb โ€” tests, evals, rubrics, constraints. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

What happens when a leading AI coding company turns its product inward? At @arize.bsky.social's Observe 2026, Cursor shared how it uses agents, evals & agent-powered workflows to build Cursor itself. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

"Kubernetes is not your sandbox." At @arize.bsky.social's Observe 2026, the Daytona team argued K8s wasn't built for agent workloads, and walked through what agent-native infrastructure actually needs. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

Building agents is harder than the demos make it look. At @arize.bsky.social's Observe 2026, Anthropic's Marius Buleandra shared why agent failures compound in production, how to design evals that catch them, and why human review still matters. Sketchnoted ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

Agents stopped being demos this year โ€” they're shipping code, fixing bugs, and running real workflows. In @arize.bsky.social's Observe 2026 keynote, the founders lay out what's next. I sketchnoted the whole keynote ๐Ÿ‘‡ ๐Ÿ”— Video link in the comments.

Bild

Do you have an AI agent? Do you actually know what it is doing? Do you know if it works? Typically the answer to the first question is yes, and for the second it's we think so, based off 'vibes'. Which is a terrible way to build and run production software. 1/2

"I genuinely don't care. Pick one." That was my contribution to a meeting last week where the team was debating two tools. And it was the most useful thing I said all day. "Strong opinions, loosely held" is the "approved" take. I think it's mostly nonsense.

Strong opinions, strongly held - and why I don't care about your tooling debate

I was in a meeting last week where the team was debating which of two tools to use for a job. Both of them do the thing.

linkedin.com

Every AI agent deployed inside an enterprise sometimes quietly disagrees with the humans running the same process. The written policy says one thing. The institutional knowledge sitting in Slack threads, hallway conversations, and the heads of long-tenure employees says another. ๐Ÿงต 1/3

Bild

One AI Question with Cam Young We asked our Strategic AI Solutions Architect: What's a ๐Ÿ”ฅ take on evals? His answer: Stop guessing and start measuring. Use "LLM-as-a-judge" for nuance, but don't ignore code-based evals for speed and human annotators for ground truth. #AI #AIStrategy #AIEvals #LLM

Claude Code users - want to be notified when Claude wants your attention? If you have a RPi and a 3.5" screen, then here's a project that puts a happy character on the screen. Bored when Claude is busy, dances when Claude needs your attention. All the code is here: github.com/jimbobbennet...

GitHub - jimbobbennett/claude-notify: Raspberry Pi + 3.5" touchscreen Claude mascot that dances when Claude Code on your Mac needs your input

Raspberry Pi + 3.5" touchscreen Claude mascot that dances when Claude Code on your Mac needs your input - jimbobbennett/claude-notify

github.com