Graziano Casto

@cstgzn.bsky.social

Developer Relations Engineer @ Akamas • Tech Lead @ CNCF TAG DevEx • Kubernetes v1.35 Release Comms Lead • Kubernetes Blog Maintainer • OSS & Cloud Native Advocate

Is it possible to reduce CPU utilization and render the HPA redundant effectively running on a single pod while maintaining the same load pattern and throughput, all while drastically decreasing application response time? 🤔 Note: Code changes and hardware upgrades are not permitted.

I’m on my 8-hour flight, waiting to land in 🇺🇸Atlanta to attend #KubeCon next week, and I was just thinking about how lucky I am to be able to travel as part of my job. I mean, doing #DevRel is often exhausting but it gives me a lot of opportunities to grow my knowledge, my network and my brand.

Next week I’ll be in Atlanta 🇺🇸 attending #KubeconNA and presenting a panel about “The future of Developer Portals” at #BackstageCon with many amazing folks! What’s your biggest pain with Developer Portals today? And what do you wish for the future? Let me know👇 sched.co/28D2T

CNCF-hosted Co-located Events North America 2025: Panel: The Future of DevPortals - Fabriz...

View more about this event at CNCF-hosted Co-located Events North America 2025

sched.co

Here is the demo repository! You can launch the model on a local cluster (using #ramalama as the inference server) or on an EKS cluster (using #vLLM). Leave a ⭐ and share it around if everything works perfectly. Open an issue to start a shitstorm if it doesn't! 😉 👉 github.com/graz-dev/llm...

Graziano Casto@cstgzn.bsky.social · 9mo ago

That was an incredible success yesterday at #KCDPorto! We shared our experience on how to manage #GenAI model inference on #Kubernetes using #Helm, #ArgoCD and #Istio. The room was packed! It's proof that interest in concrete strategies for AI in the #CloudNative environment is extremely high.