Viktor Farcic

@vfarcic.bsky.social

Inference is a memory management problem, not a compute one. Your GPU almost never runs out of arithmetic first. It runs out of room to hold conversations. And your engine already tells you the exact ceiling at startup — before a single request arrives. 🧵 https://youtu.be/g_5g1hBmAzA

YouTube Video

Inference is a memory management problem, not a compute one. Your GPU almost never runs out of arithmetic first. It runs out of room to hold conversations. And your engine already tells you the exact ceiling at startup — before a single request arrives. 🧵 https://youtu.be/g_5g1hBmAzA

youtu.be

What do CFEngine, Puppet, Chef, Ansible, Terraform & Crossplane all have in common? Same job. Different name. Five arguments. Five generations certain they were inventing something new. New video — go watch. 🏰 https://youtu.be/TF5iJbAb6Xc

YouTube Video

What do CFEngine, Puppet, Chef, Ansible, Terraform & Crossplane all have in common? Same job. Different name. Five arguments. Five generations certain they were inventing something new. New video — go watch. 🏰 https://youtu.be/TF5iJbAb6Xc

youtu.be

Docker won the internet. Every cloud, every pipeline, every deployment runs on its idea. But the company that built it? A ghost. The scariest stories aren't about failing; they're about winning so completely your victory floats away without you. 🐳 https://youtu.be/kmLztKueTcI

YouTube Video

Docker won the internet. Every cloud, every pipeline, every deployment runs on its idea. But the company that built it? A ghost. The scariest stories aren't about failing; they're about winning so completely your victory floats away without you. 🐳 https://youtu.be/kmLztKueTcI

youtu.be

There's no patch for prompt injection. It's not a bug—it's how the models work. The real question: when it happens, how much damage can it do? New video on sandboxing AI agents so you can walk away and still sleep at night. 🔒 https://youtu.be/tWb7v7Nq808

YouTube Video

There's no patch for prompt injection. It's not a bug—it's how the models work. The real question: when it happens, how much damage can it do? New video on sandboxing AI agents so you can walk away and still sleep at night. 🔒 https://youtu.be/tWb7v7Nq808

youtu.be

What if your tests read like a spec doc AND never needed babysitting? DevAssure lets you write plain-English YAML, commit it to your repo, and an AI agent clicks through your app like a real user. Zero selectors. Zero page objects. https://youtu.be/bNJoeSXUdQg

YouTube Video

What if your tests read like a spec doc AND never needed babysitting? DevAssure lets you write plain-English YAML, commit it to your repo, and an AI agent clicks through your app like a real user. Zero selectors. Zero page objects. https://youtu.be/bNJoeSXUdQg

youtu.be

The hardest part of fleet-scale inference isn't the model — it's the matchmaking. Which model runs on which GPU, across which cluster? Modelplane's scheduler handles all of it automatically. New video out now 👇 #AI #DevOps #Kubernetes https://youtu.be/Mc0BM63Pv00

YouTube Video

The hardest part of fleet-scale inference isn't the model — it's the matchmaking. Which model runs on which GPU, across which cluster? Modelplane's scheduler handles all of it automatically. New video out now 👇 #AI #DevOps #Kubernetes https://youtu.be/Mc0BM63Pv00

youtu.be

I review entire AI-built features without reading a single line of code. My agents record their own tests as video clips. I watch a short reel, confirm it works, hit merge. Two gates. Total autonomy in between. 🧵 https://youtu.be/03VwVUadsRM

YouTube Video

I review entire AI-built features without reading a single line of code. My agents record their own tests as video clips. I watch a short reel, confirm it works, hit merge. Two gates. Total autonomy in between. 🧵 https://youtu.be/03VwVUadsRM

youtu.be

The AI agent server stack that actually works: ✅ Cheap PC (~$900, one-time) ✅ Ubuntu Server ✅ Terminal agents (Claude Code, Codex) ✅ Tailscale + tmux ✅ Devbox + vals Full video breakdown 🎥 https://youtu.be/tCEhU9bQ-XY

YouTube Video

The AI agent server stack that actually works: ✅ Cheap PC (~$900, one-time) ✅ Ubuntu Server ✅ Terminal agents (Claude Code, Codex) ✅ Tailscale + tmux ✅ Devbox + vals Full video breakdown 🎥 https://youtu.be/tCEhU9bQ-XY

youtu.be

A dummy is a dummy, no matter whether they're using AI or not. AI agents are amplifiers. If you're good at your job, they make you better. If you're not, they help you cause a shitstorm at scale in minutes. https://youtu.be/_j4epp_MqF4

YouTube Video

A dummy is a dummy, no matter whether they're using AI or not. AI agents are amplifiers. If you're good at your job, they make you better. If you're not, they help you cause a shitstorm at scale in minutes. https://youtu.be/_j4epp_MqF4

youtu.be

The fix for broken code review isn't hiring more reviewers. It's flipping the ORDER. AI first → humans second. By the time a person opens the PR, the noise is already gone. Only the real judgment calls remain. Full breakdown in my latest video: https://youtu.be/KyodvaYemxM

YouTube Video

The fix for broken code review isn't hiring more reviewers. It's flipping the ORDER. AI first → humans second. By the time a person opens the PR, the noise is already gone. Only the real judgment calls remain. Full breakdown in my latest video: https://youtu.be/KyodvaYemxM

youtu.be

I run multiple AI agents with different models — Opus as orchestrator, GPT as coder, Kimi as reviewer. OpenRouter makes this effortless. One gateway, full flexibility, real cost visibility. Here's the setup 👇 https://youtu.be/YKCJwk6xLwA

YouTube Video

I run multiple AI agents with different models — Opus as orchestrator, GPT as coder, Kimi as reviewer. OpenRouter makes this effortless. One gateway, full flexibility, real cost visibility. Here's the setup 👇 https://youtu.be/YKCJwk6xLwA

youtu.be