LlamaStash v0.4.0 is out 🦙 Highlights: - SGLang backend next to vLLM for safetensors repos. 'start owner/repo --backend sglang' picks it per launch. - safer memory checks on unified-memory hosts, so a launch that would freeze the machine is refused. llamastash.dev
Deepu K Sasidharan
@deepu105.bsky.social
@jhipster co-chair. Developer 🥑 @okta. Polyglot dev/Speaker/Author. @Java_Champions. Java, Rust, JS, DevOps. ADHD.
Qwen 3.8 replaced Claude Opus for my coding. 27B and Flash Next on a 128GB Strix Halo laptop, no cloud subscription. Flash Next scores 40 on the AA index against 42 for Opus 4.8. Decode 10-15 tok/s, prefill is the real cost. deepu.tech/local-ai-qwe...
deepu.tech
LlamaStash v0.3.0 is out 🦙 Highlights: - one model, several copies, each under its own name. 'start qwen3 --name coder', and it answers to qwen3@coder on the proxy, CLI and TUI. - 'llamastash run model.yml', a preset file you can commit. llamastash.dev
that 2x speedup is the real deal
LlamaStash v0.1.0 is out 🦙 The big change: MTP speculative decoding, on by default. Roughly 2x faster decode on models that support it. Also new: pick a CUDA/ROCm/Vulkan build per launch, and one JSON shape across list and show. llamastash.dev
vllm backend support is a huge deal
LlamaStash v0.2.0 is out 🦙 The big change: a vLLM backend. The safetensors repos in your HuggingFace cache were invisible before; now they show up in list, launch from the TUI, and answer on the proxy. llama.cpp still owns GGUF. llamastash.dev
LlamaStash v0.2.0 is out 🦙 The big change: a vLLM backend. The safetensors repos in your HuggingFace cache were invisible before; now they show up in list, launch from the TUI, and answer on the proxy. llama.cpp still owns GGUF. llamastash.dev
LlamaStash — terminal-native local-LLM manager (TUI + CLI)
A fast terminal-native TUI and CLI for managing local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON output...
llamastash.dev
LlamaStash v0.1.0 is out 🦙 The big change: MTP speculative decoding, on by default. Roughly 2x faster decode on models that support it. Also new: pick a CUDA/ROCm/Vulkan build per launch, and one JSON shape across list and show. llamastash.dev
LlamaStash — terminal-native local-LLM launcher (TUI + CLI)
A fast terminal-native TUI and CLI for launching local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON outpu...
llamastash.dev
How much local LLM can you run on an AMD Strix Halo with 128GB memory? I managed to fit DeepSeek v4 Flash 284B and Gemma 4 E2B on GPU, Whisper and Qwen3.5 4B on NPU. #strixhalo #amd #deepseek #llamastash
LlamaStash v0.0.6 is out 🦙 A experimental ds4 backend runs @antirez DeepSeek-V4 GGUFs through DwarfStar (ds4) Plus: Lemonade on by default, and saved presets that auto-apply. llamastash.dev #ds4 #AI #deepseek
LlamaStash v0.0.5 is out 🦙 New: named launch presets. Tune a model's launch knobs once, name them, reuse them, per-model or per-arch. They live in plain config.yaml, so you can hand-edit, comment, and commit them to your dotfiles. llamastash.dev
LlamaStash v0.0.4 is out 🦙 - Auto launch is now the default: llama.cpp's --fit sizes context and GPU offload. - A browser UI on a stable port - Anthropic Messages API support. llamastash.dev
KDash 2.0 is out. The Kubernetes terminal dashboard now does more than watch. - Delete, edit, scale, restart, cordon, port-forward, all from the TUI - Action menu with confirm prompts - New themes + live switching Built in Rust 🦀 github.com/kdash-rs/kdash #Kubernetes #Rust #k8s #DevOps
LlamaStash is multi-backend now 🦙 (v0.0.3) llama.cpp stays the zero-overhead default. An experimental, opt-in Lemonade backend unlocks the AMD NPU, vLLM, ONNX and more. Plus vision/audio models, a LAN proxy with bearer auth, and multi-GPU support. llamastash.dev
LlamaStash — terminal-native local-LLM launcher (TUI + CLI)
A fast terminal-native TUI and CLI for launching local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON outpu...
llamastash.dev
How much overhead does an LLM launcher add? I built matched-flags benchmarks across AMD APU (Strix Halo), Apple Silicon, and NVIDIA. Wrapper overhead: within 1% of raw llama-server on every cell. Ollama and LM Studio tell a different story, especially on TTFT. deepu.tech/benchmarking...
How fast is LlamaStash? Overhead, throughput, and a fair comparison with Ollama and LM Studio | Technorage
A reproducible benchmark of LlamaStash against raw llama-server, Ollama, and LM Studio on AMD APU, Apple Silicon, and NVIDIA
deepu.tech
Today I'm releasing LlamaStash 0.0.2: a zero-overhead, terminal-native launcher for llama.cpp. One Rust binary that's a TUI, a CLI, a daemon, and an OpenAI-compatible proxy. Demo below 🧵
Arch Linux, the niri scrolling Wayland compositor, llama.cpp with ROCm, and a 27B model running fully offline on 128GB unified memory. This dev shares why local-first AI coding matters and exactly how the stack fits together. { author: @deepu105.bsky.social } dev.to/deepu105/my-...
My fully offline AI-assisted Linux development machine
My Arch Linux, Niri, and local AI coding setup on the ASUS ROG Flow Z13
dev.to
Congrats to this week's top 7 authors! @deepu105.bsky.social built a fully offline local AI Linux dev setup, and @debbie.codes documented an entire product in 4 days with an AI agent. Check out these and the rest of the posts below 👇 dev.to/devteam/top-...
Top 7 Featured DEV Posts of the Week
Welcome to this week's Top 7, where the DEV editorial team handpicks their favorite posts from the...
dev.to
I use Arch btw! 😉 My current fully offline AI-assisted Linux dev setup: 🐧 Arch Linux 🌀 niri ✨ DankMaterialShell 🤖 OpenCode 🦙 llama.cpp ⚙️ ROCm 🧠 Qwen/Gemma local models 💻 ASUS ROG Flow Z13 deepu.tech/my-fully-off...
My fully offline AI-assisted Linux development machine | Technorage
My Arch Linux, Niri, and local AI coding setup on the ASUS ROG Flow Z13
deepu.tech
KDash 1.0.0 is out 🎉 A big milestone for the terminal UI dashboard for Kubernetes. - direct shell into containers - A Troubleshoot tab - inline filter across views - aggregate logs for workloads - custom themes Release notes: github.com/kdash-rs/kda... #Kubernetes #DevOps #Rust
Learn how to secure your AI agents to prevent Excessive Agency, a top OWASP LLM vulnerability, by implementing a Zero Trust model. #auth0 auth0.com/blog/mitigat...
Mitigate Excessive Agency in AI Agents with Zero Trust Security
Mitigate Excessive Agency in AI Agents using a Zero Trust Security model. Practical guide for developers to implement RBAC, FGA, OAuth, a...
auth0.com
@deepu105.bsky.social, may the session run seamlessly and engage the audience! Access is given FGA and RAG guide the way Agents know their path #Devoxx #Akka #room7
Jfokus has always been about community + knowledge sharing 🙌 Here are more of the great speakers from 2025. @sharatchander.bsky.social @renato.cavalcanti.io @deepu105.bsky.social 👉 Want to speak in 2026? Submit now: jfokus.se/iamahero
I've been diving deep into the world of AI lately. My latest blog post explores how to build an AI agent that can call internal and external APIs using LangGraph and Auth0 Token Vault. 🗓️ You can check it out to learn how to use it! #AI #GenAI #LangGraph #ToolCalling auth0.com/blog/genai-t...
How to build an AI Assistant with LangGraph and Next.js
Learn how to build a tool-calling AI agent using LangGraph, Next.js, and Auth0. Integrate your own API as tools. Use Google
auth0.com
Published my first HyDE theme for Hyprland :) #Hyprland #HyDE #ArchLinux github.com/deepu105/hyd...
GitHub - deepu105/hyde-theme-catppuccin-macchiato: Catppuccin Macchiato with Mauve Accent for HyDE
Catppuccin Macchiato with Mauve Accent for HyDE. Contribute to deepu105/hyde-theme-catppuccin-macchiato development by creating an account on GitHub.
github.com
Join me next week at #WeAreDevelopers to Delay the AI Overlords: How OAuth and OpenFGA Can Keep Your AI Agents from Going Rogue #auth0 #AI #security
Thanks to everyone who attended my talk at the #DublinTechSummit Here are the slides from the talk, and you can find the GitHub repo link for the demo on the last slide. docs.google.com/presentation...
DublinTechSummit: Delay the AI Overlords
[intro] Hello friends, hope you are all having a great time. Today we're going to talk about how to stop your AI systems from spilling corporate secrets like a gossipy coworker after happy hour. Let’s...
docs.google.com
kdash is a TUI dashboard for Kubernetes clusters. It displays pod metrics, logs, resource YAMLs, and supports context switching, glob filters, clipboard copying and more. @deepu105.bsky.social made kdash using @ratatui.rs and is Terminal Tool of the Week! ⭐️