Deepu K Sasidharan

@deepu105.bsky.social

@jhipster co-chair. Developer 🥑 @okta. Polyglot dev/Speaker/Author. @Java_Champions. Java, Rust, JS, DevOps. ADHD.

LlamaStash v0.4.0 is out 🦙 Highlights: - SGLang backend next to vLLM for safetensors repos. 'start owner/repo --backend sglang' picks it per launch. - safer memory checks on unified-memory hosts, so a launch that would freeze the machine is refused. llamastash.dev

LlamaStash v0.2.0 is out 🦙 The big change: a vLLM backend. The safetensors repos in your HuggingFace cache were invisible before; now they show up in list, launch from the TUI, and answer on the proxy. llama.cpp still owns GGUF. llamastash.dev

LlamaStash — terminal-native local-LLM manager (TUI + CLI)

A fast terminal-native TUI and CLI for managing local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON output...

llamastash.dev

LlamaStash v0.1.0 is out 🦙 The big change: MTP speculative decoding, on by default. Roughly 2x faster decode on models that support it. Also new: pick a CUDA/ROCm/Vulkan build per launch, and one JSON shape across list and show. llamastash.dev

LlamaStash — terminal-native local-LLM launcher (TUI + CLI)

A fast terminal-native TUI and CLI for launching local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON outpu...

llamastash.dev

LlamaStash is multi-backend now 🦙 (v0.0.3) llama.cpp stays the zero-overhead default. An experimental, opt-in Lemonade backend unlocks the AMD NPU, vLLM, ONNX and more. Plus vision/audio models, a LAN proxy with bearer auth, and multi-GPU support. llamastash.dev

LlamaStash — terminal-native local-LLM launcher (TUI + CLI)

A fast terminal-native TUI and CLI for launching local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON outpu...

llamastash.dev

How much overhead does an LLM launcher add? I built matched-flags benchmarks across AMD APU (Strix Halo), Apple Silicon, and NVIDIA. Wrapper overhead: within 1% of raw llama-server on every cell. Ollama and LM Studio tell a different story, especially on TTFT. deepu.tech/benchmarking...

How fast is LlamaStash? Overhead, throughput, and a fair comparison with Ollama and LM Studio | Technorage

A reproducible benchmark of LlamaStash against raw llama-server, Ollama, and LM Studio on AMD APU, Apple Silicon, and NVIDIA

deepu.tech

I've been diving deep into the world of AI lately. My latest blog post explores how to build an AI agent that can call internal and external APIs using LangGraph and Auth0 Token Vault. 🗓️ You can check it out to learn how to use it! #AI #GenAI #LangGraph #ToolCalling auth0.com/blog/genai-t...

How to build an AI Assistant with LangGraph and Next.js

Learn how to build a tool-calling AI agent using LangGraph, Next.js, and Auth0. Integrate your own API as tools. Use Google

auth0.com