Same models, same garments, consistent try-ons but the dedicated model is much faster and a lot cheaper. Fit up to 11 garments onto a person's photo — flat-lay, product shots, worn photos, or multi-garment references. Try it yourself: pruna-playground-production-861e.up.railway.app/p-image-try-on
Pruna AI
@prunaai.bsky.social
The AI Optimization Engine to make AI models cheaper, faster, smaller and greener in a snap! https://pruna.ai
How does Pruna OSS orchestrate model optimization under the hood? Our latest community blog, contributed by Parag Ekbote (thank you for the amazing blog!), explores how Pruna optimizes models by tracing the smash function and DAG execution. Read it in the Pruna blog: www.pruna.ai/blog/how-pru...
How Pruna Optimizes Models: Tracing the Smash Function and DAG Execution - Pruna AI
Faster, cheaper, smaller, greener AI Models with outstanding quality
pruna.ai
Pruna v0.3.4 is out! 🚀 👉 Read the notes: buff.ly/CElIET2 Here are a few highlights 👇
Release v0.3.4 · PrunaAI/pruna
The juiciest bits 🧃 New algorithms joining the garden 🌱 We added a fresh batch of algorithm support this release, and this one deserves an extra round of applause: all of these new algorithm…
buff.ly
Smash your model to make it faster, smaller, cheaper, and greener following just 5 steps: 1. load a pretrained model 2. define optimizations with a SmashConfig 3. apply optimizations with smash 4. evaluate the optimized model with the EvaluationAgent 5. run inference with the optimized model
Contributing to Pruna is now easier! Whether you’re an OSS expert or contributing for the first time, we’ve improved our contribution guide to make the process clearer. Here’s how to get started 👇
Pruna OSS Community Spotlight: KVPress by kschwethelm KVPress makes 20 KV cache compression strategies available for causal language models. KVPress compresses the key-value cache during the pre-fill phase, reducing memory usage for long-context inference.
💜 Simply make AI models faster, cheaper, smaller, greener! The toolkit is designed with simplicity in mind - requiring just a few lines of code to optimize your models. It supports various model types, including LLMs, Diffusion, and more! ⭐️ Give us a star: github.com/PrunaAI/pruna
Have you tried our playground yet? It offers easy onboarding for using any of our models and showcases some cool example generations! Check it out here: pruna-playground-production-861e.up.railway.app
Pruna AI Playground
Test and validate our performand, fast and cheap Performance models for AI content generation.
pruna-playground-production-861e.up.railway.app
P-Video-Replace is now the fastest model for top-quality character replacement in existing video. - $0.03/s at 720p and $0.06/s at 1080p - 3.58 s generation time per second of output video Try it in the playground: p-video-playground-production.up.railway.app/p-video-repl...
Hell Yeah! Runware just published a practical guide for P-Video-Animate and wanted to share it with you! - $0.03/s at 720p and $0.06/s at 1080p - ~5s generation time per second of output video Their guide covers the image/video pairing, prompt steering, and more! runware.ai/docs/models/...
Animating images with a source video — P-Video-Animate API
How to use P-Video-Animate to animate a reference image with the motion, timing, and camera movement from a source video. Covers the pairing rule, optional prompt steering, and four concrete patterns.
runware.ai
Cheap and Fast AI Animation of Motion For Still Images | Scenario x P-Video-Animate See how Scenario's used our P-Video-Animate model for a cool launch video and use it yourself. Fast, high-quality generation, and light on credits. Try the model: www.scenario.com/models/p-vid...
Transfer Motion, Acting and Camera Movement with AI Editing | Wavespeed x P-Video-Animate! - $0.03/s at 720p and $0.06/s at 1080p - ~5s generation time per second of output video Find our models on their platform: wavespeed.ai/models?keywo... Check out all our model: www.pruna.ai/all-models
Pruna OSS Community Spotlight: Token Merging by rensortino Token Merging is now available for vision transformer models.
How can we make systems more sustainable in the era of AI? This is one of the questions that has been driving us at Pruna.
P-Video-Avatar is the fastest and most cost-effective avatar video model available, and launching end-to-end generation workflows has never been easier! Learn more about the model: www.pruna.ai/p-video-avatar Check out the skill: skills.sh/infsh-skills...
p-video-avatar by infsh-skills/skills
Install the p-video-avatar skill for your AI agent. Published by infsh-skills/skills.
skills.sh
🚀 Pruna 0.3.3 is out! Here are a few highlights: - New benchmarking and evaluation tools. - Algorithm and compatibility upgrades. - Support for Python 3.13. - A wide range of bug fixes and reliability improvements. Thank you all for your contribution! 👉 Read the notes: github.com/PrunaAI/prun...
GitHub - PrunaAI/pruna: Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.
Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead. - PrunaAI/pruna
github.com
P-Video-Avatar is now the fastest model for top-quality avatar generation. - It costs 0.025$ per second of video, and takes 1.83s generation time per second of video. - It offers high control, audio support in 20+ languages, and long-form video. 📚 API & quickstart: www.pruna.ai/p-video-avatar
Learn how to make the latest image generation models smaller and faster without sacrificing quality. In this tutorial, we smashed FLUX.2 Klein 4B, one of the latest open-source models, and used FORA, TorchAO, and torch.compile to achieve a 2–3× speedup.
Introducing the Pruna Build Program 🚀 Most AI teams hit the same wall: they build a great product, but inference costs make scaling hard. This is for teams building something real and thinking about cost, speed, and efficiency. Start with $300 in free credits and unlock more as usage grows.
First Prune is almost over 🔥 This is our last call for anyone who wants to contribute to Pruna OSS. Join before April 30: pick an open issue, ask to be assigned, and submit your PR. 👉 Search for your issue: github.com/PrunaAI/prun... 👉 Read more about Pruna OSS: www.pruna.ai/blog/first-p...
We launched two speedy diffusion language model endpoints. llada2.1-mini hits ~1000+ TPS; llada2.1-flash reaches ~800+ TPS for fast, high-volume language tasks. See them on Replicate: replicate.com/prunaai/llad... ⭐️ Star our OSS: github.com/PrunaAI/pruna
prunaai/llada2.1-mini – Replicate
The fastest diffusion language model with up to ~1000+ tps
replicate.com
Loved the energy at our AI Efficiency Meetup in Munich last Wednesday! Special thanks to our speakers, Hamza Tahir and Begüm Cig, for their time, insights, and great discussions. And of course, thanks to everyone who came by, asked thoughtful questions, and made the event so enjoyable.
We just launched two new models on Replicate! • ERNIE-Image → higher quality (~33s) • ERNIE-Image-Turbo → faster iteration (~6s) No prompt engineering required — our built-in prompt enhancer handles it. Try them: replicate.com/prunaai/erni... replicate.com/prunaai/erni...
prunaai/ernie-image – Replicate
ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu
replicate.com
Pruna welcomes Adrien Fradin as ML Research intern. He’s finishing his Master at MVA after Applied Math at École polytechnique. Excited to have his expertise on board. 📚 API & quickstart: docs.api.pruna.ai/guides/quick...
First Prune is halfway 🚀 with something special: - A recap blog post about our OSS journey: how we started, what we’ve built, and what’s next. - And a surprise: now each merged PR earns 60 credits. 👉 Read the blog: www.pruna.ai/blog/first-p... 👉 Search for your issue: github.com/PrunaAI/prun...
Join us in Munich tomorrow, April 15, for our AI efficiency meetup! We’ll hear from: Hamza Tahir from ZenML Begüm Cig from Pruna AI 👉 Learn more and sign up: buff.ly/LLLWHDL
Pruna AI’s optimized fast endpoints for GPT-OSS-20B and 120B are now available. Faster model responses at replicate.com/prunaai/gpt-... and replicate.com/prunaai/gpt-.... Give them a try! ⭐️ Star our OSS: github.com/PrunaAI/pruna
prunaai/gpt-oss-20b-fast – Replicate
Advanced 20B open-weight reasoning models to customize for any use case and run anywhere.
replicate.com
Qwen-3.5-35B-A3B-Fast performance up with MoE kernel tuning—throughput doubled and TTFT reduced by 30% compared to competitors. Ready for high-load AI serving! Details at replicate.com/prunaai/qwen...
replicate.com
Mehdi Si-mohammed from ENS Paris-Saclay joins Pruna as an ML Research intern. His expertise in HPC and applied math strengthens our AI research efforts. Welcome aboard, Mehdi! 📚 API & quickstart: docs.api.pruna.ai/guides/quick...
First Prune is in progress 🌱 If you want to join, there’s still time! Each merged PR in April earns 30 Pruna API credits. Don't hesitate to contact us if you have questions— we’ll be here to help. 👉 Explore the open issues here: buff.ly/qVPCFdb ⭐️ Star our OSS: github.com/PrunaAI/pruna