Pruna AI

@prunaai.bsky.social

The AI Optimization Engine to make AI models cheaper, faster, smaller and greener in a snap! https://pruna.ai

Smash your model to make it faster, smaller, cheaper, and greener following just 5 steps: 1. load a pretrained model 2. define optimizations with a SmashConfig 3. apply optimizations with smash 4. evaluate the optimized model with the EvaluationAgent 5. run inference with the optimized model

Bild

Contributing to Pruna is now easier! Whether you’re an OSS expert or contributing for the first time, we’ve improved our contribution guide to make the process clearer. Here’s how to get started 👇

Bild

Pruna OSS Community Spotlight: KVPress by kschwethelm KVPress makes 20 KV cache compression strategies available for causal language models. KVPress compresses the key-value cache during the pre-fill phase, reducing memory usage for long-context inference.

Bild

💜 Simply make AI models faster, cheaper, smaller, greener! The toolkit is designed with simplicity in mind - requiring just a few lines of code to optimize your models. It supports various model types, including LLMs, Diffusion, and more! ⭐️ Give us a star: github.com/PrunaAI/pruna

Bild

Hell Yeah! Runware just published a practical guide for P-Video-Animate and wanted to share it with you! - $0.03/s at 720p and $0.06/s at 1080p - ~5s generation time per second of output video Their guide covers the image/video pairing, prompt steering, and more! runware.ai/docs/models/...

Animating images with a source video — P-Video-Animate API

How to use P-Video-Animate to animate a reference image with the motion, timing, and camera movement from a source video. Covers the pairing rule, optional prompt steering, and four concrete patterns.

runware.ai

🚀 Pruna 0.3.3 is out! Here are a few highlights: - New benchmarking and evaluation tools. - Algorithm and compatibility upgrades. - Support for Python 3.13. - A wide range of bug fixes and reliability improvements. Thank you all for your contribution! 👉 Read the notes: github.com/PrunaAI/prun...

GitHub - PrunaAI/pruna: Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.

Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead. - PrunaAI/pruna

github.com

P-Video-Avatar is now the fastest model for top-quality avatar generation. - It costs 0.025$ per second of video, and takes 1.83s generation time per second of video. - It offers high control, audio support in 20+ languages, and long-form video. 📚 API & quickstart: www.pruna.ai/p-video-avatar

Learn how to make the latest image generation models smaller and faster without sacrificing quality. In this tutorial, we smashed FLUX.2 Klein 4B, one of the latest open-source models, and used FORA, TorchAO, and torch.compile to achieve a 2–3× speedup.

Bild

Introducing the Pruna Build Program 🚀 Most AI teams hit the same wall: they build a great product, but inference costs make scaling hard. This is for teams building something real and thinking about cost, speed, and efficiency. Start with $300 in free credits and unlock more as usage grows.

Bild

Loved the energy at our AI Efficiency Meetup in Munich last Wednesday! Special thanks to our speakers, Hamza Tahir and Begüm Cig, for their time, insights, and great discussions. And of course, thanks to everyone who came by, asked thoughtful questions, and made the event so enjoyable.

Bild