Kwindla Hultman Kramer

@kwindla.bsky.social

Low, low, low latency. Daily.co and Pipecat.ai

Gemini 2.0 Flash is competitive with GPT-4o on: - TTFT, - instruction following, - function calling, and - natural conversation dynamics. GPT-4o was ahead on all of these attributes by a wide enough margin that using any other LLM for voice AI mostly didn't make sense. Now there's competition!

Memory for voice AI agents (and composite function calling) ... There are several ways to store (and later, retrieve) conversation state. One of the simplest is just to define a couple of functions and use your local filesystem! Here, @chadbailey.net shows how to do that, using Gemini 2.0 Flash.

Sean DuBois is one of my favorite people to talk to about WebRTC, audio and video, designing good libraries, and hacking in general. Sean is the creator of Pion. Pion is an Open Source WebRTC implementation that is influential and very widely used (including at OpenAI, where Sean works).

Maslow's hierarchy of voice AI u are here ⤵️ ◻️◻️◻️◻️🟦◻️◻️◻️◻️ ◻️◻️◻️🟦🟦🟦◻️◻️◻️ ◻️◻️🟦🟦🟦🟦🟦◻️◻️ ◻️🟦🟦🟦🟦🟦🟦🟦◻️ 🟦🟦🟦🟦🟦🟦🟦🟦🟦 Network transport ▶️ Turn detection ▶️ Interruption handling ▶️ Natural voices ▶️ Tool use www.youtube.com/watch?v=tAQW...

Pipecat Flows - open source Voice AI agent builder

YouTube video by Daily

youtube.com

Today's reminder of how early we are in the generative AI/deep learning technology transition: moved a moderately complex prompt to a different LLM and 150% of my evals broke. 150% because evals I didn't even have (but, obviously, needed) broke, too.

iOS + Gemini Multimodal Live + WebRTC Filipi Fuchter added an iOS example to the Pipecat "Simple Chatbot" repo. With the Pipecat iOS SDK, you can build apps that use Gemini Multimodal Live and Gemini Flash with WebRTC, WebSockets, and HTTP networking.

I had a lot of fun talking to Eric Landau about the state of Voice AI at the end of 2024, what's coming in 2025, what the pain points are today if you're scaling voice AI agents in production, and — of course — the importance of data tooling and evals. open.spotify.com/episode/5Fjj...

Kwin Kramer | Building the Future of Real-Time AI with Daily and PipeCat: Insights on Multimodal Systems and Developer Tools

Deep Learning Leaders · Episode

open.spotify.com

Which LLM should you use for your voice agents? The team at Coval does a lot of interesting work with synthetic data and evaluations for voice AI agents. I've been working with them on evals for Gemini 2.0. They wrote up some of their results so far, here: www.coval.dev/blog/scripte...

Scripted Evaluation Framework for Large Language Models: A Controlled Approach to Comparative Analysis - My Framer Site

Coval is a simulation & evaluation platform for voice and chat agents. Start your free trial today!

coval.dev

The voice-to-voice AI Pareto frontier ... Gemini 1.5 Flash occupies an interesting place in the capabilities matrix for voice AI. It's fast, very inexpensive, has a long context window, and has native audio input. I've been experimenting with Gemini a lot. Here's an interesting Pipecat pipeline:

Bild

Team Suparova at the @supabase / @ycombinator hackathon. There was a four-participant limit on the team size. We have five, but two are robots. Last night was a very long session with lots of tiny little screws and some heavy ifconfig action.

Bild

I've been having a lot of fun writing code that uses the OpenAI Realtime API. I wrote up my notes from the past month: getting started, technical details, use cases that are the best fits for this API right now, code snippets you can borrow. www.latent.space/p/realtime-api

OpenAI Realtime API: The Missing Manual

Everything we learned, and everything we think you need to know, from technical details on 24khz/G.711 audio, RTMP, HLS, WebRTC, to Interruption/VAD, to Cost, Latency, Tool Calls, and Context Mgmt

latent.space