Analyzed a few thousand hours of agentic runs with Jev - turns out it's: • The best option I've tested at measuring progress and estimating completion • Not very good at catching models being lazy Actual prompts and results in the article: www.southbridge.ai/blog/jev-wa...
Models watching models
System One models can make it easier to track and interpret agentic runs - at scale, in real-time.
southbridge.ai