Arthi Arumugam

@themlwitch.bsky.social

In training to become a Machine Learning Witch. Built whatbroke and Hourzero. whatbroke: github.com/arthi-arumugam-git/whatbroke

ran experiment 2 with whatbroke: swapped llama3.2:3b for qwen2.5:3b in the same support agent. same size class, same prompts, temp 0.2. the diff came back with 12 changed findings. last week's 3x size downgrade only produced 3.

I swapped a tool-calling support agent from llama3.2:3b to 1b. Same scenarios, same prompts, one string changed. Then I diffed the trace files to see what the cheaper model quietly changed. 0 breaking, 3 changed, 12 info. The interesting bits are in the changed ones.