We benchmarked local LLMs vs cloud APIs as part of SD-AI (github.com/UB-IAD/sd-ai), open-source AI tooling for System Dynamics modeling at UB-IAD. Local stack won. 91% vs 89% best single cloud model. But the energy finding is the real story. 🧵
Benchmarking System Dynamics AI Assistants: Cloud Versus Local LLMs on CLD Extraction and Discussion
We present a systematic evaluation of large language model families -- spanning both proprietary cloud APIs and locally-hosted open-source models -- on two purpose-built benchmarks for System Dynamics...
arxiv.org