Almost every LLM that reaches a broad audience is quantized—compressed to run cheaply. USC's Emilio Ferrara argues this compression is treated as a safety nonevent, when it should be treated as a change to the deployed system. This has policy implications, he says.
The Model You Audit Is Not the Model You Ship
Almost every LLM that reaches a broad audience is quantized—with under-appreciated ramifications for AI safety, writes Emilio Ferrara.
techpolicy.press