MLCommons

@mlcommons.org

MLCommons is an AI engineering consortium, built on a philosophy of open collaboration to improve AI systems. Through our collective engineering efforts, we continually measure and improve AI technologies' accuracy, safety, speed, and efficiency.

LIVE NOW from the NYSE: David Kanter, Co-Founder of MLCommons & Head of MLPerf, is on theCUBE, NYSEWired, talking AI inference benchmarks. We just dropped MLPerf Endpoints v0.7 - a new way to measure GenAI service performance in real-world deployments. v1.0 later this year. Tune in: thecube.net\

Bild

Standardized benchmarking is the only way to compare AI systems fairly at scale. MLPerf Training v6.0 results are live, featuring: -11,000+ accelerator systems -New first-time submitters -Verified performance on the world's most demanding workloads 🔗 https://bit.ly/4faJ8mR

Bild

MLPerf Training v6.0 results are live! 🎉 For the first time: two Mixture-of-Experts (MoE) benchmarks reflecting where the AI training frontier actually is. 📍 DeepSeek V3 — 671B params (largest in MLPerf history) 📍 GPT-OSS 20B — 21B params Results: mlcommons.org/2026/06/mlpe... 1/4

Bild

Tonight at #VLSI2026: MLCommons' David Kanter joins the Evening Panel "AI: Grand Vision or Grand Delusion?" alongside panelists from AMD, SK Hynix, Rapidus & Oxmiq Labs. 8–10 PM, Tapa 1-3.

Bild

MLPerf Mobile v6.0 introduces new generative AI benchmarks for running LLMs (Llama 3.1 & 3.2, including the new 1B and 3B models) natively on mobile devices. Test your on-device inference performance. Available on GitHub, iOS & Android: https://bit.ly/43dlMGE

Bild

30 years of coordinated disclosure, one assumption: you can fix the thing once you find the flaw. Open-weight AI breaks that. A new version isn't a patch — every prior copy persists, indefinitely. We're helping write the standard AI evaluation needs. → https://bit.ly/43t8R3t

Bild

Meet GeoCroissant. Built on MLCommons Croissant, it adds Earth observation-specific metadata—from coordinate systems to spatial resolution—to give you better traceability and more reproducible workflows for agentic AI pipelines. https://bit.ly/3PTLywz

Bild

The median AI benchmark longevity score is 5/100. AILuminate scored 75—but even that degrades over time. To fix this, the @MLCommons AIRR team built the Continuous Prompt Stewardship System to keep risk evaluation fresh and reliable. https://bit.ly/3On4jrz

Bild

MLPerf Training v6.0 has added GPT-OSS 20B. With 21B total parameters (but only 3.6B active per token), this new sparse MoE pretraining benchmark is designed specifically for accessibility—it can run on a single 8-GPU node. https://bit.ly/4noRr14

Bild

AI Risk and Reliability certification shouldn't be a self-assessment. That's the premise behind the AILuminate Global Assurance Program (GAP). GAP gives organizations an independent path to certify that their AI systems meet established safety standards. https://bit.ly/4kIS18x

Bild

The New Wave of AI in Healthcare 2026 symposium kicks off today in NYC! 5/13 at 10:50 AM, MLCommons' Andrew Gruen, PhD will be taking the stage. If you're attending, don't miss this conversation on trust, accountability, and AI validation in medicine. https://lnkd.in/efz2t-Ja

Bild

AI software optimization is now moving faster than hardware cycles. To capture these rapid gains, MLPerf is shifting to a rolling submission cadence. David Kanter explains why this speed matters for enterprise buyers via Nutanix: https://bit.ly/3R24FVt #MLPerf #AI

Measuring AI Performance Shifts to APIs | The Forecast

MLCommons cofounder David Kanter explains how the MLPerf benchmark has been overhauled to measure AI performance via API endpoints, reflecting the shift toward rented and hybrid AI infrastructure.

nutanix.com

Submissions for MLPerf Training v6.0 are open! This round brings updates, including the introduction of large-scale MoE pretraining architectures. Whether benchmarking on a single 8-GPU node or a massive cluster, we want your results in this round. https://bit.ly/4uG3vNS

Bild

We're thrilled to welcome Flower AI to MLCommons to help shape standards for federated AI at scale. First up: MedPerf is integrating with Flower, enabling researchers to run federated clinical AI studies without moving sensitive patient data. https://bit.ly/4nt1x0T

Bild

MoE benchmarking doesn't have to require a supercomputer. MLPerf Training v6.0 introduces GPT-OSS 20B: a sparse Mixture-of-Experts pretraining benchmark that can run on a single 8-GPU node. See how the task force engineered away statistical variance (CV < 5%): https://bit.ly/3QLwvVU #MoE #AI

Bild

Mixture-of-Experts (MoE) is coming to MLPerf Training v6.0. The new DeepSeek-V3 large-scale pretraining benchmark captures critical innovations like MLA, fine-grained expert segmentation, and MTP at production scale (671B parameters). Technical details: bit.ly/49bRabO

Bild