LIVE NOW from the NYSE: David Kanter, Co-Founder of MLCommons & Head of MLPerf, is on theCUBE, NYSEWired, talking AI inference benchmarks. We just dropped MLPerf Endpoints v0.7 - a new way to measure GenAI service performance in real-world deployments. v1.0 later this year. Tune in: thecube.net\
MLCommons
@mlcommons.org
MLCommons is an AI engineering consortium, built on a philosophy of open collaboration to improve AI systems. Through our collective engineering efforts, we continually measure and improve AI technologies' accuracy, safety, speed, and efficiency.
MLPerf Endpoints v0.7 is live - a foundation release for AI inference benchmarking. Initial results from Coreweave, Google, Intel, KRAI, and NVIDIA. Four principles: Current, Comprehensive, Comparable, Commentary. Blog: https://mlcommons.org/2026/07/mlperf-endpoints-v0-7-release/
Benchmark a brain tumor AI model on real patient MRI data. No data leaves the hospital. No model weights exposed. That's MedPerf + Google Cloud Confidential Space, demonstrated live at #GoogleCloudNext 2026. mlcommons.org/2026/06/medp...
There's an AI Reliability Map. Most of it is still empty. Benchmarking clusters in a few cells. Enterprise AI readiness needs the full grid. AIRR is mapping what others skip: https://mlcommons.org/2026/04/airr-map/ #AIReadiness #EnterpriseAI #AIBenchmarking
MLCommons is introducing an Edge Agentic Inference benchmark for MLPerf Inference v6.1. Single accelerator. One user. Multi-turn tool-calling. Hard 32K context wall. Model: Qwen3.6-27B Q4_K_M Accuracy gate: BFCL v4 Deadline: 7/31/26 https://mlcommons.org/2026/07/mlperf-inference-v61-edge-agentic/
MLPerf Inference now measures multi-turn agents. 990 trajectories, Kimi K2.6 + Qwen3.6-35B-A3B, Pareto-curve performance, three-level accuracy. Built on MLPerf Endpoints. https://mlcommons.org/2026/07/agentic-inference-for-mlperf-inference/ #MLPerf #AgenticAI #LLM
AI is in doorbells, hearing aids, and factory sensors. But how do you fairly compare a $1 MCU against a neural accelerator? MLPerf Tiny v1.4: 9 orgs, 25 configs, all measured the same way — standardization makes progress possible. https://mlcommons.org/2026/07/mlperf-tiny-v1-4-results/
Standardized benchmarking is the only way to compare AI systems fairly at scale. MLPerf Training v6.0 results are live, featuring: -11,000+ accelerator systems -New first-time submitters -Verified performance on the world's most demanding workloads 🔗 https://bit.ly/4faJ8mR
MLPerf Training v6.0 results are live! 🎉 For the first time: two Mixture-of-Experts (MoE) benchmarks reflecting where the AI training frontier actually is. 📍 DeepSeek V3 — 671B params (largest in MLPerf history) 📍 GPT-OSS 20B — 21B params Results: mlcommons.org/2026/06/mlpe... 1/4
Tonight at #VLSI2026: MLCommons' David Kanter joins the Evening Panel "AI: Grand Vision or Grand Delusion?" alongside panelists from AMD, SK Hynix, Rapidus & Oxmiq Labs. 8–10 PM, Tapa 1-3.
MLPerf Mobile v6.0 introduces new generative AI benchmarks for running LLMs (Llama 3.1 & 3.2, including the new 1B and 3B models) natively on mobile devices. Test your on-device inference performance. Available on GitHub, iOS & Android: https://bit.ly/43dlMGE
30 years of coordinated disclosure, one assumption: you can fix the thing once you find the flaw. Open-weight AI breaks that. A new version isn't a patch — every prior copy persists, indefinitely. We're helping write the standard AI evaluation needs. → https://bit.ly/43t8R3t
Meet GeoCroissant. Built on MLCommons Croissant, it adds Earth observation-specific metadata—from coordinate systems to spatial resolution—to give you better traceability and more reproducible workflows for agentic AI pipelines. https://bit.ly/3PTLywz
AI systems co-design is too fragmented. Enter MLCommons Chakra (#MLSys2026): an open execution trace ecosystem to bridge software & hardware without exposing IP. Native in @PyTorch, NVIDIA NeMo, & vLLM. Read the paper & explore the traces: https://bit.ly/4vkYZEP
The median AI benchmark longevity score is 5/100. AILuminate scored 75—but even that degrades over time. To fix this, the @MLCommons AIRR team built the Continuous Prompt Stewardship System to keep risk evaluation fresh and reliable. https://bit.ly/3On4jrz
What does AI reliability actually require? It comes down to consistently following the right behavioral rules—even under adversarial attack. Meet the AI Reliability Map to guide pre-deployment testing. Explore the framework: https://bit.ly/4mG7erO #AIReliability #AI
Do tools like OpenClaw signal a turning point for mainstream AI adoption? MLCommons' Dave Graham debated that and more on the Utilizing AI podcast. What do you think? https://bit.ly/4uJj4Va #AgenticAI #AI
The Future of Agentic AI: Opportunities, Risks, and Society | Utilizing AI Episode 22
YouTube video by Utilizing AI Podcast - The Futurum Group
youtu.be
MLPerf Training v6.0 has added GPT-OSS 20B. With 21B total parameters (but only 3.6B active per token), this new sparse MoE pretraining benchmark is designed specifically for accessibility—it can run on a single 8-GPU node. https://bit.ly/4noRr14
AI Risk and Reliability certification shouldn't be a self-assessment. That's the premise behind the AILuminate Global Assurance Program (GAP). GAP gives organizations an independent path to certify that their AI systems meet established safety standards. https://bit.ly/4kIS18x
MLPerf Endpoints: decoupled client, any endpoint, zero-effort integration. Cloud or bare-metal — evaluated equally. Built for API-first GenAI. https://bit.ly/3Pjx34u #MLPerf
Great to see Microsoft highlighting the need for global collaboration on AI safety testing—and shouting out the MLCommons community’s ongoing work to expand the AILuminate benchmarks for multilingual and multimodal testing. https://bit.ly/3RdYFZG
Advancing AI evaluation with the Center for AI Standards (US) and Innovation and the AI Security Institute (UK) - Microsoft On the Issues
Today, Microsoft is announcing new agreements with the Center for AI Standards and Innovation (CAISI) in the US and the AI Security Institute (AISI) in the UK to advance the science of AI testing and ...
bit.ly
The New Wave of AI in Healthcare 2026 symposium kicks off today in NYC! 5/13 at 10:50 AM, MLCommons' Andrew Gruen, PhD will be taking the stage. If you're attending, don't miss this conversation on trust, accountability, and AI validation in medicine. https://lnkd.in/efz2t-Ja
AI software optimization is now moving faster than hardware cycles. To capture these rapid gains, MLPerf is shifting to a rolling submission cadence. David Kanter explains why this speed matters for enterprise buyers via Nutanix: https://bit.ly/3R24FVt #MLPerf #AI
Measuring AI Performance Shifts to APIs | The Forecast
MLCommons cofounder David Kanter explains how the MLPerf benchmark has been overhauled to measure AI performance via API endpoints, reflecting the shift toward rented and hybrid AI infrastructure.
nutanix.com
Submissions for MLPerf Training v6.0 are open! This round brings updates, including the introduction of large-scale MoE pretraining architectures. Whether benchmarking on a single 8-GPU node or a massive cluster, we want your results in this round. https://bit.ly/4uG3vNS
We're thrilled to welcome Flower AI to MLCommons to help shape standards for federated AI at scale. First up: MedPerf is integrating with Flower, enabling researchers to run federated clinical AI studies without moving sensitive patient data. https://bit.ly/4nt1x0T
Measuring today’s production workloads is getting harder. The Inference working group stepped up by adding GPT-OSS 120B, DeepSeek-R1, and our first text-to-video generation benchmark. https://mlcommons.org/2026/04/mlperf-inference-v6-0-results/
MoE benchmarking doesn't have to require a supercomputer. MLPerf Training v6.0 introduces GPT-OSS 20B: a sparse Mixture-of-Experts pretraining benchmark that can run on a single 8-GPU node. See how the task force engineered away statistical variance (CV < 5%): https://bit.ly/3QLwvVU #MoE #AI
Mixture-of-Experts (MoE) is coming to MLPerf Training v6.0. The new DeepSeek-V3 large-scale pretraining benchmark captures critical innovations like MLA, fine-grained expert segmentation, and MTP at production scale (671B parameters). Technical details: bit.ly/49bRabO
Security theater vs. rigorous AI benchmarking - the difference is methodology. AILuminate Jailbreak v0.7: a mechanism-first taxonomy for single-turn jailbreak attacks. Defensible. Reproducible. Auditable. https://mlcommons.org/2026/02/jailbreak-0-7/ #AILuminate #AISecurity
The New Wave of AI in Healthcare 2026 - May 12-13 in NYC. MLCommons' Andrew Gruen, PhD, is speaking on May 13. Register: https://lnkd.in/efz2t-Ja #AIinHealthcare