Now that the rich and powerful people are calling for open-weights models, it's time to pressure them to be legitimately brave and back open-corpus models.
jensen one of us, let's bring him over on bsky xcancel.com/i/status/208...
Gaël Varoquaux
@gaelvaroquaux.bsky.social
Research & code: Research director @inria ►Data, Health, & Computer science ►Python coder, (co)founder of scikit-learn, joblib, & @probabl.bsky.social ►Sometimes does art photography ►Physics PhD
Now that the rich and powerful people are calling for open-weights models, it's time to pressure them to be legitimately brave and back open-corpus models.
jensen one of us, let's bring him over on bsky xcancel.com/i/status/208...
Materials for my course on machine learning for health data: runnable examples of ML on health data, to get people thinking about model flexibility and biases gael-varoquaux.info/health_ml_tu...
An introduction to Machine Learning for health and epidemiology
Understand concepts important to Machine Learning in Health with notebooks that run on real health data, giving the practical elements to tackle the complexity of real statistical learning question...
gael-varoquaux.info
📢 Contre la présomption de légitime défense pour les forces de l'ordre ! La France est déjà le pays d'Europe comptant le plus grand nombre de personnes tuées par des agents de la force publique. Ce texte aggravera ce bilan. ✒️Signez la pétition : petitions.assemblee-nationale.fr/initiatives/...
Contre la présomption de légitime défense pour les forces de l'ordre. - Contre la présomption de légitime défense pour les forces de l'ordre. - Plateforme des pétitions de l’Assemblée nationale
Le 7 juillet 2026, l'Assemblée Nationale est appelée à se prononcer sur la proposition de loi n°691, portée par le député Eric Pauget (LR), visant à reconnaître une présomption de légitime défense pou...
petitions.assemblee-nationale.fr
A fun podcast with Maxime Gabella about our dreams about AI. The next frontier of AI is building agents that use statistical learning to tackle problem-specific challenges from data. Agents come with promises, but they need a statistical harness, as @probabl.ai's www.youtube.com/watch?v=cNpv...
Can AI Become a Real Data Scientist? | Gaël Varoquaux on scikit-learn, Probabl & Scientific Judgment
YouTube video by Maxime Gabella
youtube.com
🏹 Job alert: Postdoc on Tabular Foundation Models at Inria 📍 Palaiseau 🇫🇷 ⏰ 20 August 🔗 https://bit.ly/4eN1j0N
🧑💻🧑🏫 I'm recruiting a post-doc to work on Tabular Foundation Models, one of the hotest topics in AI, where we are at the leading edge team.inria.fr/soda/files/2... This is an opportunity to develop the next-level tabular AI, blending deep learning and tables.
team.inria.fr
This article in Le Monde about French research in general and the CNRS in particular has a few depressing graphs www.lemonde.fr/sciences/art...
You're only a real statisticien if you can say "Heteroscedasticity" without flinching
Vous avez systématiquement voté contre tout projet de loi écolo. Vous avez stigmatisé les militants en « eco-terroristes ». Vous avez coupé les crédits verts des budgets. Vous avez nommé des ministres pro énergies fossiles et pesticides. Vous avez parlé «d’écologie punitive». Ne venez pas pleurer
🎉 Scikit-learn 1.9 released: ■Solid improvements to many existing estimators: faster, more stable, handling missing values, adding GPU support… ■Also, enhanced estimator displays in notebooks, ■And callbacks that enable progress bars or monitoring of convergence blog.scikit-learn.org/updates/rele...
scikit-learn release 1.9: better numerics, new core functionality
Author: Gael Varoquaux
blog.scikit-learn.org
The CfP deadline for Compute! Paris 2026 was extended to Sunday, June 7! Just a few days left to submit a proposal on Open Source scientific compute, data science, ML & AI topics. Conference dates and venue: November 25–26, 2026, Sorbonne Université · Paris compute.events/paris2026/cf...
Call for Proposals — Compute! Paris 2026
Submit your talk proposal for Compute! Paris 2026. The Call for Proposals is open from April 15th to June 7th, 2026.
compute.events
Modern AIs tackle different questions than data science on industry or scientific applications. Statistical thinking, front and center in data science, is often hidden in AI. But it’s just as crucial. And too often, we treat data science or AI as merely a programming exercise
FAQ on NeurIPS Europe: NeurIPS Europe is an official NeurIPS 2026 satellite event taking place in Paris, France, alongside the main conference in Sydney and the other satellite event in Atlanta. NeurIPS authors can present their papers at any of the three locations, subject to space availability.
- It is now possible to pass arguments to the scorers in Data Ops, such as sample weights. - Diagrams for the Learner and parameter searches now include the full DataOp graph in their notebook repr. - It is now possible to find nodes by name in the DataOp graph.
- fuzzy_join and Joiner now allow to choose the metric that should be used for matching. - ApplyToCols now has the exclude_cols parameter, to define which columns should not be transformed.
- The TableReport now uses plot_distributions and compute_associations to control the distribution and association tabs respectively. - The cleaner now allows to control whether numeric-looking strings ("['1', '2', '3']") should be parsed to float.
✨ Skrub version 0.9.0 has been released ✨ This release adds some advanced features to the Data Ops, the has_dtype() selector, as well as some clarity improvements for the Cleaner and TableReport. Release post: github.com/skrub-data/s...
Release 0.9.0 · skrub-data/skrub
✨ Skrub version 0.9.0 has been released ✨ Main changes Scorers used by Data Ops can now take additional arguments (like sample weights). By @jeromedockes in #1995 The new methods .skb.find() and ....
github.com
#ICLR2026 paper✨️: Quantifying epistemic uncertainty of Blackbox classifiers, and link to better decisions Calibration on steroids, qualifying full prediction uncertainty with no need for Bayes, and tuning individual decisions 👇
"AI amplifies whatever is already there. Good discipline becomes great output. No discipline becomes technical debt at machine speed. Anthropic chose a direction. Go faster. Have Claude check Claude. And when it breaks, go faster still." substack.com/home/post/p-...
Claude Code's Source: 3,167-Line Function, Regex Sentiment
Anthropic claimed 100% of Claude Code is AI-written. A source leak exposed a 3,167-line function, regex sentiment analysis, and 250K wasted API calls daily
substack.com
The team of JupyterCon 2023, PyData Paris 2024 & 2025 organizes a new conference named Compute! Paris 2026 on open source computation and data. The event will take place on November 25–26, 2026 at Sorbonne Université in Paris. CfP deadline: May 24, 2026: compute.events/paris2026/cf...
Call for Proposals — Compute! Paris 2026
Submit your talk proposal for Compute! Paris 2026. The Call for Proposals is open from April 15th to May 24th, 2026.
compute.events
Perles du Sénat. Le ministre de l'ESRE : « je veux rappeler (...) qu'il n'y a là aucune discrimination politique [dans les refus d'embauche en ZRR]. La preuve en est que l'immense majorité des ZRR concerne les sciences dites dures, en particulier les technologies. » www.senat.fr/questions/ba...
Recours aux zones à régime restrictif au sein des laboratoires de recherche publics
senat.fr
New skrub release ✨️ I'am really excited about the more general ApplyToCols. I've found that it enables me to write very naturally complex data transformations on dataframes, as I combine it with skrub's selectors to choose which columns I apply transformations on. skrub-data.org/stable/refer...
ApplyToCols
Gallery examples: Getting Started Hands-On with Column Selection and Transformers
skrub-data.org
✨ skrub version 0.8.0 has been released ✨ This version includes several new features, including multiple improvements to the functionality and performance of the Data Ops, along with a few bug fixes and improvements to the docs. Changelog: skrub-data.org/stable/CHANG... Highlights below ⤵️
The minimum required version of polars has been increased from 0.20 to 1.5.
The TableReport custom filters have been improved and expanded: they can now take skrub selectors for filtering columns. The interface has also been simplified.
The has_nulls selector can now select columns based on a user-specified threshold of null values.
It is now possible to provide custom null values to the Cleaner, so that they are marked as nulls (for example, the string "unknown").
The performance of DataOps with many computational nodes has been improved. Additionally, DataOps CV splitters can now take kwargs. For example, this allows to specify groups when creating train/test splits.
The SingleColumnTransformer and RejectColumn classes allow the construction of custom-made transformers for specific use cases.
The ApplyToCols transformer is now a powerful alternative to the regular scikit-learn ColumnTransformer. It is now possible to apply any transformer to a subset of chosen columns using the skrub selectors.
✨ skrub version 0.8.0 has been released ✨ This version includes several new features, including multiple improvements to the functionality and performance of the Data Ops, along with a few bug fixes and improvements to the docs. Changelog: skrub-data.org/stable/CHANG... Highlights below ⤵️
Release history
Release 0.8.0: New Features: The eager_data_ops configuration option has been added. When set to False, no previews are computed and validation is deferred until the DataOp is actually used (e.g. w...
skrub-data.org