Rémy Lérignier

@rlerignier.bsky.social

Département de l'appui à la recherche. Service commun de documentation. Université de Poitiers. X(Twitter) : @rlerignier ; Mastodon : @rlerignier@mamot.fr.

Extrait saignant : Les grandes enquêtes du Project Information Literacy (2009, 2010) montrent que les étudiants adoptaient une stratégie prévisible et peu diversifiée, fondée sur un petit ensemble de ressources familières (cours, Google, WP), plutôt que d'explorer l'éventail des ress. dispo. 6/

Mais pour la recherche, oui, la navigation de lien en lien, la sérendipité, sont des pratiques passées de mode. Les documentalistes essaient pourtant de former les jeunes juristes à lire les notes de bas de page et aller les chercher, ce qui s'apparente à une navigation par les liens. 2/2

Très intéressant (l'article et les remarques de @precisement.bsky.social ). On arrive à un moment où on dira bientôt de la recherche web comme de la recherche en bibliothèque qu'il est loin le temps où on trouvait des choses qu'on ne cherchait pas (la beauté d'internet) et où on pouvait se perdre.

precisement.org@precisement.bsky.social · yesterday

Les IA vont-elles tuer les liens hypertextes, ces briques fondamentales du Web ? www.lemonde.fr/pixels/artic... La réponse est oui. Et cette évolution est en cours depuis très longtemps. 1/

goes on to more advanced stuff eg late interaction, learnt sparse & rerankers (cross-encoders, LLMs as rankers). Having fun adding library specific examples, warning abt common librarian confusions etc.. But as usual, the more I work, the more i want to complicate. But i try to hide it in footnotes.

Fun project, writing a "textbook" of sorts to explain to librarians what I feel is needed to understand Information retrieval.. Starts from boolean, explains lexical is not just boolean, covers ranked lexical retrival - bm25/tf-idf, moves on to simple dense embedding, how it typically trained etc

Bild

Avant dernière étape du Tour de France l'année prochaine (pure fiction ;-) : Après une terrible accélération qui a laissé sur place ses adversaires au refuge de Tête Rousse, Pogačar franchit seul en tête le sommet du Mont Blanc avec plus de 12 minutes d'avance sur ses poursuivants.

On passe déjà des plombes à cliquer sur "ne pas accepter les cookies" alors qu'il existait déjà une solution technique déployée (le "Do Not Track"). Le législateur est toujours une pintade quand il s'agit de réguler et réglementer les choses. Le problème est là.

On va passer notre vie à désactiver des agents IA et remplir des formulaires de 20 pages pour refuser l'usage de nos données personnelles par une IA.

#ESR 17 juillet 2026: 6 heures de réunion, 80 mails reçus dont 20 pour caler des réunions à la rentrée. Toujours impossible de dégager assez de temps pour finaliser une révision d'articles. Je suis désolé mais il y a 10 ans ce n'était pas comme ça.

Pensée pour l'intervenant (égaré ?) sur le plateau de CNews qui dit que le Conseil d'État ne va pas devenir un repaire de gauchistes. Manifestement, il est bien seul à "penser" ainsi sur le plateau. Le drapeau du Conseil d'état pour CNews ;-)

Drapeau rouge dont la hampe est tenu par un poing fermé

arXiv now hosts over 3 million articles – arXiv blog https://blog.arxiv.org/2026/07/09/arxiv-now-hosts-over-3-million-articles/

On July 1st, 2026, arXiv reached an important milestone – establishing ourselves as an independent nonprofit. But only a few months before, arXiv quietly passed a different milestone – arxiv.org now hosts over 3 million scientific articles. Back in 2022, arXiv founder Paul Ginsparg predicted it would likely take four and half years for arXiv to pass 3 million articles. arXiv did it in four. In April 2026 arxiv.org surpassed 3 million articles. Only the month before, arXiv broke its monthly submission record, receiving 30,045 new submissions in March 2026. This record was quickly surpassed in May 2026, when we received 31,604 new submissions and then again in June 2026, when we received 32,040 new submissions in a single month. arXiv was founded in 1991 to help scientists quickly and freely share their work with the world. arXiv was a “big idea” but we started small – one computer on one scientist’s desk at the dawn of the World Wide Web. In our very first month of existence, arXiv received less than 30 submissions. Since then, arXiv’s growth has exploded, with submissions counts continuously climbing year after year. arXiv hit its very first milestone in 2008, when we reached 500,000 papers hosted. While it took 17 years from our founding to get there, it took less than 7 years to double that milestone, reaching the 1 million article mark just after Christmas in 2014. Fast forward another 6 years, and arXiv had once again doubled its article count, reaching the 2 million mark in early 2022. In the year leading up to arXiv’s 2 million submission milestone, arXiv received 181,630 new submissions, with a monthly average of 15,135 – almost double what it was only 5 years before. arXiv’s rapid growth is fueled by many things. arXiv is open to all scientists and is used by graduate students and Nobel Prize winners alike. With over 5 million active users visiting arxiv.org every month, arXiv is where researchers discover new ideas first, every day, across the globe. arXiv is such a foundational infrastructure for sharing research that many papers appear *only* on arXiv. Researchers in subject areas like physics, math, and computer science count on arXiv to keep pace with the explosion of new research in their fields. As a community resource that is free to use, arXiv relies heavily on the scientific community, our 230 member institutions, as well as our sponsors, affiliates, and major funders. arXiv is not possible without the dedication of the researchers who use arXiv every day, hundreds of volunteer moderators, our volunteer councils, and the dedicated arXiv staff. arXiv’s submission growth shows no signs of slowing down – having passed 3 million articles in early April, three months later we’re already closing in on 3.1 million articles. At the current rate of arXiv’s monthly submission growth, it’s likely arXiv will pass 4 million articles in less than 3 years. arXiv is thankful for our early days at Los Alamos National Laboratory, and our over 25 years of growth and development with the support of Cornell University, which have created a strong community and a sustainable foundation for arXiv to continue to grow and scale independently. Thank you to everyone who helped arXiv reach 3 million articles!

blog.arxiv.org

I always like to say AI/LLM because a lot of the time, if all you can see the final result, it is hard to tell if the system is using LLM (think GPT) or some other form of "AI", traditional ML/NLP. Take for example Primo's Natural language search and Web of Science smart search both by Clarivate.(1)