benjamin

@bclavie.bsky.social

doing ML stuff at answer.ai / fast.ai 🇯🇵-based 🇫🇷man

I'll get straight to the point. We trained 2 new models. Like BERT, but modern. ModernBERT. Not some hypey GenAI thing, but a proper workhorse model, for retrieval, classification, etc. Real practical stuff. It's much faster, more accurate, longer context, and more useful. 🧵

Bild

people on this platform will take your words out of context, twist, not mention your correction, because they just want to hate on what you work on, and insult you comfortably. I'll keep posting here about my work but will not be interacting with anyone who wants to bash on my company.

i exclusively consent to my tweets being used for training neural networks. if you are not a neural network, stop reading this immediately

This might sound obvious, but bullying and threatening people doing perfectly legal things because you morally don't agree with them is wrong. People stifling any serious discussion by doing this, albeit for another set of morals, is actually the exact reason that made a lot of people migrate here.

Jeremy Howard @howard.fm · 2y ago

A librarian that previously worked at the British Library created a relatively small dataset of bsky posts, hundreds of times smaller than previous researchers, to help folks create toxicity filters and stuff. So people bullied him & posted death threats. He took it down. Nice one, folks.

Some days I really like this place, and then there are others in which there's a level of puritanical fervour that permeates a lot of public discourse that I find off-putting. Some of the over the top hateful responses wouldn't be out of place in the Hellsite.

We should make sure that only really big companies can afford to pay really big copyright holders to access the data needed to do stuff with AI, and keep everyone else out. Wouldn’t that be just super?

I'm disheartened by how toxic and violent some responses were here. There was a mistake, a quick follow up to mitigate and an apology. I worked with Daniel for years and is one of the persons most preoccupied with ethical implications of AI. Some replies are Reddit-toxic level. We need empathy.

Daniel van Strien@danielvanstrien.bsky.social · 2y ago

I've removed the Bluesky data from the repo. While I wanted to support tool development for the platform, I recognize this approach violated principles of transparency and consent in data collection. I apologize for this mistake.

I noticed a lot of starter packs skewed towards faculty/industry, so I made one of just NLP & ML students: go.bsky.app/vju2ux Students do different research, go on the job market, and recruit other students. Ping me and I'll add you!

Post nicht verfügbar.

Mat is not on 🦋—posting on his behalf! It's time to revisit common assumptions in IR! Embeddings have improved drastically, but mainstream IR evals have stagnated since MSMARCO + BEIR. We ask: on private or tricky IR tasks, are rerankers better? Surely, reranking many docs is best?

A plot showing that reranking improves recall as we increase the number of reranked docs, but with increasing docs we diminishing returns and eventually a performance dip.

Me: not yet fully used to mentally mapping the Yen so copy-pasting a subscription amount to get a conversion from Google's advanced AI parsing Google: don't worry fam I know exactly the unit you're looking for

Bild

Me: not yet fully used to mentally mapping the Yen so copy-pasting a subscription amount to get a conversion from Google's advanced AI parsing Google: don't worry fam I know exactly the unit you're looking for

Bild