@datascienceweekly.bsky.social
When it comes to Beatles references, scientists can't let it be arstechnica.com/science/2026...
When it comes to Beatles references, scientists can't let it be
New study detects 3,200 references to Beatles song titles and lyrics in academic papers.
arstechnica.com
📝 "Chart Checklist" Level up your data visuals with this expert guide. 👤 Helena Jambor (@helenajambor.bsky.social) 🔗 https://helenajamborwrites.netlify.app/posts/25-6_NCB/ #rladies #rstats #datavis #article #howto
Well philosophy is willing to welcome all these undergrads
I hereby declare the "learn to code" era officially dead: Big declines in the number of people studying computer science in the last year or two 📉 Chart from this week’s edition of our newsletter on AI and the labour market www.ft.com/content/9183...
There is now StanCon 2026 Uppsala playlist www.youtube.com/playlist?lis... At the moment talks by @mjskay.com, @aseyboldt.bsky.social, @paulbuerkner.com, @charlesm993.bsky.social are up, and we'll keep uploading more during the week
StanCon 2026 Uppsala - YouTube
youtube.com
Blog post about building bibliographic superwork clusters to help in catalog discovery. An experiment in using a local LLM to help label how books relate to each other: thisismattmiller.com/post/superwo...
Building Bibliographic Superwork Clusters for Discovery with Local LLMs
Judging work relationships in 88K clusters
thisismattmiller.com
you 👏🏼 cannot 👏🏼 automatically 👏🏼 detect 👏🏼 plagiarism > the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text link.springer.com/article/10.1... 1/
Testing of detection tools for AI-generated text - International Journal for Educational Integrity
Recent advances in generative pre-trained transformer large language models have emphasised the potential risks of unfair use of artificial intelligence (AI) generated content in an academic environme...
link.springer.com
I'd like to announce that at @openathena.ai, @jder.bsky.social and I helped @m2lines.bsky.social release Samudra 2. We scaled this neural ocean emulator to train on 16x the size of data in bytes on the same hardware budget. We can now skillfully predict 8 years of the ocean on a single GPU at a 1/4°
🌊 Samudra 2: A Fast, Cheap AI Ocean Model, Now at the Scale That Matters
M²LInES’ neural ocean emulator now runs multi-year simulations at eddy-permitting resolution on a single GPU, turning a supercomputer-scale…
medium.com
It’s been 10 years since this post and nothing’s changed 😭 www.ethanrosenthal.com/2016/07/20/l...
I'm all about ML, but let's talk about OR | Ethan Rosenthal
You’ve studied machine learning, you’re a dataframe master for massaging data, and you can easily pipe that data through a bunch of machine learning libraries. You go for a job interview at a SAAS com...
ethanrosenthal.com
Tools from operations research solve so many problems and the fact they’re not super widely deployed really tells you that many bottlenecks are not technical
This post is another success of the heuristic "if you've written 10 tweets on a topic, flesh them out into a post on the real internet."
I spent the last month obsessed with finding colors that can't be displayed on a conventional screen. This is what I found. moultano.wordpress.com/2026/06/19/w...
Here it is! A digital archives resource list for the curious and nerdy 💚 I'll keep drafting my own essays and guides, but this is where I'll link people to external material. allmyfriendsarestories.neocities.org/AParchive/re...
Resources
Resources related to digital archives.
allmyfriendsarestories.neocities.org
I’m making a resource page about digital archives that goes from “I’m just curious” to “I wanna start archiving for fun” to “I’m lost in the sauce and willing to read standards” Do y’all have topics or questions you’d like to see me cover? #digitalarchives #digitalpreservation #DH
Second preprint, with Elizabeth Bonawitz, one of my favorite studies children search the hypothesis space more broadly and variably than adults. osf.io/preprints/ps...
OSF
osf.io
I have been keeping a running timeline of data engineering acquisitions since 2022. It is getting hard to keep up. www.ssp.sh/brain/data-...
Data Engineering Acquisitions (2022-2026)
Consolidation in the Data Engineering market is happening quickly. Tools from the Modern Data Stack get unified into bigger Data Platforms.
ssp.sh
Many people dream of traveling to faraway places to experience new cultures and sights. Yes, you learn a ton, you learn new places, and the culture is hard to learn when not traveling. Travel locally, where you are.
Travel Locally, Where You Are
Many people have the dream of traveling to faraway places to experience new culture and sightseeing.
ssp.sh
Context is king, especially in the age of AI, where everyone is connecting their agents to their data stack. But writing SQL was never the hard part. Making it accurate and trustworthy against your warehouse has always been.
Beyond the Semantic Layer: Building a Context Layer for the Agentic Era
A context layer puts your warehouse schema, joins, metric definitions, and business knowledge in one reviewable place so data agents query governed context instead of guessing field names. A look at h...
kaelio.com
data folk: what's your goto open source data viz tool? is it still basically superset or metabase? (i just wanna point it at some data and throw some charts and maybe a map together; deffo not going the "just vibe something" route :-P )