Aaron Tay

@aarontay.bsky.social

Currently on organizing committee of FORCE2026, Singapore, 3-5 June 2026. I'm librarian + blogger from Singapore Management University. Social media, bibliometrics, analytics, academic discovery tech.

The first paragraph explaining how OpenAlex is article first rather than indexing by journal is how Google Scholar and many 200M+ academic web search engines work. It's not intuitive to librarians who are used to traditional databases. Its also why which journals you index is hard to answer for GS

infoDOCKET@infodocket.bsky.social · last wk.

OpenAlex Announces a "Major Cleanup of Journal Records" (via #OpenAlex) blog.openalex.org/a-major-clea... #databases #scholcomm #journals

I think the fundamental split is not between people who are pro-ai and anti-ai. Is more between people who think it is actually powerful/capable and those who think it's just hype. If you disagree on this matter no amount of debate will bridge the gap...

Anna Mills@annamillsoer.bsky.social · last wk.

"Society is not prepared for AI, and that terrifies me. When I speak with politicians, or anyone outside frontier labs, it seems many people still think the capabilities might be marketing." -- anonymous engineer who signed the call for a pause or slowdown. +

I have heard similar stories of researchers asking Fable/Sol to text mine public data (or so they think) and they later realised it somehow managed to hack ir bypass paywalls to access the data...

Anna Mills@annamillsoer.bsky.social · 7d ago

Really helpful article by @klonick.bsky.social. We need regulation of the companies processes for testing and develoment of these models, not just a kill switch in case they go rogue. www.lawfaremedia.org/article/the-...

1/ Rereading some information retrieval literature from late 2010s for maybe the 5th time. It's amazing how much more sense papers make now. I tend to oversimplify when talking about IR is to act like the story began with transformers->BERT then bi-encoder/cross-encoder.

Fun project, writing a "textbook" of sorts to explain to librarians what I feel is needed to understand Information retrieval.. Starts from boolean, explains lexical is not just boolean, covers ranked lexical retrival - bm25/tf-idf, moves on to simple dense embedding, how it typically trained etc

Bild

Working out which views are which isn't obvious to the novice. A long time ago, many *freshmen* started asking me about the ins and outs of bibliometrics! Why? Because they wanted to compare the citation counts of authors on opposing sides to provide "evidence" on who was right (for an assignment)!

Aaron Tay@aarontay.bsky.social · last wk.

12/ Some positions are central. Some are respectable minorities. Some are emerging. Some are obsolete. Some survive mainly because journalists, politicians or influencers keep reviving them.

1/Information literacy is often described as teaching students to “evaluate information”. But we actually teach several different kinds of evaluation, aimed at solving different epistemic problems. The issue I feel is we rarely make those differences explicit.

Last thread i promise. 1/ Powerful AI makes the inside view extraordinarily seductive. Ask an LLM for the strongest case for one side Ask it to expose methodological weaknesses. Ask it to rebut every objection. Soon you may feel that you have mastered the debate.

Aaron Tay@aarontay.bsky.social · last wk.

18/ Information literacy should teach students to question authority without pretending that all authorities are interchangeable. “Authority is constructed” must not become “expertise is optional” - end for now

1/ Information literacy has a confidence problem. We often teach novices to “evaluate the evidence” as though a checklist can turn them into temporary experts.

[Read]A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews - largest collection of Systematic review extracted from OpenAlex. 301,871 SRs more than 4x bigger than next biggest set!

Bild

Hot take. I was never that impressed by peer review papers that was basically 1. run a topic/journal search, 2. generate "analysis" that describes pub years, type of content cited (eg book, article), citation counts etc . Writing a paper to test if you can use AI to do that is....