Rogue Scholar

@rogue-scholar.wisskomm.social.ap.brid.gy

[bridged from https://wisskomm.social/@rogue_scholar on the fediverse by https://fed.brid.gy/ ]

Rogue Scholar follows the Principles of Open Scholarly Infrastructure (POSI)

With this blog post, the science blog archive Rogue Scholar starts the formal process to adhere to the Principles of Open Scholarly Infrastructure (POSI). To do so, an organization has to perform a self-audit of its compliance with the principles, with a focus on _principles_ and not hard _rules_. POSI was updated to version 2.0 this October, with the changes marked up in a separate document. In the coming weeks, I hope to finish this document with feedback from blogs participating in Rogue Scholar, other users, and the Rogue Scholar Advisory Board that meets in November. Below is an overview of where I think Rogue Scholar already adheres to POSI (🟢), some work is still needed (🟡), or where major gaps exist (🔴). I then describe the major gaps and the work that is needed in more detail. ### Governance 🟢 Coverage across the scholarly enterprise 🔴 Stakeholder governed 🟢 Non-discriminatory membership 🟡 Transparent governance 🟡 Cannot lobby 🔴 Living will 🟡 Regular review of purpose and community value ### Sustainability 🟡 Transparent operations 🟢 Time-limited funds are used only for time-limited activities 🔴 Goal to generate surplus 🔴 Establish and maintain financial reserves guided by policy 🟢 Mission-consistent revenue generation 🟢 Revenue generated from services, not data 🔴 Volunteer labour 🔴 Transition planning ### Insurance 🟢 Open source 🟡 Ensure open and secure data accessibility within legal and ethical constraints🟡 Available and preserved 🟡 Patent non-assertion 🟢 Prioritise interoperability and open standards to ensure continuity and resilience ### Stakeholder governed _A board-governed organisation drawn from the stakeholder community builds confidence that the organisation will make decisions driven by community consensus and a balance of interests._ Rogue Scholar has had an Advisory Board since January 2024, but no board-governed organization structure. This is the most critical shortcoming for Rogue Scholar, and work is underway to address this in the coming months. ### Living will _To build trust, organisations should establish and communicate clear commitments regarding their long-term stewardship responsibilities, including the principles by which assets, data, resources, services, and staff would be responsibly transferred to a successor or the organisation or service wound down._ This is another major shortcoming of Rogue Scholar, but can't be addressed until stakeholder governance is in place. By using paid services from the Internet Archive (content archiving) and Crossref (metadata registration and archiving), central aspects of the Rogue Scholar science blog archive will live on even if the service stops to exist. ### Goal to generate surplus _To weather economic, social and technological volatility, organisations and services need financial resources beyond immediate operating costs._ As a Diamond open access infrastructure with no fees to authors or readers, Rogue Scholar struggles to generate surplus even for immediate operating costs. Thanks to the reliance on open source software and the automated workflows ingesting blog content and metadata, operating costs are moderate, and donations are a big help. Going forward Rogue Scholar will make it easier for organizations and individuals to contribute financially, linked to the planned board-governed organizational structure mentioned above. ### Establish and maintain financial reserves guided by policy _Organisations and services should have a clear policy on maintaining financial reserves, including the purpose, minimum and maximum level, and governance of these funds._ Again, this is a consideration for a future board-governed organizational structure. ### Volunteer labour _Organisations that rely on volunteers and their labour should recognise this as a valuable resource for the organisation’s long-term viability, and factor it into sustainability planning and risk management._ Volunteer labour is a central element for Rogue Scholar operations, as all participating blogs are maintained without contributions (time and/or infrastructure) from Rogue Scholar, which focuses on archiving existing scholarly blogs. Running the Rogue Scholar archive also depends on volunteer labour, but good progress has been made migrating the service to the InvenioRDM platform. Rogue Scholar benefits from the work of the InvenioRDM community both for new features and bug fixes. Going forward, more functionalities can be migrated to InvenioRDM, in particular, the extraction of content and metadata from blog feeds, and volunteer labour contributions to Rogue Scholar can be specified. ### Transition planning _Organisations that are heavily dependent on a limited number of individuals should take steps to reduce their dependence on these individuals, including via transition and succession planning, so that the organisation is not at risk of collapse in the event of their departure._ Rogue Scholar was started by me in 2023, and it depends heavily on my personal involvement. While I have no intentions to reduce my involvement in Rogue Scholar, the organization has to evolve and involve more people, e.g. evaluating new blog submissions or working on funding opportunities, and start work on transition and succession planning. ### Blog posts of other organizations following POSI These organizations following POSI have written blog posts about it that have been archived in Rogue Scholar. These blog posts provide good context to what Rogue Scholar tries to achieve with adopting POSI: * Crossref (December 2, 2020) https://doi.org/10.64000/hzemx-j7n79 * ROR (December 16, 2020) https://doi.org/10.71938/n0kg-4k60 * JOSS (February 14, 2021) https://doi.org/10.59349/m5h23-pjs71 * OpenCitations (August 9, 2021) https://doi.org/10.59350/p5czb-8ff81 * DataCite (August 29, 2021) https://doi.org/10.5438/vy7h-g464 * EuropePMC (February 21, 2022) https://doi.org/10.59350/qpyx2-xj167 * Liberate Science (August 2, 2022) https://doi.org/10.59350/n09bh-vbj58 * PKP (May 29, 2024) https://doi.org/10.59350/pkp.11288 There are more blog posts about POSI (e.g. from Crossref), you can find them in the Open Infrastructure community. If your organization has committed to POSI and you have written a blog post about it, reach out if you want your blog archived by Rogue Scholar. Please use Slack, email, Mastodon, or Bluesky if you have any questions or comments regarding Rogue Scholar following the Principles of Open Scholarly Infrastructure, ideally until November 4. Rogue Scholar is a scholarly infrastructure that is free for all authors and readers. You can support Rogue Scholar with a one-time or recurring donation or by becoming a sponsor. ## References 1. POSI Adopters. (2025). _The Principles of Open Scholarly Infrastructure v2.0_. https://doi.org/10.14454/G8WV-VM65 2. Eve, M. P. (2023, July 28). Rules vs. Principles in POSI. _Martin Paul Eve_. https://doi.org/10.59348/7sgt5-1qy45 3. Fenner, M. (2024, February 8). Introducing the Rogue Scholar Advisory Board. _Front Matter_. https://doi.org/10.53731/9yf86-p8541 4. Bilder, G. (2020, December 2). Crossref’s Board votes to adopt the Principles of Open Scholarly Infrastructure. _Crossref Blog._https://doi.org/10.64000/hzemx-j7n79 5. ROR Leadership Team. (2020). _Aligning ROR with the Principles of Open Scholarly Infrastructure_. https://doi.org/10.71938/N0KG-4K60 6. Katz, D. S., Smith, A. M., Niemeyer, K., Huff, K., & Barba, L. A. (2021, February 14). JOSS’s Commitment to the Principles of Open Scholarly Infrastructure. _Journal of Open Source Software Blog_. https://doi.org/10.59349/m5h23-pjs71 7. Di Giambattista, C. (2021, August 9). OpenCitations’ compliance with the Principles of Open Scholarly Infrastructure. _OpenCitations Blog_. https://doi.org/10.59350/p5czb-8ff81 8. Buys, M. (2021). _DataCite’s commitment to The Principles of Open Scholarly Infrastructure_. https://doi.org/10.5438/VY7H-G464 9. Europe PMC Team. (2022, Februar 21). Europe PMC adopts the Principles of Open Scholarly Infrastructure. _Europe PMC News Blog._ https://doi.org/10.59350/qpyx2-xj167 10. Liberate Science. (2022, August 2). Principles of Open Scholarly Infrastructure evaluation (2022). _Liberate Science_. https://doi.org/10.59350/n09bh-vbj58 11. Stranack, K. (2024, May 29). PKP Signs on to the POSI. _Public Knowledge Project_. https://doi.org/10.59350/pkp.11288

blog.front-matter.io

Rogue Scholar now registers DOIs with InvenioRDM

This week, the Rogue Scholar science blog archive started registering DOIs and metadata with Crossref using the InvenioRDM repository platform rather than relying on external tooling. InvenioRDM has of course supported DOI registration with DataCite for a long time, so this adds another option for repositories hosting reports, preprints, dissertations, or other textual documents. Rogue Scholar has registered content with Crossref since its launch two years ago (more than 45,000 blog posts to date), and migrated to the InvenioRDM repository platform 12 months ago, but it always relied on external tooling for Crossref integration. This new Crossref integration into InvenioRDM simplifies the DOI registration workflow. The main difference for Rogue Scholar blog authors and readers is that automatic DOI registration has become faster, happening within a few minutes after a blog post on a participating blog is published or updated. Rogue Scholar checks the RSS feeds of participating blogs for new or updated content every 10 minutes. When new or updated content is found, uploading the post to InvenioRDM and registering/updating a DOI and metadata with Crossref now happens in 1-2 minutes (depending on how many posts have to be updated at the same time). Further improvements in the time it takes to register blog posts with InvenioRDM and Crossref will depend on improving the monitoring of RSS feeds, e.g. by adding a _push_ mechanism (triggered by the blog when content is updated) instead of the current _pull_ mechanism. ### Versioning The InvenioRDM platform supports versioning of content, which is particularly important for datasets and software, but also relevant for text documents such as blog posts. Examples include major new versions of content, possibly after having received feedback via peer review using the Publish, Review, Curate workflow. Other use cases are corrections and retractions. Crossref added version information to its metadata schema with the 5.4.0 release in March, and I am currently testing this functionality in the Rogue Scholar Staging server: One challenge is that the InvenioRDM Crossref integration supports multiple DOI prefixes, so this is a bit more work compared to the DataCite integration only supporting one DOI prefix. In addition, versioning adds complexity to DOI suffix generation. Luckily, the built-in InvenioRDM tooling uses the same DOI naming scheme that Rogue Scholar has used since its launch, but going forward, Rogue Scholar will no longer support custom DOI suffixes that I introduced for WordPress and Substack blogs. Several blogs participating in Rogue Scholar do their own DOI registrations, and those DOIs (what InvenioRDM calls externally managed DOIs) will not be versioned. ### InvenioRDM Crossref DOI registration uses the `commonmeta-py` Python package to generate Crossref XML metadata and a fork of the `invenio-rdm-records` package to integrate with InvenioRDM. After completing versioning support and more testing, the functionality will be merged into the InvenioRDM core functionality, and the discussion with the other InvenioRDM maintainers has started. Soon, every InvenioRDM instance will be able to use Crossref DOI registration with a simple configuration setting (and Crossref member credentials), very similar to how DataCite DOI registration works. ## References 1. Fenner, M. (2024, September 2). Rogue Scholar migrates to InvenioRDM. _Front Matter_. https://doi.org/10.53731/sdazp-kzn55 2. Feeney, P. (2025, March 19). Version 5.4.0 metadata schema update now available. _Crossref Blog_. https://doi.org/10.13003/325070 3. Fenner, M. (2025, January 16). Persistent identifiers, random strings, and checksums. _Front Matter_. https://doi.org/10.53731/6kfyy-nq280 4. Marcum, C. S. (2025, April 8). Drinking from the firehose? Write more and Publish Less (Version 2). _Upstream_. https://doi.org/10.54900/vr8ax-nz653 5. Marcum, C. S. (2025, April 8). Peer-Review for a Blog Post? My Experience with MetaROR. _Upstream_. https://doi.org/10.54900/bymaz-4fw37 6. Hendricks, G., Lammey, R., & Rittman, M. (2022). _Towards a connected and dynamic scholarly record of updates, corrections, and retractions_. MetaArXiv. https://doi.org/10.31222/osf.io/6z7s3

blog.front-matter.io

Rogue Scholar relaunches today

The science blog archive Rogue Scholar relaunched today with a number of exciting new features, including major new software version, new hardware, new look and feed, and new authentication. ### Major new software version Version v13.0 of the InvenioRDM open source repository platform was released today. After running release candidate versions for a few weeks, Rogue Scholar today was relauched with version v13.0. There are numerous changes in this new version described in detail in the release notes, and this will facilitate additional new features and bug fixes going forward. Many of the major changes of v13.0 happened in the backend and are only visible to administrators, including an improved administration panel or audit logs. Other improvements, e.g. subcommunities and collections, have to be enabled and will happen in the next few months. ### New hardware Rogue Scholar is running on new hardware, which makes the service faster, easier to update, and cheaper to run. Using the Kamal deployment tool, Rogue Scholar now runs on dedicated hardware rented from Hetzner and located in Germany instead of via the cloud provider Fly.io. ### New look and few With this relaunch I fixed several long-standing issues with Rogue Scholar blog post list views: * Show the DOI to allow users to jump directly to the blog post instead of needing to go to the Rogue Scholar archived version first, * Show the language (15% of Rogue Scholar posts are in languages other than English), you can also filter search results by language, * Show optional feature images, as is common for RSS feed readers, * Support (a subset of) HTML in titles, e.g. superscript, bold or italic. ### New authentication No authentication is required to read Rogue Scholar content, and blog posts are automatically imported from participating blogs. Rogue Scholar user accounts currently have limited functionality, mainly allowing blog authors to submit blog posts to one or more topic communities, such as R (the programming language), book review, or interviews. Going forward, I will work with the Rogue Scholar community to improve what you can do with user accounts, including registering a new blog with the platform – currently still requiring an external form. For this functionality, Rogue Scholar user accounts have to be easy to manage and secure. The built-in functionality of the InvenioRDM platform for local accounts allows users to self-manage their accounts and reset their passwords. And it allows administrators to block accounts that misbehave. Rogue Scholar has for a while supported login via ORCID accounts. Today I have launched another authentication option, login via passkeys, using a self-hosted Pocket ID service. Passkeys are both easier to use and safer than usernames/passwords, and after a transition period to allow users to link their existing local accounts to ORCID and/or passkeys, Rogue Scholar will disable local accounts on September 15. Please use Slack, email, Mastodon, or Bluesky if you have any questions or comments regarding this major update. ## References 1. Fenner, M. (2025, July 7). Upgrading to InvenioRDM v13. _Front Matter_. https://doi.org/10.53731/dd5h7-z5y55 2. Fenner, M. (2025, June 27). Kamal deploys InvenioRDM Starter to production. _Front Matter_. https://doi.org/10.53731/m7gng-jmm19

blog.front-matter.io

Kamal deploys InvenioRDM Starter to production

InvenioRDM is the open source turn-key research data management platform, with detailed documentation available here. InvenioRDM Starter facilitates deployment and configuration of InvenioRDM, allowing you to run InvenioRDM on your local computer within 15 min. This is achieved by providing a) a prebuilt Invenio-App-RDM Docker image, and b) a Docker Compose configuration file with sensible defaults. Starting this week, InvenioRDM starter can also be used to deploy InvenioRDM to production, using the Kamal tool. Kamal is similar to Docker Compose, but adds important functionality, including automatic remote builds, zero-downtime deployments, and deployments to multiple servers. Kamal is a command-line utility with a YAML configuration file, and much simpler to use than Kubernetes or commercial Docker container orchestration services such as Amazon Elastic Container Service (Amazon ECS). Kamal can deploy InvenioRDM to your hardware or to a virtual machine provided by your organization or a cloud provider. Whereas Kubernetes is a good option for large InvenioRDM installations, smaller InvenioRDM instances benefit from simpler deployment tools both in terms of cost and required maintenance. The science blog archive Rogue Scholar managed by Front Matter is a good example of an InvenioRDM repository that can benefit from simpler deployment options. As the next major release of InvenioRDM (v13.0) will happen in the next few weeks, Rogue Scholar needs to prepare for the upgrade, and I have this week launched a Rogue Scholar staging instance at https://staging.rogue-scholar.org using Kamal and a virtual machine provisioned by Hetzner and located in Germany. The setup was mostly straightforward, except for the integration with the Kamal proxy server, which turned out to be very painful. In the end I had to set the InvenioRDM`APP_ALLOWED_HOSTS` ENV variable to `None` and patch the REST API cross site request forgery (CRSF) check to not check the request host. This needs more discussion but appears safe, as all requests must go through the Kamal proxy, where the host header is already checked. More work is needed on the staging server, including regular automatic backups of the database, and setting up monitoring (logs and metrics). The instance is running the latest stable release (v12.1.0), but I will soon be able to install the latest v13 release candidate – v13.0.0rc2 was released three days ago. Kamal was released in 2023 by 37signals, the company behind the Basecamp and Hey services, and one of the major contributors to the Rails platform. Kamal is installed as a Ruby gem, but is not specific to Rails or Ruby. Kamal can be seen as the successor to the Capistrano deployment tool, also originally written by 37signals, but Kamal is working with Docker containers. When I was the technical lead of the Article Level Metrics project at the publisher PLOS 2012-2015 (at the time Docker was not yet adopted for production deployments), I made heavy use of Capistrano. InvenioRDM Starter now includes a Kamal configuration option, and I deployed an InvenioRDM instance to https://demo.front-matter.io using Kamal. Feel free to play around, but only admin accounts can create records – use the official InvenioRDM demo instance (also linked in the footer) if you want to create and/or update records. I will spend the next few weeks refining the Kamal setup and documentation, so that InvenioRDM Starter is ready for Kamal deployments when InvenioRDM v13.0 is officially released. ## References Fenner, M. (2024, June 17). Announcing InvenioRDM Starter Beta. _Front Matter_. https://doi.org/10.53731/jxecm-0me48 Fenner, M. (2015, July 29). Thank you PLOS. _Front Matter_. https://doi.org/10.53731/r294649-6f79289-8cvzn

blog.front-matter.io

Rogue Scholar Authorship Guidelines

Rogue Scholar archives the content of currently more than 150 science blogs with more than 40,000 blog posts. In this blog post, I want to clarify the guidelines that Rogue Scholar tries to follow regarding authorship. Rogue Scholar blog posts are scholarly content and thus follow the same basic guidelines as other scholarly outputs, such as journal articles, preprints, or book chapters. ### Authorship All authors are expected to have made substantial contributions to the submitted work and to be accountable for the work both before and after publication. Those who contributed to the work but do not meet the criteria for authorship can be mentioned in the Acknowledgments. ### Artificial Intelligence (AI) AI tools cannot meet the requirements for authorship, as explained by the Committee on Publication Ethics (COPE): > AI tools cannot meet the requirements for authorship as they cannot take responsibility for the submitted work. As non-legal entities, they cannot assert the presence or absence of conflicts of interest nor manage copyright and license agreements. And COPE recommends that: > Authors who use AI tools in the writing of a manuscript, production of images or graphical elements of the paper, or in the collection and analysis of data, must be transparent in disclosing in the Materials and Methods (or similar section) of the paper how the AI tool was used and which tool was used. Rogue Scholar blogger Mark Dingemanse recently made a strong case for why synthetic text is incompatible with science blogging. ### Contributor Roles For blog posts with multiple authors, Rogue Scholar plans to add support for the Contributor Role Taxonomy (CRediT). And for blog posts handled by an editor or undergoing peer review, Rogue Scholar also wants to add those roles. ### Possible Actions Rogue Scholar is an archive of science blog posts, the content is originally published elsewhere, and the decision for publication was taken by the blog authors. In rare cases, blog authors might retract a blog post or post a correction, and that information should also be communicated by Rogue Scholar and via the DOI metadata. When blog posts don't follow the above guidelines, e.g. when inappropriately using AI Tools, Rogue Scholar staff, after consultation with the Rogue Scholar Advisory Board, will decide on appropriate actions, including retraction. ## References 1. McNutt, M. K., Bradford, M., Drazen, J. M., Hanson, B., Howard, B., Jamieson, K. H., Kiermer, V., Marcus, E., Pope, B. K., Schekman, R., Swaminathan, S., Stang, P. J., & Verma, I. M. (2018). Transparency in authors’ contributions and responsibilities to promote integrity in scientific publication. _Proceedings of the National Academy of Sciences_ , _115_(11), 2557–2560. https://doi.org/10.1073/pnas.1715374115 2. _Authorship and AI tools_. (2024). Committee on Publication Ethics. https://doi.org/10.24318/cCVRZBms 3. Holcombe, A. O. (2019). Contributorship, Not Authorship: Use CRediT to Indicate Who Did What. _Publications_ , _7_(3), 48. https://doi.org/10.3390/publications7030048 4. Marcum, C. S. (2025, April 8). Peer-Review for a Blog Post? My Experience with MetaROR. _Front Matter_. https://doi.org/10.54900/bymaz-4fw37 5. Dingemanse, M. (2025, May 2). Why synthetic text is incompatible with science blogging. _Front Matter_. https://doi.org/10.59350/63b1y-1js90

blog.front-matter.io

DOI registration workflow for a science blog (version 2)

_This post is an updated version of the_ _DOI registration workflow for a science blog_ _post I published in September 2023. It reflects the best practices used by the Rogue Scholar science blog archive and contains one important announcement._ In previous blog posts such as the one published earlier, I discussed the various elements involved in registering a DOI for a science blog post. Briefly, the Rogue Scholar service takes advantage of the fact that blogs * use RSS feeds (or the Atom or JSON Feed format) and/or JSON APIs to distribute content and metadata at the time of publication, * these feeds contain the most important metadata needed for publication – such as title, authors, publication date, and * addition metadata (such as abstract and references) can be automatically extracted from the full-text content included in the feed. DOI registration itself has technical (generating metadata that conforms to a specific schema) and business (membership in a DOI registration agency such as Crossref) requirements that are not trivial, so ideally and unless the blog is publishing a lot of content similar to a journal, it is handled by a dedicated service — Rogue Scholar. This basic workflow can be optimized in many ways, such as including funding information, but one fundamental issue remains to be solved: how does the blog learn about the DOI registered for a new post and automatically add it to the blog? There are two basic approaches: a) generate a random DOI and communicate this back to the blog, or b) let the blog pick the DOI, following some basic rules. Most importantly that the DOI is unique, but ideally is a relatively short string without special characters that can easily copy/pasted, and that the DOI is opaque, i.e. contains no meaning that becomes problematic over time. Before January 2025, Rogue Scholar was using the first workflow, i.e. generate a random DOI and communicate this back to the blog via the Rogue Scholar API and website. ## Canonical URL As much as possible Rogue Scholar takes advantage of technologies that have existed for a long time and are not specific to scholarly content. That's why the service works with existing blogs that use standard blogging software - currently eleven different platforms, the most popular being Wordpress, Blogger, and Hugo. These platforms don't know about DOIs without extra work, but they all know about a similar concept: canonical URLs. Wikipedia explains: > A **canonical link element** is an HTML element that helps webmasters prevent duplicate content issues in search engine optimization by specifying the "canonical" or "preferred" version of a web page. It is described in RFC 6596, which went live in April 2012. The problem canonical URLs are addressing is duplicate content at different locations that can confuse search engines such as Google or Bing. This is related to the problem persistent identifiers such as DOIs are addressing for the scholarly community: accessing content over long periods of time that may change its location on the web (its URL), with two inter-related strategies: * **URL redirection**. DOIs redirect to a target URL that can be changed by the publisher, * **Persistence**. The publisher of scholarly content makes an extra effort to make sure content doesn't disappear (link rot), or significantly change (content drift). Obviously, canonical URLs are not DOIs, but they provide a standard way for a science blog to add a DOI to a post. ## Backends Science blogs provide a backend to store content and metadata, including the canonical URL. This can either be a database (as in the case of Wordpress or Ghost) or a file (as in the case of Hugo and many other static site generators). ### Wordpress Wordpress doesn't know about canonical URLs out of the box, but they can be added via a plugin, the most popular for this being Yoast SEO (which comes in free and paid versions). After installing and activating the plugin you can add a canonical URL in a new Yoast SEO section of the post editor: Alternatively, you can fiddle with your Wordpress configuration to add a custom field for the canonical URL. ### Ghost The Ghost blogging platform has a canonical URL field for every post, which you can access from the post settings sidebar: ### Hugo Hugo and other Open Source static site generators give you a lot of flexibility with metadata. If you add a `canonicalUrl` field to the blog post Front Matter, you can reuse it for the canonical URL (with some additional work). The canonical URL or DOI is now stored with the blog post, but also exposed to web crawlers. The format is `<link rel="canonical" href="``https://doi.org/10.53731/gvb08-7kc16``">`. ## Frontends To display the canonical URL aka DOI on your blog frontend, you have to modify your blog theme, the popular themes for Wordpress, Ghost, and Hugo don't really support displaying the canonical URL out of the box, as they are primarily intended for web crawlers and not humans. You should follow the Crossref DOI display guidelines, when thinking about how to display the DOI for your blog post, i.e. always be displayed as a clickable full URL link. Rogue Scholar displays DOIs like this: This blog (using the Ghost platform) displays DOIs like this in a sidebar: ## DOI registration workflow The changes to the backend and frontend explained above are good enough for occasional blog posts or to get started with Rogue Scholar. After a blog post is published, Rogue Scholar will register a DOI within 20 minutes and show that DOI on the website or via API. You can then copy/paste that DOI into your new canonical URL field. A simple improvement would be notifications of new DOI registrations by email, similar to what Crossref is sending to Front Matter as the Crossref member: <?xml version="1.0" encoding="UTF-8"?> <doi_batch_diagnostic status="completed" sp="ds5"> <submission_id>1590342900</submission_id> <batch_id>8a637b09-fda6-4980-baa1-147497683bd9</batch_id> <record_diagnostic status="Success"> <doi>10.53731/w6nzs-jta75</doi> <msg>Successfully added</msg> <citations_diagnostic> <citation key="ref1" status="resolved_reference">10.53731/gvb08-7kc16</citation> <citation key="ref2" status="resolved_reference">Cite to nonCR doi: 10.5281/zenodo.1324300</citation> <citation key="ref3" status="resolved_reference">10.1371/journal.pone.0115253</citation> <citation key="ref4" status="resolved_reference">10.59350/p000s-pth40</citation> <citation key="ref5" status="resolved_reference">10.53731/r79x921-97aq74v-ag5a2</citation> </citations_diagnostic> </record_diagnostic> <batch_data> <record_count>1</record_count> <success_count>1</success_count> <warning_count>0</warning_count> <failure_count>0</failure_count> </batch_data> </doi_batch_diagnostic> But maybe including a clickable link to the DOI just registered and some basic metadata that were registered (as it takes a few hours until the metadata show up in the Crossref REST API). For blogs with a more frequent publication frequency (e.g. weekly or daily) this workflow should be automated. One important consideration is whether the blog should know the DOI that will be registered in advance, avoiding the round trip with Rogue Scholar and Crossref, and allowing customizations of the DOI name, such as `10.53731/front-matter.2023-09-19`. The biggest advantage would be that the DOI name can be shared in advance of publication, e.g. for press releases, or to reference in other content. While these considerations are reasonable and not new for DOIs in general, for the science blog use case the workflow should be simple and I want to follow these principles: * Rogue Scholar DOIs will be generated as a short random 10-character string upon DOI registration. Rogue Scholar users or staff can't modify the DOI names that will be generated. Rogue Scholar DOIs are cool DOIs. * If you see a Rogue Scholar DOI, it can be used (immediately as a link, accessing the metadata after a few hours). Rogue Scholar is not offering DOIs that are not or not fully registered, i.e. DOIs for pending publications (Crossref) or draft DOIs (DataCite). * DOI registration happens with the Rogue Scholar service talking to the Crossref API, participating blogs don't need to install or develop functionality to generate Crossref metadata and/or interact with the Crossref API. While this workflow was a reasonable start, it was overly complicated and required an extra effort by the science blog. So in January 2025, Rogue Scholar started a new workflow: If the blog generated the DOI string containing the same random 10-character string, and added this string to the RSS feed, Rogue Scholar would use that string for DOI registration. Ten blogs are already participating in that workflow and the experience the past three months has been very positive. As always, the devil is in the details, and on one occasion the checksum of the provided DOI string was not valid. The limitation of this workflow is that it requires the blog to send the intended DOI string in the RSS feed. Which works nicely for static site generators, but for database-driven blogging platforms this may not possible. So this week Rogue Scholar is launching a new workflow. ### Generating DOI strings from the id/guid in the blog post feed Blogging platforms that are not static site generators but database-driven use a unique identifier for blog posts provided by the database. This can be long and complicated, as is the case for Blogger, Substack, or Ghost, but in the case of Wordpress the **post_id** is a simple number that increases with every post. And the feed contains this `id/guid` together with the hostname of the blog as URL, e.g. `https://svpow.com/?p=23496`. Every blog in Rogue Scholar has a unique identifier, which is used internally and to identify the blog communities, typically based on the domain name, so the **Sauropod Vertebra Picture of the Week** (svpow) blog can be found here. The combination makes a relatively short, globally unique identifier that can be used for the DOI string: `https://doi.org/10.59350/svpow.23496` Rogue Scholar added support for this DOI format for Wordpress blogs this week. This feature is currently in beta testing, please reach out if you want to be an early adopter. If there are no surprising issues, I expect this feature to roll out for all Rogue Scholar Wordpress blogs on May 15. And if your blog uses a static site generator (e.g. Hugo, Jekyll, or Quarto), you can also reach out if you want to pre-assign DOIs in the random format. They still made sense here, as static site generators don't automatically generate unique persistent IDs for posts (they generate permalinks, which depend on the configuration and may change over time). ## References Fenner, M. (2023, September 22). DOI registration workflow for a science blog. _Front Matter_. https://doi.org/10.53731/w6nzs-jta75 Fenner, M. (2023, September 19). Streamlining the archiving of science blog posts. _Front Matter_. https://doi.org/10.53731/gvb08-7kc16 Fenner, M. (2025, January 16). Persistent identifiers, random strings, and checksums. _Front Matter_. https://doi.org/10.53731/6kfyy-nq280

blog.front-matter.io

Working with the Research Organization Registry (ROR) Data Dump

The commonmeta Go library has seen a major update this week that dramatically simplifies working with the Research Organization Registry (ROR) data dump, including conversion to other serialization formats (e.g. JSON Lines) and metadata formats (InvenioRDM), and integrated affiliation matching. ### File Download ROR metadata are updated regularly (typically about once a month) and made available as a file download via the Zenodo repository under a Creative Commons Zero waiver. The single file is a compressed zip archive with the metadata in JSON and CSV formats, each for v1 and v2 of the ROR schema. There are two challenges with the ROR data dump file download: while there is a stable DOI for the latest version of the data, that DOI resolves to the dataset landing page, and there is no easy way to automatically get to the file download URL for automatic downloads of new versions. The other challenge is that the compressed zip archive contains four archived files that can't be downloaded individually. The complete archive is 58.6 MB, whereas the zipped v2 JSON would be 17.7 MB (the uncompressed file is 256.9 MB). In the commonmeta library, the download URL of the most recent ROR data dump and the file names in that zip archive are hard-coded. commonmeta can automatically fetch the full zip archive and selectively extract the v2 JSON. When using commands that require ROR metadata, commonmeta looks for a data dump in zipped Avro format (more on Avro below) in the folder where the command is run. If that file isn't found, commonmeta looks for a v2 JSON file from the data dump, and if that file isn't found either, fetches the data dump from Zenodo, extracts the v2 JSON file, and generates the compressed Avro file. For example a local lookup of an organization via its ROR identifier: commonmeta convert https://ror.org/04jvcky17 Running this command transparently downloads, extracts, and converts the latest ROR data dump and looks up the metadata for https://ror.org/04jvcky17 (Newport Festivals Foundation). This takes only a few seconds (depending on your network connection), and going forward uses ROR data stored locally. ### File Formats Commonmeta can automatically convert the JSON data of all 115K ROR records into other serialization formats, currently JSON Lines, YAML, CSV, and Avro. And optionally compress them as zip archive, for example: commonmeta list --from ror --file mydata.csv.zip This command in a few seconds generates a compressed CSV file (almost) identical to the CSV provided in the ROR data dump. Whereas JSON, JSON Lines, and YAML are straightforward to work with, the CSV format has limitations and shows only a (large) subset of the metadata. Avro is another special format; it is focused on efficient storage and transmission over network connections, but is not human-readable. Avro uses a schema in JSON format, which has advantages over the schema-less JSON, YAML, and CSV, including data validation and smaller file sizes. These are the file sizes for the ROR data dump (the JSON file is slightly smaller than the original file because null values were omitted). Format | Size (MB) | Size ZIP (MB) ---|---|--- JSON | 182.3 | 17.0 JSON Lines | 108.1 | 14.5 YAML | 125.5 | 15.3 CSV | 33.5 | 10.3 Avro | 41.6 | 13.2 Other criteria besides file size are the speed of reading and writing files in this format, readability by humans, and supported data types. CSV is very readable, but is only useful as output format, as data types other than text and numbers are not supported. Avro generates the smallest files, but is more complicated to work with as it requires a schema. YAML is very human-readable, whereas JSON Lines works well over network connections as it is easier to stream than JSON. commonmeta allows working with all these formats for ROR data (CSV only as output as it doesn''t include all metadata), so you can for example give a JSON Lines file to a colleague and she can use it as commonmeta input: commonmeta list mydata.jsonl --from ror --file mydata2.csv ### Metadata formats and filtering The ROR schema describes the metadata needed for the affiliation use case in scholarly works. A related metadata schema is used by the InvenioRDM repository platform that the Rogue Scholar blogging platform also uses. Here metadata vocabularies are typically described in YAML and use a subset of the ROR metadata, both for author affiliations and for funders. One small twist is that one metadata field is different: affiliations support acronyms for names, whereas funders support the country code. commonmeta can handle this and also generate a subset of organizations that are of type `funder` : commonmeta list --from ror --to inveniordm --file funders.yaml One challenge in the current version of the InvenioRDM platform (v12.0) is that importing large vocabularies (e.g. all ROR data) is slow and error-prone. In a previous blog post I suggested to only import the affiliations needed, but that workflow is slow and complicated. The upcoming version v13.0 of InvenioRDM has better handling of large vocabulary imports, in the meantime commonmeta supports the generation of smaller vocabularies that can be imported in batches, e.g. commonmeta list --from ror --to inveniordm --file affilations_ror.yaml -n 10000 --page 1 This command generates a YAML file in a format InvenioRDM understands and containing only 10,000 organizations. By installing commonmeta on the InvenioRDM server and running this command repeatedly and importing the YAML in batches we can overcome the limitations of the v12.0 vocabulary import. commonmeta currently supports two filters to generate subsets of the ROR data: by organization type (e.g. funder or university) and/or by country. Please reach out if you are interested in other metadata formats for organizations and/or filters. ### Queries commonmeta supports simple queries of the local ROR data by ROR ID or external ID (Crossref Funder ID, GRID, ISNI, Wikidata): commonmeta convert Q7713086 More complex queries are currently not possible with local data, but commonmeta integrates with the ROR API to support affiliation matching. commonmeta match --from ror "The Alfred Hospital" This will return a single ROR record if a match with a score of 0.9 or higher is found by the ROR API. The affiliation matching can be combined with looking up metadata from Crossref, DataCite, or InvenioRDM, a core functionality of commonmeta. If affiliation names but no ROR ids are provided, commonmeta can automatically merge the information found by affiliation matching. To indicate that the metadata was found by ROR and not the publisher, commonmeta adds an `assertedBy` field with the value `ror` to the response. Affiliation identifiers provided by the publisher will have a `publisher` value in the response. The following query returns a random sample of 50 publications from Crossref member 31795 (Front Matter) with affiliation matching applied and the results stored in commonmeta schema format. commonmeta list --from crossref --member 31795 --sample=true -n 50 --match=true --file matching.json ### Conclusions The work on using Go to read, format, and integrate ROR metadata was inspired by a session at the recent InvenioRDM Partner meeting in Hamburg. InvenioRDM is written in Python and Javascript/React, but Go is a great alternative for simple installations (commonmeta is a single 5 MB binary) and performance-critical functions (e.g. converting the ROR data dump into the InvenioRDM YAML format). Rust is another language that people use in these situations, and more work is needed to compare the relative strengths and weaknesses. More work is also needed to decide on the best file format for storing scholarly metadata at scale. Avro looks promising, but needs to be compared with JSON in more detail. Another interesting newer format is Parquet. It is a column-oriented file format in contrast to the formats described here, which are row-based. This makes some things harder but other things much easier, and this becomes more critical as the number of metadata records grows from 100K (ROR) to the millions (DataCite, Crossref). Finally, this work demonstrates that a good proportion of metadata work can be done locally, working with data dumps rather than high frequencies of API calls. ## References Research Organization Registry. (2025). _ROR Data_ (Version v1.63) Dataset]. Zenodo. [https://doi.org/10.5281/ZENODO.6347574 Martin Fenner. (2025). _front-matter/commonmeta: V0.19.4_ (Version v0.19.4) Computer software]. Zenodo. [https://doi.org/10.5281/ZENODO.15256488 Fenner, M. (2025, April 7). Where I simplified ROR affiliation metadata handling. _Front Matter_. https://doi.org/10.53731/ymbv8-7jm78 Fenner, M. (2025, March 19). Rogue Scholar meets the InvenioRDM community. _Front Matter_. https://doi.org/10.53731/1aw0b-pr243

blog.front-matter.io

Rogue Scholar Newsletter March 2025

This is the third issue of the monthly newsletter from the Rogue Scholar science blog archive. The newsletter reports on new blogs that have joined the platform, important technical updates in Rogue Scholar infrastructure, community updates, and other news relevant to Rogue Scholar users. ## Blogs added to Rogue Scholar Four blogs from four different subject areas were added in February. Welcome everybody! ### John Arundel (Bitfield Consulting) _Computer and information sciences, English._ https://bitfieldconsulting.com/posts/ ### Roger Beecham's blog Social and economic geography _, English._ https://www.roger-beecham.com ### Análise Quantitativa das Mudanças Sociais _Social science, Portuguese._ https://aqms.substack.com/ ### Blogposts on autosys _Computer and information sciences, English._ https://autosys.informatik.haw-hamburg.de/blog/ ## Technical Updates In March, I continued work on the statistics page, which was renamed to the Rogue Scholar dashboard. The data for this page comes from Rogue Scholar search facets, which have been greatly expanded, including filtering by publication year: Work has started to show the full-text content (stored in the database since Rogue Scholar launched and available via API) on record landing pages. This helps with archiving and searching the full-text content. The feature is currently undergoing extensive testing and will launch on April 14. Users with **manager** permissions for communities (blog, subject area, or topic) can see the full-text already. If you do, please provide feedback. ### Community Update InvenioRDM is the repository platform that powers Rogue Scholar as well as more than 20 other repositories, including Zenodo. Last week, about 40 people met in Hamburg for the annual partner meeting to discuss ongoing development, new features, and the timeline for the release of version v13 of the platform. On the first day, we had short presentations from about ten InvenioRDM instances in production, highlighting unique functionalities. I shared my slides two weeks ago. Please use Slack, email, Mastodon, or Bluesky if you have any questions or comments regarding this monthly newsletter. Rogue Scholar is a scholarly infrastructure that is free for all authors and readers. You can support Rogue Scholar with a one-time or recurring donation or by becoming a sponsor. ## References Fenner, M. (2025, March 10). Working on the Rogue Scholar dashboard. _Front Matter_. https://doi.org/10.53731/wtvvs-f4h04 Fenner, M. (2025, March 19). Rogue Scholar meets the InvenioRDM community. _Front Matter_. https://doi.org/10.53731/1aw0b-pr243 Fenner, M. (2025). _Rogue Scholar InvenioRDM Workshop 2025_. https://doi.org/10.5281/ZENODO.15050863

blog.front-matter.io