Open Context

@opencontext.hcommons.social.ap.brid.gy

Open Context (https://opencontext.org) reviews, edits, annotates, publishes and archives research data and digital documentation for archaeology and […] [bridged from https://hcommons.social/@opencontext on the fediverse by https://fed.brid.gy/ ]

RE: https://scholar.social/@ekansa/116811273528024594 This essay discusses some of the costs of #LLM #AI driven mass web-scraping and how it impacts #nonprofit #opendata services like Open Context. #datacuration #rdm #datamanagement

Eric Kansa@ekansa.scholar.social.ap.brid.gy · last mo.

Yay! My essay on the impacts of Large Language Model (#LLM) #AI in #archaeology was just published: https://doi.org/10.11141/ia.71.15 It looks at #bots and mass scraping on the infrastructure supporting #opendata and #openaccess. It also looks at the incentives that encourage the […]

Just making sure we have updated a full copy of the data we publish. Internationally. Just in case. https://doi.org/10.5281/zenodo.14728229 #OpenData #archaeology

Open Context Database SQL Dump

Open Context (https://opencontext.org) publishes free and open access research data for archaeology and related disciplines. An open source (but bespoke) Django (Python) application supports these data publishing services. The software repository is here: https://github.com/ekansa/open-context-py The Open Context team runs ETL (extract, transform, load) workflows to import data contributed by researchers from various source relational databases and spreadsheets. Open Context uses PostgreSQL (https://www.postgresql.org) relational database to manage these imported data in a graph style schema. The Open Context Python application interacts with the PostgreSQL database via the Django Object-Relational-Model (ORM). This database dump includes all published structured data organized used by Open Context (table names that start with 'oc_all_'). The binary media files referenced by these structured data records are stored elsewhere. Binary media files for some projects, still in preparation, are not yet archived with long term digital repositories. These data comprehensively reflect the structured data currently published and publicly available on Open Context. Other data (such as user and group information) used to run the Website are not included.    IMPORTANT This database dump contains data from roughly 190+ different projects. Each project dataset has its own metadata and citation expectations. If you use these data, you must cite each data contributor appropriately, not just this Zenodo archived database dump.

zenodo.org