Pınar Tözün

@pinartozun.bsky.social

Head of Data, Systems, and Robotics Section & Associate Professor at IT University of Copenhagen https://www.pinartozun.com/ https://itu-dasyalab.github.io/RAD/ http://distortedpollyanna.blogspot.com/

[not work] Since the end of March, I have been in some sort of camp, first to find a new place to live and then to prepare for the Danish citizenship test. It all got sorted out in the end, even though it was a rollercoaster. We start with the latter. distortedpollyanna.blogspot.com/2026/06/indf...

Indfødsretsprøven

Study material for indfødsretsprøven. There I was last Wednesday in a sports hall in Frederiksberg with many other foreigners, sitting on a ...

distortedpollyanna.blogspot.com

How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI Sophia N. Wilson, Sebastian Mair, Mophat Okinyi, Erik B. Dam, Janin Koch, Raghavendra Selvan https://arxiv.org/abs/2602.00056 #AI #thereIsNoAI #thereIsInParticularNoSustainableAI

How Hyper-Datafication Impacts the Sustainability Costs in Frontier AI

Large-scale data has fuelled the success of frontier artificial intelligence (AI) models over the past decade. This expansion has relied on sustained efforts by large technology corporations to aggregate and curate internet-scale datasets. In this work, we examine the environmental, social, and economic costs of large-scale data in AI through a sustainability lens. We argue that the field is shifting from building models from data to actively creating data for building models. We characterise this transition as hyper-datafication, which marks a critical juncture for the future of frontier AI and its societal impacts. To quantify and contextualise data-related costs, we analyse approximately 550,000 datasets from the Hugging Face Hub, focusing on dataset growth, storage-related energy consumption and carbon footprint, and societal representation using language data. We complement this analysis with qualitative responses from data workers in Kenya to examine the labour involved, including direct employment by big tech corporations and exposure to graphic content. We further draw on external data sources to substantiate our findings by illustrating the global disparity in data centre infrastructure. Our analyses reveal that hyper-datafication does not merely increase resource consumption but systematically redistributes environmental burdens, labour risks, and representational harms toward the Global South, precarious data workers, and under-represented cultures. Thus, we propose Data PROOFS recommendations spanning provenance, resource awareness, ownership, openness, frugality, and standards to mitigate these costs. Our work aims to make visible the often-overlooked costs of data that underpin frontier AI and to stimulate broader debate within the research community and beyond.

arxiv.org

3, 2, 1... LIFT OFF!! 🚀 ✨🛰️✨ For få minutter siden blev den danske studentersatellit 𝗗𝗜𝗦𝗖𝗢-𝟮 sendt afsted ud i rummet! Satellitten er udviklet af studerende fra AU, SDU og ITU med støtte fra bl.a. Industriens Fond.

Gør AI udviklere overflødige? Ikke ifølge ITU-professor Peter Sestoft og uddannelseschef Luís Cruz-Filipe. "Der er stadigvæk brug for ekstremt dygtige softwareudviklere, som kan styre de her AI‑agenter og bruge dem til at lave ordentligt software, som faktisk kan bruges ansvarligt." #dktech #itpol

Dansk it-professor: Du misforstår noget, hvis du tror, at AI vil gøre behovet for it-folk mindre

Interview: Vibe coding og jagten på hurtige AI-kompetencer har skabt både begejstring og bekymring for programmørens og uddannelsernes snarlige død. Men det er en misforståelse, siger to de fremtræden...

computerworld.dk

This week our partners in DEEP project (funded by @snsf.ch) hosted us in Switzerland. We had talks and lively discussions at UNIL, EPFL, HEIG-VD. Thanks to Pamela Delgado, Busra Karatay Demiray, & Hector Satizabal for hosting us. A brief article on the DEEP project: en.itu.dk/About-ITU/Pr...

Research project to boost sustainable AI

The environmental impact of AI is significant and growing. The researchers behind the new research project, DEEP, have received a 5,6 million DKK grant for a project focused on reducing the cost and t...

en.itu.dk

MSc thesis work of Emil Houlborg and Andreas Tietgen @itu.dk will be presented at CIDR conference next week. Huge congrats to them! Modern storage I/O landscape is very diverse, yet not so accessible. We created a new file system extension for @duckdb.org with xNVMe to help with this accessibility.

Bild

Azim had a stellar committee which included Daniel Lemire (TELUQ 🇨🇦), Pinar Tözün (ITU 🇩🇰) and Viktor Leis (TUM 🇩🇪). They gave talks at the Dutch Seminar on Data Systems Design (DSDSD) on SIMD-accelerated parsing, xNVMe and escaping from the insanity of SQL. (video will be posted on dsdsd.da.cwi.nl)

Bild

Looking forward to the South Bay Systems event after CIDR on Jan 21! I will be presenting our CIDR'26 paper on integrating xNVMe into DuckDB --> vldb.org/cidrdb/paper... Thanks for the invitation @southbaysystems.xyz @alexmillerdb.bsky.social !!

vldb.org

South Bay Systems@southbaysystems.xyz · 8mo ago

Our next event will be on January 21st, featuring speakers from (the just-finishing) CIDR! Come to Databricks to hear about: * DuckDB on xNVMe by @pinartozun.bsky.social of ITU * Spilling in QP by Maximilian Kuschewski of TUM * NPUs in DBs by Alexander Baumstark of TU-Ilmenau luma.com/8a54z94d