a neat thing about hubble is that you can get a point-in-time* snapshot of the whole network. vs doing your own backfill crawl, which will give you samples spread over a >24h period any researchers who want this, feel free to reach out
24,347,088,288 records in 41,172,356 repos in the reachable network fits in 1.86TiB in hubble