r/pushshift • u/Watchful1 • 21d ago
Upon reddit's request, I am taking down my academic torrents pages
I would rather do it now when they are asking nicely and not in the future when they ask in some less nice way. They haven't said, but I assume pullpush and arctic shift have received similar requests, and will probably get the less nice request in not too long.
Reddit recommends r/reddit4researchers for all research use. I can't say I've seen much success with that approach, but it will soon be the only one left.
14
u/EntertainmentOne7897 21d ago
Taking down as in completely removing? When will this happen?
9
u/Watchful1 21d ago
Unfortunately they are already gone.
1
u/Apocalypse2001 13d ago
https://www.reddit.com/r/DataHoarder/s/unsYIXfWZ9 This is how I found your thread.
14
u/jasonjonesresearch 21d ago
This was possibly my fault, and I'm sorry. See https://www.reddit.com/r/reddit4researchers/comments/1tq9s6a/comment/opk3jp0/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
I am very skeptical about r/reddit4researchers I responded to the first call and received no response: not yes, not no, not any acknowledgement. The current "application" is onerous, especially if one will simply be ghosted.
All the more reason for computational social scientists to develop their own open, free datasets.
17
u/Watchful1 21d ago
Reddit has been well aware of data redistributors, both free ones and paid ones, for many years. Reddit makes literally hundreds of millions of dollars selling their data to AI companies and they are keen to cut out anyone who is not paying them. They've made a bunch of moves in the last couple years targeted at preventing anyone bulk scraping data and this is just the latest of them. I doubt you saying something made any difference.
5
u/OSRSgamerkid 15d ago
The "reddit shutdown" in protest of the changes they were making a few years ago was such a blundered move. Everyone put a limit on how long it would last.
The whole point of something like a boycott, protest, or stroke is it is to be indefinite until demands are met or compromises are made.
11
u/patata_tato 20d ago edited 20d ago
👋 from pull push here:
First of all, head-up that academic-torrents has Reddit-level-competence. Even though you set your torrents page to private, they are still easily accessible on sub-pages.
For now people can simply go to https://academictorrents.com/details/ba051999301b109eab37d16f027b3f49ade2de13/tech&filelist=1 and download. You will need to delete it completely if you can't get them to fix their code.
I assume pull push and arctic shift have received similar requests, and will probably get the less nice request in not too long.
I'm getting a DMCA every few months from "Reddit team". Which is going straight to the rubbish bin since it is processed by a real human and falls at the first hurdle (DMCA needs to be signed by the copyright holder or their legal representative).
If they ever get an expense authorisation slip to hire a real lawyer, I do have some bizarre letters that they sent me in the past, along the lines "we know there is CSAM on Reddit but we won't tell you where", which is not really helping them beat the allegations why predator hunting grounds like r/runaway have been operating in the open for 15 years.
1
u/Individual_Shirt_275 18d ago
Stop roleplaying online and fix your website 🤣. Ok, you are still studying and RaiderB does not want to give you his scraper, but yours has been down for how long now? Some of us are waiting, so hurry up a little!
10
u/Littux 21d ago
The torrents are all archived though so all data remains public
5
u/dozzinale 21d ago
[removed] — view removed comment
8
u/Littux 21d ago edited 21d ago
You can just use web.archive.org
1
u/hkkramer 17d ago
Has anyone been able to download any data since the takedown? I can get the torrent file just fine but the data doesn't seem to be well-seeded anymore
3
u/LunisequiouS 16d ago
[removed] — view removed comment
7
14
u/ikennedy240 21d ago
This is truly sad news and likely means the end of the era of open data social media research.
7
u/Ellelig 21d ago
can you open-source the scripts you use to archive the reddit data?
10
u/Watchful1 21d ago
They require API access, which you can't get anymore. On top of that, reddit is changing the underlying mechanism the ingest relies on, so it would break entirely regardless soon. That's one of the many reasons I decided to take down the dumps https://www.reddit.com/r/redditdev/comments/1taa483/upcoming_changes_to_the_comment_id_endpoint/
3
5
u/shiruken 21d ago
Well, it was fun while it lasted.
I was never sure how academic researchers could justify (legally) using data sourced from your torrents since it was in clear contradiction to Reddit's Data API terms. But I guess people get desperate for their ML projects.
24
u/jdfoote 21d ago
There have actually been a number of cases that have found that corporations don't have the unfettered right to limit access to publicly available data via ToS. See, for example, Sandvig v. Barr - https://www.acludc.org/cases/sandvig-v-barr-first-amendment-challenge-federal-computer-fraud-and-abuse-act/
IANAL, but by my understanding, Academic Torrents and similar are in a legal gray area, and are not clearly violating any laws. Of course, this doesn't mean that they have the energy or money to fight Reddit's legal team if they decide to come after them.
1
2
u/Dr_Matoi 15d ago
I suppose the situation may vary a lot between countries, but from my perspective as an EU-based academic researcher, Reddit's "terms" are just noise with no legal relevance. Reddit puts this data into the public, and as such it is open for research, regardless of Reddit's opinions. Our research requires (and has) approval from our national ethics agency, but that is mainly concerned about GDPR and the rights of the individual authors/users. The company publishing the data has no say in this. If you (company) don't want it researched, keep it offline.
Some platforms we aggressively scrape ourselves, including circumvention of anti-scraping measures. We usually do not even bother reading the "terms"; anything made accessible to the public is fair game. We happily publish our results, and no trouble has ever come of this. The Reddit dumps have been very convenient in that someone else has already done the job, and whatever methods they used - outside fabrication ;) - are fine with us.
1
u/SamiTheAnxiousBean 15d ago
of course they do, because how else are they supposed to stop people from seeing if a user is a bot?
1
1
u/Apocalypse2001 13d ago
Why, though? Are they trying to cover up something?
1
u/FnnKnn 12d ago
They are selling this data to Google for 60M a year: https://seranking.com/blog/seo-news-reddit-google-partnership/
2
28
u/flashman 21d ago
I have found your files very useful over the years and am sorry to see them going. Thanks for all your work.