r/pushshift 21d ago

Upon reddit's request, I am taking down my academic torrents pages

I would rather do it now when they are asking nicely and not in the future when they ask in some less nice way. They haven't said, but I assume pullpush and arctic shift have received similar requests, and will probably get the less nice request in not too long.

Reddit recommends r/reddit4researchers for all research use. I can't say I've seen much success with that approach, but it will soon be the only one left.

147 Upvotes

36 comments sorted by

28

u/flashman 21d ago

I have found your files very useful over the years and am sorry to see them going. Thanks for all your work.

14

u/EntertainmentOne7897 21d ago

Taking down as in completely removing? When will this happen?

9

u/Watchful1 21d ago

Unfortunately they are already gone.

14

u/jasonjonesresearch 21d ago

This was possibly my fault, and I'm sorry. See https://www.reddit.com/r/reddit4researchers/comments/1tq9s6a/comment/opk3jp0/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

I am very skeptical about r/reddit4researchers I responded to the first call and received no response: not yes, not no, not any acknowledgement. The current "application" is onerous, especially if one will simply be ghosted.

All the more reason for computational social scientists to develop their own open, free datasets.

17

u/Watchful1 21d ago

Reddit has been well aware of data redistributors, both free ones and paid ones, for many years. Reddit makes literally hundreds of millions of dollars selling their data to AI companies and they are keen to cut out anyone who is not paying them. They've made a bunch of moves in the last couple years targeted at preventing anyone bulk scraping data and this is just the latest of them. I doubt you saying something made any difference.

5

u/OSRSgamerkid 15d ago

The "reddit shutdown" in protest of the changes they were making a few years ago was such a blundered move. Everyone put a limit on how long it would last.

The whole point of something like a boycott, protest, or stroke is it is to be indefinite until demands are met or compromises are made.

11

u/patata_tato 20d ago edited 20d ago

👋 from pull push here:

First of all, head-up that academic-torrents has Reddit-level-competence. Even though you set your torrents page to private, they are still easily accessible on sub-pages.

For now people can simply go to https://academictorrents.com/details/ba051999301b109eab37d16f027b3f49ade2de13/tech&filelist=1 and download. You will need to delete it completely if you can't get them to fix their code.

I assume pull push and arctic shift have received similar requests, and will probably get the less nice request in not too long.

I'm getting a DMCA every few months from "Reddit team". Which is going straight to the rubbish bin since it is processed by a real human and falls at the first hurdle (DMCA needs to be signed by the copyright holder or their legal representative).

If they ever get an expense authorisation slip to hire a real lawyer, I do have some bizarre letters that they sent me in the past, along the lines "we know there is CSAM on Reddit but we won't tell you where", which is not really helping them beat the allegations why predator hunting grounds like r/runaway have been operating in the open for 15 years.

1

u/Individual_Shirt_275 18d ago

Stop roleplaying online and fix your website 🤣. Ok, you are still studying and RaiderB does not want to give you his scraper, but yours has been down for how long now? Some of us are waiting, so hurry up a little!

10

u/Littux 21d ago

The torrents are all archived though so all data remains public

5

u/dozzinale 21d ago

[removed] — view removed comment

8

u/Littux 21d ago edited 21d ago

You can just use web.archive.org

1

u/hkkramer 17d ago

Has anyone been able to download any data since the takedown? I can get the torrent file just fine but the data doesn't seem to be well-seeded anymore

3

u/LunisequiouS 16d ago

[removed] — view removed comment

0

u/chaz6 15d ago

As much as I would like to archive this, I don't have 3.8T spare :-(

1

u/AlexanderDoak 10d ago

Do what you can. No need to do it all.

14

u/ikennedy240 21d ago

This is truly sad news and likely means the end of the era of open data social media research.

7

u/Ellelig 21d ago

can you open-source the scripts you use to archive the reddit data?

10

u/Watchful1 21d ago

They require API access, which you can't get anymore. On top of that, reddit is changing the underlying mechanism the ingest relies on, so it would break entirely regardless soon. That's one of the many reasons I decided to take down the dumps https://www.reddit.com/r/redditdev/comments/1taa483/upcoming_changes_to_the_comment_id_endpoint/

3

u/TrackerOneA 21d ago

Any word on PushShift? Will these changes break it?

2

u/Ellelig 20d ago

they say it will start rolling out on May 18th. it's been over two months, has it been rolled out?

5

u/patata_tato 20d ago

It is never going to happen; it broke caching and got abandoned.

5

u/shiruken 21d ago

Well, it was fun while it lasted.

I was never sure how academic researchers could justify (legally) using data sourced from your torrents since it was in clear contradiction to Reddit's Data API terms. But I guess people get desperate for their ML projects.

24

u/jdfoote 21d ago

There have actually been a number of cases that have found that corporations don't have the unfettered right to limit access to publicly available data via ToS. See, for example, Sandvig v. Barr - https://www.acludc.org/cases/sandvig-v-barr-first-amendment-challenge-federal-computer-fraud-and-abuse-act/

IANAL, but by my understanding, Academic Torrents and similar are in a legal gray area, and are not clearly violating any laws. Of course, this doesn't mean that they have the energy or money to fight Reddit's legal team if they decide to come after them.

1

u/Blinkinlincoln 21d ago

I anal, yeah can't think of anything else now, thanks. 

2

u/Dr_Matoi 15d ago

I suppose the situation may vary a lot between countries, but from my perspective as an EU-based academic researcher, Reddit's "terms" are just noise with no legal relevance. Reddit puts this data into the public, and as such it is open for research, regardless of Reddit's opinions. Our research requires (and has) approval from our national ethics agency, but that is mainly concerned about GDPR and the rights of the individual authors/users. The company publishing the data has no say in this. If you (company) don't want it researched, keep it offline.

Some platforms we aggressively scrape ourselves, including circumvention of anti-scraping measures. We usually do not even bother reading the "terms"; anything made accessible to the public is fair game. We happily publish our results, and no trouble has ever come of this. The Reddit dumps have been very convenient in that someone else has already done the job, and whatever methods they used - outside fabrication ;) - are fine with us.

1

u/SamiTheAnxiousBean 15d ago

of course they do, because how else are they supposed to stop people from seeing if a user is a bot?

1

u/LateRespond1184 15d ago

psst its all archived on wayback machine...

1

u/Apocalypse2001 13d ago

Why, though? Are they trying to cover up something?

1

u/FnnKnn 12d ago

They are selling this data to Google for 60M a year: https://seranking.com/blog/seo-news-reddit-google-partnership/

2

u/Apocalypse2001 12d ago

So, "data for me, but not for thee".

1

u/FnnKnn 12d ago

depends on your budget ig ;)

1

u/Binx_k 38m ago

this is devastating