This contrive aims to make the offset of a mapping of post-fall crosswise Reddit. I am looking at to take apart how contentedness flows between subreddits, by inquisitory for content that has been cross-posted (as referenced in the title). Reddit was founded punt in October 2006 by Alexis Ohanian, Steve Huffman and Aaron Swartz. It is a website which congregated many online communities of varied tastes together in ace locate. Users are able-bodied to state texts, pic and contact posts, which are voted up or belt down by other users. The primary ferment that divine this protrude was created by anvaka on github.
I plan to add together functionality to the codification so that the turn of multiplication a subreddit is cross-posted to tooshie be counted. In the almost future, I search to promote engage mechanisation of ECL code, so that totally 2,500 subreddits tooshie be analyzed. Once this is done, I design to make a visualization correspondence (similar to the represent referenced earlier, created by Anvaka) of subreddit depicted object sharing. Represent naturally-occurring inter-subreddit depicted object communion patterns on Reddit by analyzing how posts are “cross-posted” between subreddits based on 2.5 million posts across the top 2,500 subreddits. Uses ECL and HPCC Systems. My original plan was to have each subreddit’s [subredditname]_master_results.thor file function as a child of a master class, which could then be denormalized so that each sub’s cross-posted subs could be seen in a single master file. My code was written so that when an event (a file arrive in the Landing Zone with type .CSV) was triggered, the scheduler could automatically run all subsequent files and generate the child class.
They created a visualization of related subreddits, using the “Redditors who commented in this subreddit, also commented to…” suggestions that generate in subreddit side bars. Flair in Reddit is a type of ‘tag’ which can be used to sort posts and users also can have it beside their username to broadcast a message or show an emoji. For example in sports-related subreddits, users can have their club logos beside their username to show the club they support. Both have a common thread, which is that before any analysis can be done, manual code for cleaning has to be written. When ECL Watch sprays a .CSV file, it doesn’t actually create record sets, split rows into columns, or do anything beyond outputting rows in the most basic single column format. The primary work that aided in the coding of this project was created by HPCC Systems on github. They created a training ECL program to analyze and predict New York City Traffic Data, using data pulled from CSV files. The top page of the ECL Language Reference can be found at hpccsystems.com/training/documentation/ecl-language-reference/html. While ECL proves to be a highly efficient language for data mining, there are not many existing resources for it outside of HPCC System’s purview.
On my current HPCC account, though, I could not get access to the ECL Scheduler, and the number of CSV files to go through was to great to do it manually. I explored using FUNCTIONMACRO, MACRO, and embedding both Python and JavaScript, but by the end of the project I was still unable to successfully find a method to do this. Before beginning my analysis, orgy porn videos I had to obtain my datasets and get them ready to upload to the HPCC Landing Zone. Here is a complete list of my datasets and bundles used, as well as relevant information for each. It is important to note that datasets relied heavily on the fact that Reddit generates unique base36 IDs for every comment, user, post, and subreddit. Subreddit IDs are formatted as t5_[ID], and posts are formatting as t3_[ID]. Finally, due to the lack of a master file, I was unable to create the visual map I had first set out to make. An interesting way to use this project is to generate cross-posting subs and then compare them to anvaka’s program to see how users move between subreddits based on both cross-posting and comments.
My results were of a reasonable quality, and certainly provide interesting analysis, but limitations in ECL, programmer knowledge, and available resources lead to a few issues. HPCC has a slew of administrators and experts that are constantly on their help forum, and most issues that couldn’t be resolved via the ECL Language Reference could be found here. In particular, the data importing and cleaning portion of this project was built while referencing this code. This sub had significantly longer titles than many other subs, so title values are cropped.
Content found in one subreddit is often considered to be relevant to another subreddit by users. However, the social rules of Reddit discourage you from taking someone else’s content and reposting it somewhere else on the massive set of forums (creating a “repost” in the eyes of the community). As a way to share content without bringing out the mob that is often generated when a stolen repost is discovered, users have begun using the phrase “cross-posted”. This phrase can take a variety of structures, but always contains some form of cross-posted and a subreddit name and is almost always found at the end of a post title. Figure 1 shows a few real examples parsed from the dataset this project used. The goal of this project is to map naturally-occurring inter-subreddit content sharing patterns on Reddit.com by analyzing how posts are shared (or “cross-posted”) between subreddits. When posts are shared in another subreddit, community habit is to give credit to the original poster with the phrase “cross-posted from /r/subredditname” or something following this pattern posted in the comment section of the post. Looking at the patterns of cross-posting, I want to determine the relationships between subreddits as determined by the content shared between them. I will then compare this to an existing relational map of subreddits on Reddit to determine if content sharing predominantly occurs along relational aligns or outside of them.
This code was run on /r/aww, /r/DIY, /r/AskComputerScience, /r/dataisbeautiful, and /r/catpictures. Although (as outlined in the above) there were additional outputs generated along the steps, these are the ones that are the most useful from a data analysis view point. This data really shows how the type of subreddit affects content, and it begins to show how the subreddits relate to each other as they share content.
As such, most of the resources used to help code this project are found on their website or their github. As a supplementary reference to this, Guru Bandasha’s accompanying blogpost on the DataSeers website was used to understand how the code fit together. Here, you can see just a small sample of the ways that the cross-posting phrase can vary. This was one of the main challenges of the data analysis portion of this project. In addition to the above, there are another three pages of results like this. Out of 1000 posts, more than a tenth are cross posts (an extremely high rate, especially compared to the other subs looked at). I have omitted the following pages of post results in the interest of saving space.
