Tech

Dataset reveals how Reddit communities are adapting to AI

Share
Share
Dataset reveals how Reddit communities are adapting to AI
Number of subreddits in the Longitudinal Subreddit Set with AI rules. Subreddits are bucketed into deciles based on their subscriber count at the time of the second crawl. Labels above the bars indicate percent change between crawls. Credit: arXiv (2024). DOI: 10.48550/arxiv.2410.11698

Researchers at Cornell Tech have released a dataset extracted from more than 300,000 public Reddit communities, and a report detailing how Reddit communities are changing their policies to address a surge in AI-generated content.

The team collected metadata and community rules from the online communities, known as subreddits, during two periods in July 2023 and November 2024. The researchers will present a paper with their findings at the Association of Computing Machinery’s CHI conference on Human Factors in Computing Systems being held April 26 to May 1 in Yokohama, Japan.

One of the researchers’ most striking discoveries is the rapid increase in subreddits with rules governing AI use. According to the research, the number of subreddits with AI rules more than doubled in 16 months, from July 2023 to November 2024.

“This is important because it demonstrates that AI concern is spreading in these communities. It raises the question of whether or not the communities have the tools they need to effectively and equitably enforce these policies,” said Travis Lloyd, a doctoral student at Cornell Tech and one of the researchers who initiated the project in 2023.

The study found that AI rules are most common in subreddits focused on art and celebrity topics. These communities often share visual content, and their rules frequently address concerns about the quality and authenticity of AI-generated images, audio and video. Larger subreddits were also significantly more likely to have these rules, reflecting growing concerns about AI among communities with larger user bases.

“This paper uses community rules to provide a first view of how our online communities are contending with the potential widespread disruption that is brought by generative AI,” said co-author Mor Naaman, professor at the Jacobs Technion-Cornell Institute at Cornell Tech, and of information science in the Cornell Ann S. Bowers College of Computing and Information Science. “Looking at actions of moderators and rule changes gave us a unique way to reflect on how different subreddits are impacted and are resisting, or not, the use of AI in their communities.”

As generative AI evolves, the researchers urge platform designers to prioritize the community concerns about quality and authenticity exposed in the data. The study also highlights the importance of “context-sensitive” platform design choices, which consider how different types of communities take varied approaches to regulating AI use.

For example, the research suggests that larger communities may be more inclined to use formal, explicit rules to maintain content quality and govern AI use. In contrast, closer-knit, more personal communities may rely on informal methods, such as social norms and expectations.

“The most successful platforms will be those that empower communities to develop and enforce their own context-sensitive norms about AI use. The most important thing is that platforms do not take a top-down approach that forces a single AI policy on all communities,” Lloyd said. “Communities need to be able to choose for themselves whether they want to allow the new technology, and platform designers should explore new moderation tools that can help communities detect the use of AI.”

By making their dataset public, the researchers aim to enable future studies that can further explore online community self-governance and the impact of AI on online interactions.

The findings are published on the arXiv preprint server.

More information:
Travis Lloyd et al, AI Rules? Characterizing Reddit Community Policies Towards AI-Generated Content, arXiv (2024). DOI: 10.48550/arxiv.2410.11698

GitHub: github.com/sTechLab/AIRules

Journal information:
arXiv


Provided by
Cornell University


Citation:
Dataset reveals how Reddit communities are adapting to AI (2025, April 7)
retrieved 7 April 2025
from

This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no
part may be reproduced without the written permission. The content is provided for information purposes only.

Share

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Articles
Midjourney V7 gives the AI image-maker power, speed, and correctly shaped hands
Tech

Midjourney V7 gives the AI image-maker power, speed, and correctly shaped hands

Midjourney has released a new AI model for producing images, Midjourney V7...

Details on material composition now available for Germany’s entire building stock could promote sustainability
Tech

Details on material composition now available for Germany’s entire building stock could promote sustainability

by Heike Hensel, Leibniz-Institut für ökologische Raumentwicklung e. V. View of an...

Is AI truly creative? Study shows how visibility of process shapes perception
Tech

Is AI truly creative? Study shows how visibility of process shapes perception

In the study, participants were initially asked to evaluate the creativity of...

Encryption method for key exchange enables tap-proof communication to fend off future quantum tech threats
Tech

Encryption method for key exchange enables tap-proof communication to fend off future quantum tech threats

Credit: Pixabay/CC0 Public Domain Quantum computers are a specter for future data...