Edit Content
Click on the Edit Content button to edit/add the content.

Reddit accuses Perplexity of data scraping

Reddit accuses Perplexity of data scraping

Reddit has accused artificial intelligence firm Perplexity of data scraping. According to reports, the platform filed a lawsuit against Perplexity for continuing to use its content to train its AI model despite prior warnings not to scrape the platform’s content.

As AI systems increasingly rely on publicly available online content to train and generate answers, companies like Reddit are trying to draw firm lines over what is considered “public” and “proprietary” data. Reddit has now filed a lawsuit against Perplexity, accusing it of illegally collecting data through its platform. According to court documents filed Wednesday in a Manhattan federal court, Reddit said Perplexity ignored instructions not to scrape its content and continued to use Reddit data to generate AI answers.

Reddit exposes alleged data theft

The complaint says Reddit had explicitly blocked Perplexity from collecting its data, but the AI company’s “answer engine” still produced results containing Reddit content. “The increase was so dramatic that an outside observer hypothesized that the increase was due to Perplexity entering a licensing deal with Reddit,” the lawsuit said. “In truth, there is no license between Perplexity and Reddit.”

To prove its suspicion, Reddit designed a clever digital test. It created a “trap” post that could only be found by Google’s search engine. Google has a legitimate content-licensing deal with Reddit, and so any company without such a deal should have been unable to access the post. The company described it as the online equivalent of a “marked bill,” noting that if Perplexity’s system reproduced the contents of that hidden post, Reddit would know it had gone around its safeguards, possibly by pulling data through Google’s search results, known as SERPs.

Within hours, the supposedly private test post began showing up in responses generated by Perplexity’s AI tool. “The only way that Perplexity could have obtained that Reddit content and then used it in its ‘answer engine’ is if it and/or its co-defendants scraped Google SERPs,” the lawsuit stated. Reddit named three data-scraping companies in the suit: Oxylabs UAB, AWM Proxy, and SerpApi. It accused them of helping Perplexity gain unauthorized access to Reddit’s posts, or of selling Reddit’s data to Perplexity.

Meanwhile, Perplexity has rejected Reddit’s allegations. According to Jesse Dwyer, a spokesperson for Perplexity, the company “will not tolerate threats against openness and the public interest.” The company also said in a Reddit post after the lawsuit was filed that it “does not train AI models on content.” Representatives of the other companies named in the lawsuit also issued statements.

A spokesperson for SerpApi said it plans to “vigorously defend” itself in court. Oxylabs’ chief governance and strategy officer, Denas Grybauskas, said his company was “shocked and disappointed,” adding that Oxylabs “has always been and will continue to be a pioneer and an industry leader in public data collection.”

In August, internet infrastructure company Cloudflare revealed it had conducted a similar test to see if Perplexity was following web-crawling rules. Cloudflare said it created pages marked with code telling Perplexity’s bots not to access them, but it still found the AI company’s crawlers visiting the restricted pages. Cloudflare’s CEO, Matthew Prince, made headlines by comparing Perplexity’s behavior to that of “North Korean hackers.” “Some supposedly ‘reputable’ AI companies act more like North Korean hackers,” Prince wrote on X.

Share:

More Posts

Public Companies Are About To Surpass Satoshi’s Bitcoin Holdings

Public Companies Are About To Surpass Satoshi’s Bitcoin Holdings

Bitcoin held by publicly traded companies is just 8,501 BTC short of matching Satoshi’s 1,096,358 BTC holdings. Strategy remains the largest public company by digital asset portfolio, with 671,268 BTC. ETFs and funds have long overtaken the Bitcoin creator’s portfolio with their combined 1,496,189 BTC. Various governments worldwide hold an estimated 647,014 BTC. Public treasury

Solana Recovers Above the Crucial $120 Threshold

Solana Recovers Above the Crucial $120 Threshold

// Price Reading time: 2 min Published: Dec 24, 2025 at 17:37 Solana’s (SOL) price has fallen below the moving average lines, but the price range has remained steady above the $120 support and below the moving average lines. Solana price long-term prediction: ranging Buyers were unable to sustain bullish momentum above the

Here's an Early Release from Custody

Here’s an Early Release from Custody

Former Alameda Research CEO Caroline Ellison, sentenced to two years in prison for her role in the misuse of clients’ funds at cryptocurrency exchange FTX, will be released in a matter of weeks following an update from US federal authorities. As of Wednesday, Ellison’s release from federal custody will be Jan. 21, according to information

Send Us A Message

©2025, thefreecurrencyconverter. All Rights Reserved by thefreecurrencyconverty.com

👥 Visitors:

[post-views]