4.5B Posts Scraped from TikTok
TheOnlyWayUp
53 points
61 comments
September 03, 2026
Related Discussions
Found 5 related stories in 76.4ms across 5,468 title embeddings via pgvector HNSW
- Scrap (2006) tosh · 361 pts · August 22, 2026 · 48% similar
- The Government Is Monitoring Anti-Flock TikTok and Instagram Accounts Jimmc414 · 101 pts · August 15, 2026 · 48% similar
- The Government Is Monitoring Anti-Flock TikTok and Instagram Accounts Cider9986 · 14 pts · August 12, 2026 · 48% similar
- Detecting scraper bots through scroll behaviour theanonymousone · 32 pts · August 20, 2026 · 44% similar
- In less than a year, Google made scraping 10× harder Growtika · 16 pts · August 31, 2026 · 43% similar
Discussion Highlights (16 comments)
smallerize
There's no way this dataset is going to survive on HF, right? It will be hit with so many DMCA takedowns.
ckugblenu
This is being posted all over the place. on multiple subreddits and stuff. Why?
Retr0id
Can any brave soul wade through the LLM prose to provide a human-readable summary?
igor_nast
Is this even useful in any way?
magicmicah85
>Is it legal? It is against TikTok's terms of service. It is sold for research and educational use. Oh, ok. Otherwise, very detailed deconstruction to scrape their API. Lots of layers of registration and creating a request that looks like it is valid client.
pr337h4m
There aren't any actual videos in the dataset though.
gdevenyi
Like gazing into hell
raver1975
That data does not belong to the public. Why do you think it is OK to steal from TikTok?
moinism
> Everything described here is a private Go repository. One-time payment, permanent access, complete source. > $699 one time · lifetime access Not open-source apparently. And I cant find the reddit post but I think I read that videos/assets are not actually pre-downloaded, they have to be requested through Tiktok API using the provided code. So if Tiktok patches, the code will need updates too.
VulgarExigency
Absolutely no consideration for the people who, when posting things to Tiktok, would prefer for them not to get scraped.
nomilk
> Three things to notice, because each one bites later: Very LLMish language!
vachina
The AI keeps mentioning how a HTTP 200 can silently pollute your dataset. Why not just check contents of body? Usually APIs follow strict JSON contract for successful queries, alert or throw an error when that changes.
VCFundedGenYer
Textbook example of what not to post on the internet. Unreadable incoherent LLM hallucination slop.
1vuio0pswjnm7
1787990085 | X and Meta's past data scraping lawsuit losses and how it would relate to nitter | https://www.reuters.com/legal/musks-x-corp-loses-lawsuit-aga... | https://news.ycombinator.com/item?id=49487864 Perhaps Google can succeed where others like Meta and X have failed, or perhaps not https://storage.courtlistener.com/recap/gov.uscourts.cand.46...
bartleeanderson
Well, most engineers are familiar with GIGO. Garbage in... Err, thats where my comment stops. Garbage in. Yup.
negura
Very impressive figuring out all 4 checks. I'm not even sure this was reverse engineered, rather than leaked. I wish it mentioned anything about the methodology they used