Creepy Crawlies

jay_kyburz 22 points 2 comments August 30, 2026
people.kernel.org · View on Hacker News

Discussion Highlights (2 comments)

never_inline

CGit being a html-generator with arbitrary combinations seems particularly vulnerable. I am in a similar situation - I have a hobby web tool hosted on AWS lambda's free tier behind cloudfront which is quite economical (I usually stay below limit). But over last 2 months I am getting many thousand requests a day. In theory there aren't many crawlable pages in my site but these bots are hammering the search endpoints with arbitrary queries which are somewhat related to the subject matter (must be powered by some weak LMs because they're not valid queries, just generated in a plausible way). I have seen quite a few websites put behind anubis or some other sort of verification system in last 6 months or around. This satire from Krazam [1] was on point. [1]: https://m.youtube.com/watch?v=IU4ByUbDKNc&pp=0gcJCYsCo7VqN5t...

jwilk

https://news.ycombinator.com/item?id=49491791 (over 200 comments at the moment)

Semantic search powered by Rivestack pgvector
4,990 stories · 44,964 chunks indexed