Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

The issue with scrapping is the intensity and volume of bots.

I think that nobody would care if I use wget or curl for few pages, e.g. because I would like to read a site as offline or archive it.

Btw average age of any page is 10 years. Deletion or structural change after acquisition is common, Signal vs Noise site recent wipe out could serve as an example why we need to archive sites.



A lot of websites want "bot defense" due to high volume scrapers, and that "bot defense" often also ends up blocking low-volume wget/curl and polite crawlers like Common Crawl's CCBot.


Cloudflare can verify certain bots when they come from known ip addresses. So if your site is using cloudflare it can let CCBot if it has done the verification.


cloudflare routinely denies my human-piloted browser now, on many sites.


you're not a bot thats irrelevant


I wonder if you'd have more reliable internet browsing with a browser that was a CF verified bot.


> I think that nobody would care if I use wget or curl for few pages

If only you were the only one doing it...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: