Web-scraping AI bots disrupt scientific databases (Nature news, 2025)¶
Reference
Citation: "Web-scraping AI bots cause disruption for scientific databases and journals." Nature (news), 2 June 2025. Type: news report. Link: nature.com.
What it is¶
A news report on automated bots harvesting open scientific resources for AI training data at volumes that degrade or break access for legitimate users. The image repository DiscoverLife (about 3 million species photographs) began taking millions of daily hits in early 2025 and slowed to the point that it no longer loaded. It gives named, dated instances of the load problem that the COAR survey (coar-ai-bots-2025) reports in aggregate.
Role in the record¶
- Grounds BP02: a concrete case that open resources face heavy automated load, and that scraping a resource with no clean interface degrades service for everyone.
- Corroborates the Failures log entry on agents falling back to scraping or legacy endpoints when no task-shaped interface exists.
Atom-level for/against detail and quotes are in the provenance data (assets/provenance.yml), keyed by practice atom.