Now that the preview release of musdb’s new UI has finally been released, there’s time to write about the other sides of museum-digital. More usual struggles. Which is to say: New, larger waves of AI scrapers and bots that once again started to impair our services starting July this year.
We maintain that hypocrisy is not for us. I personally have been involved in discussions about the FAIR principles from 2015 onwards. And in 2015 they were already running under that label for a while. As a formalized document, they’ve been around since 2016, but disregarding the label, the principles and surrounding discussions go back, I guess, at least to the 1990s.
In the cultural sector, with many institutions being primarily tax-payer funded, we want to provide our data and services in a Fair, Accessible, Interoperable, and Reusable manner. Relevant here are primarily accessible, interoperable, and _reusable
Realistically, accessibility refers to two things: Immediate, universal barriers like a requirement for authentication should be avoided where possible. On the other hand, user-specific barriers should be reduced as much as possible as well: Human-readable pages should e.g. be designed with sufficient contrast.
If one sees other services and machines as users, which may be reductionist but is logically much more coherent than doing otherwise, then interoperable refers simply to a special form of accessibility: As data should be provided in a way that as many human users can use as possible, it should also be served in formats that machines can read. Ideally in ways that machines can read without further adjustments – following open standards.
Reusable then builds upon these. As the data is now freely accessible by others, setting appropriate licenses allows for the creative reuse of the data.
Ten years after the FAIR principles were formalized, we finally have someone who reuses our data on a massive scale. We should rejoice. Unfortunately all our previous conceptions of how that reuse might take place turned out to be wrong.
Celebrating AI Scrapers?
AI scrapers flooding web services, and especially larger, well-established infrastructures with requests has been an ongoing story for some two years now. Last year museum-digital was met with a first large wave of AI scrapers. See the previous blog posts (1, 2, 3).
On the one hand, using our data for training AI models is undoubtedly reuse. It is even productive reuse: It’s a small contribution to the training of technology that, at its current quality, sounded like science fiction 10 years ago.
On the other hand, this type of reuse takes place without attribution. AI scrapers work based on scale; Machine-readable APIs are uninteresting to scrapers. They suck up as much data as possible, as quickly as possible. Any page-specific adjustments, such as finding and using the API, would reduce the speed of scraping. The requests are so numerous, that they routinely crash servers. Last but not least, those prominently benefitting from the AI hype are far from sympathetic people.
But principles are principles.
It, again, cannot be denied that what AI scraping is eventually aimed at is reuse. Even though they use HTML rather than the APIs we lovingly carved for them, they interoperate with our services in the way most accessible to them.
The problem is, that they are not playing fair. And if the previously unlikely number of requests leads to servers crashing and shutting down, then all principles were for naught. A service that is providing no data at all cannot provide them _FAIR_ly either.
At museum-digital we have seen the new requests as an opportunity to improve our software. If it can withstand the onslaught of thousands of bots a second, it is confirmedly battle-tested and stable. Improving our publishing software on the other hand also meant cutting out overly resource-hungry functionalities. Last year we dropped most publicly available, server-side PDF generation capabilities, we moved the IIIF APIs to an separate configuration that is specced to be able to fall over without impairing the rest of our services. We improved the parsing of search terms and introduced rate-limiting (how many requests a single IP can perform for a given time).
The new wave of AI scrapers since July is even bigger than last year’s. Those improvements were not sufficient to keep our primary server, which hosts the production databases, stable.
From Each According to Their Ability: Load Shedding
We continued on last year’s course however, analyzing the scrapers’ requests, their influence on server stability (which is to say, which requests were especially resource-intensive) and adjusted our setup accordingly. Note, that we did not scale up: museum-digital runs on exactly the same bare-metal servers today as it did two years ago.
In analysis, we identified two types of especially resource-intensive browsing behaviors that are almost entirely limited to scrapers (as well as some very engaged researches, to whom we express our regret for now regularly banning them):
- Scrapers follow links on search pages, especially the facet search, and end up combining more and more search parameters. It is unlikely that a human would perform a search for objects that “Are related to Berlin, and related to Germany, and related to Institution X, but not related to Institution Y, and also related to a time between 1990 and 2000.” For an automated scraper it is normal behavior.
- If a published object page has been updated, a snapshot of its current state is automatically saved to provide an archived version of the page at a given time. This is thought mainly for researches to be able to cite a given state of the page. In everyday operation, archive pages should be barely used. They are also delisted from search machines. But saving snapshots again and again leads to the existence of a large number of links and separate pages scrapers can scrape to no benefit to anybody.
Both functionalities are legitimately useful but scale badly.
We hence introduced load shedding: Whenever a search is performed or an archive page is accessed, the current server load is evaluated. If it is above a certain threshold, the server does not perform complicated search queries (depending on how high load is, it is restricted to a minimum of three search parameters) or does not load the archive page. Instead a warning is presented, that the given action is currently unavailable and a custom HTTP error code is sent. If the same IP causes that error code to be sent twice, the IP is blocked for a while.
As of the time of writing writing, we have thus blocked a total of 131,014,530 IPs in two month, with 5,802,771 being currently blocked.
This approach follows our principles: We try to be fair with everybody. But if users are not fair and do not follow advise, we don’t feel obliged to continue playing fair either.
Unfortunately this approach works on the level of blocking whole IPs. As “residential proxies” – back in the days we called them botnets infecting home routers – have become an ever larger problem, it is not unlikely that actually interested, normal users also unknowingly host an AI scraper at home. I fear I’ve already encountered one such case, were a colleague could not access museum-digital for seemingly no good reason from her home network. Otherwise, it is surprisingly effective – the number of requests has barely been reduced, but our services are essentially stable.
The only crash since happened the day before yesterday and was caused by the combination of bots and an error in our controlled vocabularies – the latter of which is entirely our fault.
Numbers Games: Logging Strategy
After the above-mentioned actions had effectively returned stability to our publicly accessible portals, one issue remained for internal services: File uploads were incredibly slow. It turned out that the large number of requests being permanently logged overburdened the SSD. We have hence disabled the general access logging, restricting ourselves to logging requests that either cause errors or trigger slow responses. This reduces the ability to react well-informed to new waves of requests and actual attacks. Thankfully those also commonly trigger actual errors we can still identify.
Fairness, FAIRness
With these actions we have again averted the need to use more drastic strategies that would limit accessibility to users – human and machine alike – like blanket bans on whole regions of the world or the installation of software such as Anubis. And, again, we have not yet needed to improve our hardware (and pay the price for that), even though that might be in order eventually. It might be wise for other reasons as well.
A holistic picture of one’s setup – software, hardware, systems administration – and the capability to think and act upon these together helps a lot.
But, if anything, hosting in a FAIR, principled way still works in these trying times.
(Post image generated using Anima Base v1)




