<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Development | museum-digital: blog</title>
	<atom:link href="https://blog.museum-digital.org/category/development/feed/" rel="self" type="application/rss+xml" />
	<link>https://blog.museum-digital.org</link>
	<description>A blog on museum-digital and the broader digitization of museum work.</description>
	<lastBuildDate>Thu, 01 Oct 2026 12:38:19 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	

<image>
	<url>https://blog.museum-digital.org/wp-content/uploads/2020/01/cropped-mdlogo-code-512px-32x32.png</url>
	<title>Development | museum-digital: blog</title>
	<link>https://blog.museum-digital.org</link>
	<width>32</width>
	<height>32</height>
</image> 
<atom:link rel="search" type="application/opensearchdescription+xml" title="Search museum-digital: blog" href="https://blog.museum-digital.org/wp-json/opensearch/1.1/document" />
	<item>
		<title>Braving the Waves. Surviving in Times of AI.</title>
		<link>https://blog.museum-digital.org/2026/10/01/braving-the-waves-surviving-in-times-of-ai/</link>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 12:30:59 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Bots]]></category>
		<category><![CDATA[Performance]]></category>
		<category><![CDATA[Web Development]]></category>
		<category><![CDATA[Web Hosting]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4755</guid>

					<description><![CDATA[Now that the preview release of musdb&#8217;s new UI has finally been released, there&#8217;s time to write about the other sides of museum-digital. More usual struggles. Which is to say: New, larger waves of AI scrapers and bots that once again started to impair our services starting July this year. We maintain that hypocrisy is <a href="https://blog.museum-digital.org/2026/10/01/braving-the-waves-surviving-in-times-of-ai/" class="more-link">...</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Now that the <a href="https://blog.museum-digital.org/2026/09/30/pre-release-of-musdbs-new-ui/">preview release of musdb&#8217;s new UI has finally been released</a>, there&#8217;s time to write about the other sides of museum-digital. More usual struggles. Which is to say: New, larger waves of AI scrapers and bots that once again started to impair our services starting July this year.</p>



<p class="wp-block-paragraph">We maintain that hypocrisy is not for us. I personally have been involved in discussions about the <em>FAIR</em> principles from 2015 onwards. And in 2015 they were already running under that label for a while. As a formalized document, they&#8217;ve been around since 2016, but disregarding the label, the principles and surrounding discussions go back, I guess, at least to the 1990s.</p>



<p class="wp-block-paragraph">In the cultural sector, with many institutions being primarily tax-payer funded, we want to provide our data and services in a <em>F</em>air, <em>A</em>ccessible, <em>I</em>nteroperable, and <em>R</em>eusable manner. Relevant here are primarily <em>accessible</em>, <em>interoperable</em>, and _reusable</p>



<p class="wp-block-paragraph">Realistically, accessibility refers to two things: Immediate, universal barriers like a requirement for authentication should be avoided where possible. On the other hand, user-specific barriers should be reduced as much as possible as well: Human-readable pages should e.g. be designed with sufficient contrast.</p>



<p class="wp-block-paragraph">If one sees other services and machines as users, which may be reductionist but is logically much more coherent than doing otherwise, then <em>interoperable</em> refers simply to a special form of accessibility: As data should be provided in a way that as many human users can use as possible, it should also be served in formats that machines can read. Ideally in ways that machines can read without further adjustments &#8211; following open standards.</p>



<p class="wp-block-paragraph"><em>Reusable</em> then builds upon these. As the data is now freely accessible by others, setting appropriate licenses allows for the creative reuse of the data.</p>



<p class="wp-block-paragraph">Ten years after the FAIR principles were formalized, we finally have someone who reuses our data on a massive scale. We should rejoice. Unfortunately all our previous conceptions of how that reuse might take place turned out to be wrong.</p>



<h2 class="wp-block-heading">Celebrating AI Scrapers?</h2>



<p class="wp-block-paragraph">AI scrapers flooding web services, and especially larger, well-established infrastructures with requests has been an ongoing story for some two years now. Last year museum-digital was met with a first large wave of AI scrapers. See the previous blog posts (<a href="https://blog.museum-digital.org/2025/12/09/updates-ai-scrapers-and-resilience/">1</a>, <a href="https://blog.museum-digital.org/2025/12/22/cleaning-out-our-closet/">2</a>, <a href="https://blog.museum-digital.org/2025/12/29/trimming/">3</a>).</p>



<p class="wp-block-paragraph">On the one hand, using our data for training AI models is undoubtedly reuse. It is even productive reuse: It&#8217;s a small contribution to the training of technology that, at its current quality, sounded like science fiction 10 years ago.</p>



<p class="wp-block-paragraph">On the other hand, this type of reuse takes place without attribution. AI scrapers work based on scale; Machine-readable APIs are uninteresting to scrapers. They suck up as much data as possible, as quickly as possible. Any page-specific adjustments, such as finding and using the API, would reduce the speed of scraping. The requests are so numerous, that they routinely crash servers. Last but not least, those prominently benefitting from the AI hype are far from sympathetic people.</p>



<p class="wp-block-paragraph"><strong>But principles are principles.</strong></p>



<p class="wp-block-paragraph">It, again, cannot be denied that what AI scraping is eventually aimed at is reuse. Even though they use HTML rather than the APIs we lovingly carved for them, they interoperate with our services in the way most accessible to them.</p>



<p class="wp-block-paragraph">The problem is, that they are not playing fair. And if the previously unlikely number of requests leads to servers crashing and shutting down, then all principles were for naught. A service that is providing no data at all cannot provide them _FAIR_ly either.</p>



<p class="wp-block-paragraph">At museum-digital we have seen the new requests as an opportunity to improve our software. If it can withstand the onslaught of thousands of bots a second, it is confirmedly battle-tested and stable. Improving our publishing software on the other hand also meant cutting out overly resource-hungry functionalities. Last year we dropped most publicly available, server-side PDF generation capabilities, we moved the IIIF APIs to an separate configuration that is specced to be able to fall over without impairing the rest of our services. We improved the parsing of search terms and introduced rate-limiting (how many requests a single IP can perform for a given time).</p>



<p class="wp-block-paragraph">The new wave of AI scrapers since July is even bigger than last year&#8217;s. Those improvements were not sufficient to keep our primary server, which hosts the production databases, stable.</p>



<h2 class="wp-block-heading">From Each According to Their Ability: Load Shedding</h2>



<p class="wp-block-paragraph">We continued on last year&#8217;s course however, analyzing the scrapers&#8217; requests, their influence on server stability (which is to say, which requests were especially resource-intensive) and adjusted our setup accordingly. Note, that we did <strong>not</strong> scale up: museum-digital runs on exactly the same bare-metal servers today as it did two years ago.</p>



<p class="wp-block-paragraph">In analysis, we identified two types of especially resource-intensive browsing behaviors that are almost entirely limited to scrapers (as well as some very engaged researches, to whom we express our regret for now regularly banning them):</p>



<ul class="wp-block-list">
<li>Scrapers follow links on search pages, especially the facet search, and end up combining more and more search parameters. It is unlikely that a human would perform a search for objects that &#8220;Are related to Berlin, and related to Germany, and related to Institution X, but not related to Institution Y, and also related to a time between 1990 and 2000.&#8221; For an automated scraper it is normal behavior.</li>



<li>If a published object page has been updated, a snapshot of its current state is automatically saved to provide an archived version of the page at a given time. This is thought mainly for researches to be able to cite a given state of the page. In everyday operation, archive pages should be barely used. They are also delisted from search machines. But saving snapshots again and again leads to the existence of a large number of links and separate pages scrapers can scrape to no benefit to anybody.</li>
</ul>



<p class="wp-block-paragraph">Both functionalities are legitimately useful but scale badly.</p>



<p class="wp-block-paragraph">We hence introduced load shedding: Whenever a search is performed or an archive page is accessed, the current server load is evaluated. If it is above a certain threshold, the server does not perform complicated search queries (depending on how high load is, it is restricted to a minimum of three search parameters) or does not load the archive page. Instead a warning is presented, that the given action is currently unavailable and a custom HTTP error code is sent. If the same IP causes that error code to be sent twice, the IP is blocked for a while.</p>



<p class="wp-block-paragraph">As of the time of writing writing, we have thus blocked a total of 131,014,530 IPs in two month, with 5,802,771 being currently blocked.</p>



<p class="wp-block-paragraph">This approach follows our principles: We try to be fair with everybody. But if users are not fair and do not follow advise, we don&#8217;t feel obliged to continue playing fair either.<br>Unfortunately this approach works on the level of blocking whole IPs. As &#8220;residential proxies&#8221; &#8211; back in the days we called them botnets infecting home routers &#8211; have become an ever larger problem, it is not unlikely that actually interested, normal users also unknowingly host an AI scraper at home. I fear I&#8217;ve already encountered one such case, were a colleague could not access museum-digital for seemingly no good reason from her home network. Otherwise, it is surprisingly effective &#8211; the number of requests has barely been reduced, but our services are essentially stable.</p>



<p class="wp-block-paragraph">The only crash since happened the day before yesterday and was caused by the combination of bots and an error in our controlled vocabularies &#8211; the latter of which is entirely our fault.</p>



<h2 class="wp-block-heading">Numbers Games: Logging Strategy</h2>



<p class="wp-block-paragraph">After the above-mentioned actions had effectively returned stability to our publicly accessible portals, one issue remained for internal services: File uploads were incredibly slow. It turned out that the large number of requests being permanently logged overburdened the SSD. We have hence disabled the general access logging, restricting ourselves to logging requests that either cause errors or trigger slow responses. This reduces the ability to react well-informed to new waves of requests and actual attacks. Thankfully those also commonly trigger actual errors we can still identify.</p>



<h2 class="wp-block-heading">Fairness, FAIRness</h2>



<p class="wp-block-paragraph">With these actions we have again averted the need to use more drastic strategies that would limit accessibility to users &#8211; human and machine alike &#8211; like blanket bans on whole regions of the world or the installation of software such as <a href="https://github.com/TecharoHQ/anubis/">Anubis</a>. And, again, we have not yet needed to improve our hardware (and pay the price for that), even though that might be in order eventually. It might be wise for other reasons as well.</p>



<p class="wp-block-paragraph">A holistic picture of one&#8217;s setup &#8211; software, hardware, systems administration &#8211; and the capability to think and act upon these together helps a lot.<br>But, if anything, hosting in a <em>FAIR</em>, principled way still works in these trying times.</p>



<p class="wp-block-paragraph"><em>(Post image generated using Anima Base v1)</em></p>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Pre-Release of musdb&#8217;s New UI</title>
		<link>https://blog.museum-digital.org/2026/09/30/pre-release-of-musdbs-new-ui/</link>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 14:53:43 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[User interface]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4730</guid>

					<description><![CDATA[For the last months, a new, fully reworked user interface for musdb has been in the works. Two previous blog posts discussed the reasons and the scope of the overhaul (1, 2). The latter of these also came with our target timeline for the release. Which is to say: It is time for the first <a href="https://blog.museum-digital.org/2026/09/30/pre-release-of-musdbs-new-ui/" class="more-link">...</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">For the last months, a new, fully reworked user interface for musdb has been in the works. Two previous blog posts discussed the reasons and the scope of the overhaul (<a href="https://blog.museum-digital.org/2026/05/05/musdb-reimagined-a-preview/">1</a>, <a href="https://blog.museum-digital.org/2026/09/14/musdb-reimagined-update-release-schedule/">2</a>). The latter of these also came with our target timeline for the release. Which is to say: It is time for the first pre-release.</p>



<p class="wp-block-paragraph">Starting tonight, an overlay will appear, guiding users of musdb to try out the new UI. Which begs the question: Is the new UI ready for general use. Realistically there can be no general answer to that question. We have tried to prioritize the development of commonly used functionalities, while many less commonly used sections of musdb are still entirely missing in the new UI. Before a (nearly) comprehensive list and categorization of the features that are thus far missing in the new UI is presented, let&#8217;s look at a few key new features and benefits of the new UI.</p>



<h2 class="wp-block-heading">New Architecture, New Possibilities</h2>



<p class="wp-block-paragraph">As has been outlined in the previous blog posts multiple times, the most challenging and yet basic aspect of musdb&#8217;s UI rewrite is the move to a different architectural approach. In the old UI, each page is its own script, with all these scripts using the same functions in the background. Opening a page &#8211; or e.g. saving a form &#8211; thus requires a full page reload, including a reload of all the data that remains unchanged. On the face of it, this sounds inefficient in terms of performance and bandwidth use. It is. It is also noticable in immediately perceivable characteristics like the view jumping back to the top of the page when a form is submitted.</p>



<p class="wp-block-paragraph">Refactoring musdb&#8217;s backend code to a more reusable, testable and logical structure is an endeavour that has been ongoing for over half a decade now. The occassional references to musdb&#8217;s growing API in the blog were the immediate benefits of it &#8211; functions that were used to provide the regular, human-readable user interface were also made available for machine-to-machine communications.</p>



<p class="wp-block-paragraph">The new UI is then the result of the final completion of that refactoring: It is a Typescript application built entirely on top of musdb&#8217;s API. On the one hand, this means that potentially any action that users can perform using musdb&#8217;s graphical UI will also be available for automated use using the API. On the other hand, it simplifies things in development: As we are ourselves using the API everywhere, good maintenance is simplfied a lot (see <a href="https://en.wikipedia.org/wiki/Eating_your_own_dog_food">also</a>).</p>



<p class="wp-block-paragraph">Given this new architecture, the user interface is entirely processed in the browser. If forms are submitted, only the changed values are transmitted to the server. The required bandwidth is thus reduced greatly, once the application has loaded. Similarly, navigating only reloads those sections of a page that actually change, reducing computing load for server and client alike (and thus increasing performance).</p>



<h3 class="wp-block-heading">Automation via Webhooks</h3>



<p class="wp-block-paragraph">Only possible due to the new architecture has been the addition of a webhook system in musdb, which offers a simple way to react to updates in musdb in an automatic way. Webhooks can be configured in the institution-level settings pages of the new UI. Note that, as the webhook system requires the use of the new architecture, webhooks configured in the new UI will not take effect when one continues to work in the old one (while both are available). See also: <a href="https://en.wikipedia.org/wiki/Webhook">Wikipedia on webhooks</a>.</p>



<h2 class="wp-block-heading">Design Overhaul</h2>



<p class="wp-block-paragraph">Doing a re-design from the ground up naturally allows for the creation of a much more streamlined user interface. Going by the (albeit limited) feedback we have thus far received on the new UI, we have indeed succeeded in creating one. Similar functionalities that were spread over different sections of a page or different pages altogether have been moved into the same section of a page; redundant or confusing options have been removed. Where things could remain as before &#8211; as even the old design received its fair share of praise despite its shortcomings &#8211; they remain roughly the same.</p>



<p class="wp-block-paragraph">The previous blog posts have already provided some in-depth glances into the new UI. Importantly for this post, the design overhaul will include the removal of some redundant, badly thought-out or effectively counter-productive funtcionalities.</p>



<h2 class="wp-block-heading">List of Features Not Yet Implemented</h2>



<p class="wp-block-paragraph">In the following, a simple list &#8211; in some cases with further remarks &#8211; of the features the old UI offered, that are not or not yet implemented in the new UI will follow. If some feature is not to be found in the list, it is most likely already implemented.</p>



<h3 class="wp-block-heading">Features That Will Certainly Be Implemented, But Are Not Yet</h3>



<h4 class="wp-block-heading">Whole Pages</h4>



<ul class="wp-block-list">
<li>The general search page for object images has not been implemented yet</li>



<li>Visitor counting</li>
</ul>



<h4 class="wp-block-heading">Navigation</h4>



<ul class="wp-block-list">
<li>Notification system</li>
</ul>



<h4 class="wp-block-heading">Misc Tools</h4>



<p class="wp-block-paragraph">There can be accessed in the old UI using the &#8220;Tools&#8221; subsection of the dashboard or the similarly named &#8220;Tools&#8221; submenu of the navigation.</p>



<ul class="wp-block-list">
<li>QR-code generator</li>



<li>Calendar</li>



<li>Process list</li>



<li>To-Do list</li>



<li>Knowledge management / mini-wiki</li>



<li>Hyperlink checker</li>



<li>URL Shortening</li>



<li>Consistency checks for migrating data from free text fields to controlled ones</li>
</ul>



<h4 class="wp-block-heading">Features Available in Multiple Sections of musdb</h4>



<ul class="wp-block-list">
<li>Notes (for any page)</li>



<li>Nextcloud integration</li>



<li>UI for users with less permissions<br><em>The UI should hide submission buttons etc. where a user has no permissions to perform the relevant action. This is by far only implemented in very few cases (while permissions are enforced in the backend).</em></li>
</ul>



<h4 class="wp-block-heading">Object Pages</h4>



<ul class="wp-block-list">
<li>Overview pages
<ul class="wp-block-list">
<li>Sorting menu</li>



<li>Watch list</li>



<li>List view / editing most fields in a tabular view</li>
</ul>
</li>



<li>Sidebar
<ul class="wp-block-list">
<li>Warning if an object is currently reserved</li>



<li>Object record editing status (e.g. open, locked, archived)</li>



<li>Export settings for XML exports; quick exports work</li>



<li>Improvement suggestions</li>
</ul>
</li>



<li>Basic data tab
<ul class="wp-block-list">
<li>Editing translations</li>



<li>Tag suggestions based on the object image</li>



<li>Tag suggestions based on the object description</li>



<li>Ability to add sources / citations for events</li>
</ul>
</li>



<li>Addendum tab
<ul class="wp-block-list">
<li>Data field-specific setting for publication status</li>
</ul>
</li>



<li>Administration tab
<ul class="wp-block-list">
<li>Deaccession (administration tab)</li>
</ul>
</li>



<li>Location tab
<ul class="wp-block-list">
<li>Links to loan management</li>
</ul>
</li>



<li>Most of the restoration / conservation tab</li>



<li>Data history tab
<ul class="wp-block-list">
<li>Versioning UI</li>
</ul>
</li>



<li>Custom / user-defined object editing interace</li>



<li>Comparing objects</li>
</ul>



<h4 class="wp-block-heading">Institution pages</h4>



<ul class="wp-block-list">
<li>Institution editing page
<ul class="wp-block-list">
<li>Tab for self-categorization of institutions / survey</li>
</ul>
</li>



<li>Institution-level settings
<ul class="wp-block-list">
<li>Configuration of scheduled, automatic generation of exports</li>
</ul>
</li>
</ul>



<h4 class="wp-block-heading">Collection pages</h4>



<ul class="wp-block-list">
<li>Hierarchy &amp; sorting</li>



<li>Editing page</li>



<li>Tagging for collections</li>
</ul>



<h4 class="wp-block-heading">Series</h4>



<ul class="wp-block-list">
<li>Series editing pages
<ul class="wp-block-list">
<li>Hierarchy</li>



<li>Sorting objects of the series</li>



<li>Sidebar
<ul class="wp-block-list">
<li>Setting for visibility of the series on the public institution page</li>



<li>Option to group related vocabulary entries for revision in nodac</li>
</ul>
</li>
</ul>
</li>
</ul>



<h4 class="wp-block-heading">User Settings</h4>



<ul class="wp-block-list">
<li>Editing pages for other users
<ul class="wp-block-list">
<li>Configuration of user permissions</li>
</ul>
</li>



<li>Own account settings
<ul class="wp-block-list">
<li>Export of own account data</li>



<li>WebDAV settings (for preparing independently run data imports)</li>



<li>Notification settings</li>
</ul>
</li>
</ul>



<h4 class="wp-block-heading">Contacts / Address Book</h4>



<ul class="wp-block-list">
<li>Contact editing pages
<ul class="wp-block-list">
<li>vCard export</li>



<li>List of linked literature entries (to be implemented on the literature search page)</li>
</ul>
</li>
</ul>



<h4 class="wp-block-heading">Spaces</h4>



<ul class="wp-block-list">
<li>Hierarchy</li>



<li>Editing pages
<ul class="wp-block-list">
<li>Sensor data display</li>
</ul>
</li>
</ul>



<h4 class="wp-block-heading">Exhibitions</h4>



<ul class="wp-block-list">
<li>Editing pages
<ul class="wp-block-list">
<li>Tab for publications linked to the exhibition</li>
</ul>
</li>
</ul>



<h4 class="wp-block-heading">Vocabularies / Backgrounds</h4>



<ul class="wp-block-list">
<li>Unreleased catalogue raisonne feature</li>



<li>Statements</li>
</ul>



<h3 class="wp-block-heading">Features That Are Implemented But Certainly Need to Get Better</h3>



<ul class="wp-block-list">
<li>Object editing page
<ul class="wp-block-list">
<li>Linking recently linked entities on base tab</li>
</ul>
</li>
</ul>



<h3 class="wp-block-heading">Features That Will Be Removed</h3>



<ul class="wp-block-list">
<li>General page features
<ul class="wp-block-list">
<li>Bookmarks<br><em>Browsers are much better at bookmarking than musdb. And musdb runs in a browser anyway.</em></li>
</ul>
</li>



<li>Object page
<ul class="wp-block-list">
<li>Tab: Provenance research<br><em>Feedback from provenance researchers (and common sense and experience) informed us that the current implementation is altogether ill-conceived. A whole separate, general section of musdb for managing research projects might eventually be added to actually do what we aimed to do here.</em></li>



<li>Entering free text data from an appropriate prepared QR-code</li>



<li>&#8220;Ask an expert&#8221; tool</li>
</ul>
</li>



<li>Article section (never really used)</li>



<li>Podcasting features (never published)</li>



<li>Series page
<ul class="wp-block-list">
<li>Offline HTML catalogue<br><em>This solution has not been properly maintained in a decade. With a generally wider availability of internet connections, API-based solutions are also much more flexible.</em></li>



<li>Configuration for display modes of published series pages</li>
</ul>
</li>



<li>Chat</li>



<li>Setting for the start page of a user upon logging in<br><em>The primary and most reliable means by which we can communicate updates to musdb (and other museum-digital tools) is the blog, whose feed is embedded into the dashboard. Making sure that everybody occassionally visits the dashboard thus also means that the likelihood that users know about updates is increased. On the other hand understanding users&#8217; bug reports regarding login and startup is simplified if there is a more unified workflow around logging in.</em></li>
</ul>



<h4 class="wp-block-heading">Features to be Discussed</h4>



<ul class="wp-block-list">
<li>Object page
<ul class="wp-block-list">
<li>Social media buttons setting. With the selection of available social media buttons and current usage trends, it might be wiser to simply scrap the option.</li>
</ul>
</li>
</ul>



<h2 class="wp-block-heading">What Does That Mean For the Roadmap?</h2>



<p class="wp-block-paragraph">As one can quite easily see, a lot is left to be done. But the bulk of the work has been done and we are perfectly on schedule. For anybody who relies on the features that are still missing, the old version will remain available until around November 22. For the majority of users, the current state of musdb should actually already support most or all of their requirements better than the old UI.</p>



<p class="wp-block-paragraph">Until November 22 both versions are available for users to slowly transition over to the new version. We are curious for feedback and continuously push new updates to the new version.</p>



<p class="wp-block-paragraph"><em>(Post image generated using Anima Base v1)</em></p>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>musdb Reimagined &#8211; Update &#038; Release Schedule</title>
		<link>https://blog.museum-digital.org/2026/09/14/musdb-reimagined-update-release-schedule/</link>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Mon, 14 Sep 2026 15:56:53 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[Batch editing]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[User interface]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4682</guid>

					<description><![CDATA[In May 2026 we announced an upcoming large-scale rewrite of musdb&#8217;s user interface. Now, it is taking shape and we can announce an actual timeline. Recap: Scope and Reasons of the Rewrite musdb as a tool has been around since sometime around 2009 or 2010 &#8211; soon after the wider inception of museum-digital. Back then, <a href="https://blog.museum-digital.org/2026/09/14/musdb-reimagined-update-release-schedule/" class="more-link">...</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">In <a href="https://blog.museum-digital.org/2026/05/05/musdb-reimagined-a-preview/">May 2026 we announced</a> an upcoming large-scale rewrite of musdb&#8217;s user interface. Now, it is taking shape and we can announce an actual timeline.</p>



<h2 class="wp-block-heading">Recap: Scope and Reasons of the Rewrite</h2>



<p class="wp-block-paragraph">musdb as a tool has been around since sometime around 2009 or 2010 &#8211; soon after the wider inception of museum-digital. Back then, its scope was to provide an input interface for the recording of publishable museum object data. Since then, the basic architecture to the user interface has remained largely unchanged. While the scope grew to musdb aiming to be a full-featured collection and museum management system, many features that were added were only added &#8220;on top&#8221; &#8211; far enough out of the way to not disturb legacy users or to break existing workflows. In a nutshell: musdb in its current form is quite good, but it&#8217;s an application whose user interface has grown without a clear roadmap and a systematic reconsideration and streamlining of the user interface for over a decade.</p>



<p class="wp-block-paragraph">Since 2009, web development has changed. We do not need to worry about the Internet Explorer anymore, browser compatibility is much better, and web applications can be much more interactive with modern browsers and hardware. To fully make use of these larger advances in the development of musdb, a general change to its architecture is necessary however.</p>



<p class="wp-block-paragraph">As the architecture changes from a traditional web site with many independent pages to a single page application, the rewrite is indeed that: Instead of generating pages on the server and only displaying them in the browser, musdb&#8217;s user interface will be fully generated in the browser, with only the relevant data being queried from the server using dedicated, documented APIs. Even to keep the same functionality, we would need to re-implement its presentation. This means that the rewrite is also a perfect opportunity to streamline the user interface and to reconsider previously unused or overseen functionalities.</p>



<p class="wp-block-paragraph">On the other hand: With the new user interface changing everything display-related on a technical level, that does not mean that much has to change on a conceptual level. The rewritten musdb will feature the same basic division of pages &#8211; lists, editing pages, and pages for adding new entries -, it features the same idea of a tabbed editing pages and &#8211; with few exceptions &#8211; the same data fields and order thereof.</p>



<h2 class="wp-block-heading">A Sneak Preview</h2>



<p class="wp-block-paragraph">Where much stays the same, taking a small glimpse into things that will change might be in order. In the following there are a few such previews with small screen casts. Most of these focus on overview / list pages, which are already in a release-ready form. </p>



<h3 class="wp-block-heading">Navigation</h3>



<figure class="wp-block-video"><video height="1080" style="aspect-ratio: 1920 / 1080;" width="1920" controls src="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-navigation-customization-eng.mp4"></video><figcaption class="wp-element-caption">The screencast shows how menu entries can now be toggled into and out of the quick access navigation. The quick access navigation and the list of other sections of musdb have now been moved closer together.</figcaption></figure>



<h3 class="wp-block-heading">New Filter Options</h3>



<figure class="wp-block-video"><video height="1080" style="aspect-ratio: 1920 / 1080;" width="1920" controls src="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-search-filter-eng.mp4.crf28.mp4"></video><figcaption class="wp-element-caption">As had already begun in the old UI of musdb, more and more types of overviews now feature additional fine-grained filter settings.</figcaption></figure>



<h3 class="wp-block-heading">Display Modes of Overview Pages</h3>



<figure class="wp-block-video"><video height="1080" style="aspect-ratio: 1920 / 1080;" width="1920" controls src="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-search-toggle-display-eng.mp4.crf28.mp4"></video><figcaption class="wp-element-caption">All overview pages in the new UI of musdb feature the same three display modes: grid, list, and table.</figcaption></figure>



<h3 class="wp-block-heading">Batch Editing On All Overview Pages</h3>



<figure class="wp-block-video"><video height="1080" style="aspect-ratio: 1920 / 1080;" width="1920" controls src="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-search-batch-edit-eng.mp4.crf28.mp4"></video><figcaption class="wp-element-caption">All overview pages now feature the basic functionality of selecting entries for batch editing. The actual batch editing capabilities will vary depending on the respective type of entities listed. For now, all publishable entities can be published or unpublished in batch this way. Further batch editing functionalities will be simple to implement upon request.</figcaption></figure>



<h3 class="wp-block-heading">Interactive Overviews</h3>



<figure class="wp-block-video"><video height="1080" style="aspect-ratio: 1920 / 1080;" width="1920" controls src="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-publish-from-list-eng.mp4"></video><figcaption class="wp-element-caption">List views of publishable entries in the database now show whether the respective entries are published directly within the overview in a uniform way. Clicking on the publication status toggles it and publishes hidden entries or hides published ones.</figcaption></figure>



<h2 class="wp-block-heading">Timeline</h2>



<p class="wp-block-paragraph">Now, when will the new UI be done? And when will it be released?</p>



<p class="wp-block-paragraph">As of now, the development is progressing smoothly and the required building blocks for almost all functionalities are there. Depending on what features one uses, the new UI can already replace the old for most common use cases. To fully do so for common use cases means that at least object editing pages need to be implemented fully. This should be the case as of late September. In the following, seminars and a time period where both versions can be used side-by-side will be had to ensure a somewhat smooth transition from the old to the new UI.</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><tbody><tr><td>Sept. 14, 2026</td><td>Blog post</td></tr><tr><td>Sept. 28 &#8211; Oct. 4</td><td>A dialogue will appear on the dashboard of musdb giving notice about the upcoming re-design. It will include a link to the new version.</td></tr><tr><td>Oct. 30</td><td>The new version will be presented at the German user meetup (online). The update will be finished at this point.</td></tr><tr><td>Nov. 1</td><td>The login flow will be updated to link users to the new version first. Here they will see a dialogue similar to the old one, giving notice about the re-design and allowing users to use the old version for the time being.</td></tr><tr><td>Nov. 2 (1 PM)</td><td>Seminar on the update (German)<br><em>Hosted by the museum-digital Deutschland e.V.</em></td></tr><tr><td>Nov. 6 (1 PM)</td><td>Seminar on the update (English)<br><em>Hosted by the museum-digital Deutschland e.V.</em></td></tr><tr><td>Nov. 16 (10 AM)</td><td>Seminar on the update (German)<br><em>Hosted by the museum-digital Deutschland e.V.</em></td></tr><tr><td>Nov. 22</td><td>The dialogue on the dashboard disappears, the old<br>version is removed.</td></tr></tbody></table></figure>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>



<p class="wp-block-paragraph"></p>



<p class="wp-block-paragraph">(Post image generated using Anima Base v1)</p>
]]></content:encoded>
					
		
		<enclosure url="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-navigation-customization-eng.mp4" length="242294" type="video/mp4" />
<enclosure url="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-search-filter-eng.mp4.crf28.mp4" length="266321" type="video/mp4" />
<enclosure url="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-search-toggle-display-eng.mp4.crf28.mp4" length="307178" type="video/mp4" />
<enclosure url="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-search-batch-edit-eng.mp4.crf28.mp4" length="310874" type="video/mp4" />
<enclosure url="https://blog.museum-digital.org/wp-content/uploads/2026/09/musdb-publish-from-list-eng.mp4" length="133430" type="video/mp4" />

			</item>
		<item>
		<title>musdb Reimagined &#8211; A Preview</title>
		<link>https://blog.museum-digital.org/2026/05/05/musdb-reimagined-a-preview/</link>
					<comments>https://blog.museum-digital.org/2026/05/05/musdb-reimagined-a-preview/#respond</comments>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Tue, 05 May 2026 16:32:34 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[User interface]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4667</guid>

					<description><![CDATA[musdb is almost as old as museum-digital itself. From its beginning sometime around 2009 or 2010, it has significantly grown in size and scope. Originally, it was created as an input interface for the publication of museum objects. Today, it aims to be a full fledged collection and museum management solution, capable of covering publishable <a href="https://blog.museum-digital.org/2026/05/05/musdb-reimagined-a-preview/" class="more-link">...</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">musdb is almost as old as museum-digital itself. From its beginning sometime around 2009 or 2010, it has significantly grown in size and scope. Originally, it was created as an input interface for the publication of museum objects. Today, it aims to be a full fledged collection and museum management solution, capable of covering publishable object and exhibition data, internal aspects of collection management such as climate sensor readings and loan management as well as providing general knowledge and project management capabilities.</p>



<p class="wp-block-paragraph">Over all that time however, the basic architecture of musdb remained basically unchanged, with new functionality fit in or added on top where it seemed suitable. Additionally we have been struggling with the duplicate implementation of several functionalities at different times, where one implementation offers a very simple access to the respective data (e.g. a free text field), while the other offers a representation of the same data in a much more fine-grained, precise, and more interlinked way, whose requirements are however harder to fulfill during imports and that requires a more thorough documentation generally.</p>



<p class="wp-block-paragraph">We routinely get very positive feedback from colleagues working in museums, that musdb is comparatively easy to use and logical. The age of the architecture as well as the many small inconsistencies in terms of the user interface that have crept in over time however suggest that it is time for a more radical reimagining of what musdb looks and works like.</p>



<p class="wp-block-paragraph">This post is a first small preview to what musdb should look like starting in the third quarter of 2026. After a more thorough discussion of the reasons for rewriting significant parts of the publications, some first glimpses into the new version are provided at the end of this post.</p>



<h2 class="wp-block-heading">Reasons</h2>



<h3 class="wp-block-heading">Technical</h3>



<p class="wp-block-paragraph">musdb is primarily a PHP application. As such, the internal logic as well as most of the user interface rest on the server. Almost all of the user interface is rendered into HTML on the server and then sent to the browser to generate what a user can view. Some functionalities that require additional interactivity, are presented in overlays, etc. are implemented in JavaScript and processed directly in the browser, with the JavaScript drawing its data both from the HTML as well as APIs (some internal, some documented for internal and external use).</p>



<p class="wp-block-paragraph">Speaking in modern web development terms, PHP covers model, view, and controller on the server, with JavaScript extending the view somewhere half-way. Logically, this is problematic. In brutal practicality, it means that musdb&#8217;s JavaScript components are a collection of small scripts always dependent of some external (PHP/server-side) preparation and not logically interacting with each other. This prevents lots of the benefits of a more structured programming from being used altogether.</p>



<p class="wp-block-paragraph">Our aim with the technical architecture of a reimagined musdb is to design musdb as a modern single page application with all data being provided by clearly documented APIs. This will come with several benefits:</p>



<ul class="wp-block-list">
<li>The view layer is entirely processed entirely in the client, providing a clear logical division between the different tasks</li>



<li>With a modern build infrastructure, code can be reused much more than thus far, providing a better maintainability in the long run</li>



<li>musdb will come with barely any browser-level page reloads. If an API is queried, the application can clearly communicate what is being loaded and for what purpose.</li>



<li>As musdb&#8217;s UI will be entirely built based on the same API as can be used by external developers, that API will be fully feature-complete by the end of the process.</li>
</ul>



<h3 class="wp-block-heading">User interface</h3>



<p class="wp-block-paragraph">The general structure of musdb&#8217;s user interface has proven itself to be quite good. However, as musdb has grown over the years, a lot of small inconsistencies and opportunities for improvement have crept in.</p>



<p class="wp-block-paragraph">Take the navigation alone:</p>



<ul class="wp-block-list">
<li>To access object groups, one can always hover over the three dots sub-menu of the navigation section on the very top right. If one enables the menu option &#8220;object group&#8221; in the personal settings (accessible by clicking on one&#8217;s name), one can also quickly access it via the quick access navigation (the second line of the navigation). As such, simply accessing a central part of musdb is spread over three different sections of the navigation.</li>



<li>There are two question mark icons in the navigation. One, in a speech bubble, links to communication channels and contact persons that one can use to seek help. The other provides links to the handbook. On the one hand, there is a clear logical differentiation between the two. On the other, both look similar and follow rather similar aims. That the two sub-menus are divided rather than being merged into a single sub-menu is not intuitively apparent.</li>



<li>Clicking on the page title or the colored circle at the right of the navigation opens the context-independent search function. Both bear no resemblance to an icon associated with searching and actually follow entirely different primary aims by themselves.</li>
</ul>



<p class="wp-block-paragraph">Many similar issues exist around musdb. In most cases, these somehow make sense and are logical in themselves &#8211; but they still contradict a seamless, intuitive use of the application.</p>



<p class="wp-block-paragraph">Additionally, there are some functionalities that in themselves contradict a more streamlined use of musdb. One such example also accessible via the navigation is the option for users to set their preferred page to be referred to right after they log in. musdb has been built more and more on the premise, that users will start their use of the application on the dashboard, and thus some features &#8211; like requiring users to enable two-factor authentication for regional administrators &#8211; can be bypassed by selecting any non-dashboard start page. The option to select a non-standard start page will hence almost certainly be scrapped in the new version of musdb.</p>



<h2 class="wp-block-heading">Preview</h2>



<p class="wp-block-paragraph">For now, only the basic scaffolding, most of the navigation and some of the overview pages have been implemented in the new version. Screenshots of these can be seen below. Importantly and yet missing, each different section of musdb is planned to come with an icon representing it in the navigation (top left).</p>



<p class="wp-block-paragraph">Note also, that nothing about the new UI is set in stone. These are first drafts. More will follow soon.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="539" src="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_Dashboard_20260504_10h14m08s-1024x539.webp" alt="" class="wp-image-4676" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_Dashboard_20260504_10h14m08s-1024x539.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_Dashboard_20260504_10h14m08s-300x158.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_Dashboard_20260504_10h14m08s.webp 1515w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">The dashboard in the new UI of musdb will offer a reduced, more streamlined and &#8211; importantly &#8211; subdivided selection of tabs. At the top, the new navigation with the yet missing icons is missing (all icons thus far being a location pin).</figcaption></figure>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="993" height="376" src="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-quick-access1_20260505_18h02m50s.webp" alt="" class="wp-image-4671" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-quick-access1_20260505_18h02m50s.webp 993w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-quick-access1_20260505_18h02m50s-300x114.webp 300w" sizes="auto, (max-width: 993px) 100vw, 993px" /><figcaption class="wp-element-caption">The navigation has been reduced into two basic sections: On the left, there is the access to the different main modules of musdb. This section on the one hand comes with links directly visible and accessible from any page. In the same section of the navigation, a hamburger menu offers access to those modules not enabled for quick access. Toggling entries into or out of the quick access navigation is directly accessible in the same place.</figcaption></figure>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="346" height="517" src="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-quick-access2_20260505_18h02m55s.webp" alt="" class="wp-image-4670" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-quick-access2_20260505_18h02m55s.webp 346w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-quick-access2_20260505_18h02m55s-201x300.webp 201w" sizes="auto, (max-width: 346px) 100vw, 346px" /><figcaption class="wp-element-caption">The screenshot depicts those menu options not enabled for quick access.</figcaption></figure>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="366" height="532" src="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-ask_20260504_10h14m38s.webp" alt="" class="wp-image-4672" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-ask_20260504_10h14m38s.webp 366w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-ask_20260504_10h14m38s-206x300.webp 206w" sizes="auto, (max-width: 366px) 100vw, 366px" /><figcaption class="wp-element-caption">The two help menus in the navigation have been merged into one. As we saw when we first introduced the option to create user profiles in musdb, providing additional information on contact people makes users more likely to actually ask for help if they need it. Hence, some data from the user profiles is directly integrated into the navigation.</figcaption></figure>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="399" height="258" src="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-settings_20260504_10h14m34s.webp" alt="musdb's new UI, first draft, screenshot: account settings in navigation" class="wp-image-4669" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-settings_20260504_10h14m34s.webp 399w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_navigation-settings_20260504_10h14m34s-300x194.webp 300w" sizes="auto, (max-width: 399px) 100vw, 399px" /><figcaption class="wp-element-caption">The navigation features one&#8217;s profile picture, if one has uploaded any. Hovering over it gives access to the account settings. In this, the new UI tracks many popular websites and services.</figcaption></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="636" src="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-grid_20260505_18h02m13s-1024x636.webp" alt="" class="wp-image-4675" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-grid_20260505_18h02m13s-1024x636.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-grid_20260505_18h02m13s-300x186.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-grid_20260505_18h02m13s.webp 1527w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Overview and search pages roughly follow the current logic. Some exceptions aside, search results can be displayed as a grid, a list, or in a table. This screenshot shows them in grid mode.<br>It is also the first in this list of screenshots to display the new musdb UI in dark mode.</figcaption></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="609" src="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-list_20260505_18h02m00s-1024x609.webp" alt="" class="wp-image-4674" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-list_20260505_18h02m00s-1024x609.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-list_20260505_18h02m00s-300x179.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-list_20260505_18h02m00s.webp 1529w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Overview and search pages roughly follow the current logic. Some exceptions aside, search results can be displayed as a grid, a list, or in a table. This screenshot shows them in list mode.</figcaption></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="607" src="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-table_20260505_18h02m06s-1024x607.webp" alt="" class="wp-image-4673" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-table_20260505_18h02m06s-1024x607.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-table_20260505_18h02m06s-300x178.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-table_20260505_18h02m06s-1536x911.webp 1536w, https://blog.museum-digital.org/wp-content/uploads/2026/05/Screenshot-musdb_exhibition-overview-table_20260505_18h02m06s.webp 1543w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Overview and search pages roughly follow the current logic. Some exceptions aside, search results can be displayed as a grid, a list, or in a table. This screenshot shows them in table mode.</figcaption></figure>
]]></content:encoded>
					
					<wfw:commentRss>https://blog.museum-digital.org/2026/05/05/musdb-reimagined-a-preview/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>State of Development, December 2025</title>
		<link>https://blog.museum-digital.org/2026/01/12/state-of-development-december-2025/</link>
					<comments>https://blog.museum-digital.org/2026/01/12/state-of-development-december-2025/#respond</comments>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Mon, 12 Jan 2026 17:15:11 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Importer]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[nodac]]></category>
		<category><![CDATA[IIIF]]></category>
		<category><![CDATA[Imports]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[Object editing (musdb)]]></category>
		<category><![CDATA[Object images]]></category>
		<category><![CDATA[Object search (musdb)]]></category>
		<category><![CDATA[Single image view]]></category>
		<category><![CDATA[System administration]]></category>
		<category><![CDATA[User interface]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4616</guid>

					<description><![CDATA[December 2025 was an interesting month for museum-digital. An update to the PHP version used as well as a flood of requests by what is most likely AI scrapers forced us to make changes for improved stability, reducing and reformulating features rather than adding new ones and working on matters of systems administration over purely <a href="https://blog.museum-digital.org/2026/01/12/state-of-development-december-2025/" class="more-link">...</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">December 2025 was an interesting month for museum-digital. An update to the PHP version used as well as a flood of requests by what is most likely AI scrapers forced us to make changes for improved stability, reducing and reformulating features rather than adding new ones and working on matters of systems administration over purely matters of code quite often. Add to that the long-promised update of the terms of use for German museums to more structured and lawyer-approved ones, and you get yet more small changes that do not directly concern the work of museums with museum-digital but rather improve the necessary context.</p>



<h2 class="wp-block-heading">musdb</h2>



<h3 class="wp-block-heading">Object overview</h3>



<p class="wp-block-paragraph">In the default tile view of the object overview page, hovering over an object image thus far revealed the object&#8217;s name. As object names are often too long to display fully and inventory numbers are the primary means of identifying an object in most museums, this preview text has now been extended to include the inventory numer.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="570" src="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-object-list-1024x570.webp" alt="Screenshot in the object overview." class="wp-image-4613" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-object-list-1024x570.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-object-list-300x167.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-object-list-1536x855.webp 1536w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-object-list-2048x1140.webp 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Hovering over an object image in the tile view now also displays the inventory number.</figcaption></figure>



<h3 class="wp-block-heading">User management</h3>



<h4 class="wp-block-heading">New Options for Managing User Accounts: Disabling Accounts &amp; Setting Account Expiry Dates</h4>



<p class="wp-block-paragraph">Two new options on user editing pages allow disabling logins on an account and setting an expiry date for the account. Both can be useful for administration: If a new worker joins the museum for a project with a clear-cut limitation on funding and time, one can now set the account expiry at the beginning of the project to the end of it. The accounts will then automatically be deleted when the project ends. Similarly, colleagues that leave service temporarily but for a prolonged time (e.g. for a sabbatical) and will not need to use their accounts for that time can have their accounts disabled.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="398" src="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-options-1024x398.webp" alt="Screenshot of the user editing page in musdb." class="wp-image-4611" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-options-1024x398.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-options-300x116.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-options-1536x596.webp 1536w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-options-2048x795.webp 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Two new options allow disabling user accounts and setting expiry dates for user accounts.</figcaption></figure>



<h4 class="wp-block-heading">List of Terms of Use</h4>



<p class="wp-block-paragraph">A new tab on a user&#8217;s (own) account settings page provides the option to list all usage agreements / terms of use a user has agreed to in the context of their use of museum-digital / musdb and when they agreed to them.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" src="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-agreement-list-1024x576.webp" alt="Screenshot of the user editing page." class="wp-image-4612" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-agreement-list-1024x576.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-agreement-list-300x169.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-agreement-list-1536x864.webp 1536w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_musdb-user-agreement-list.webp 1949w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">A new tab on the user page lists all user agreements for musdb that the user has agreed to and when they did so.
<br></figcaption></figure>



<h3 class="wp-block-heading">Imports</h3>



<h4 class="wp-block-heading">Limiting Report Mail Size</h4>



<p class="wp-block-paragraph">When a user runs imports themselves using the <a href="https://de.handbook.museum-digital.info/import/importe-selbst-durchfuehren.html">WebDAV upload</a>, the end of the import process &#8211; no matter if it is successful or fails &#8211; is marked by the sending of a report via mail. This report usually contains a list of noteworthy operations that happened during the import, e.g. which objects of which inventory number were imported to which object in musdb, identified by its ID. As imports grow, this list of operation grows. To not encounter issues sending the report, it is henceforth limited to a maximum of 2 MB or 10000 lines.</p>



<h4 class="wp-block-heading">Dry-run Mode</h4>



<p class="wp-block-paragraph">Sometimes it is useful to try running an import to see if it will actually work but not actually process any data. This option has been available in the importer command line interface for a while, among others powering <a href="https://quality.museum-digital.org/">museum-digital:qa</a>. It is now available in the import configuration for self-run imports as well using the setting <code>dry-run</code>. Enabling the setting accordingly stops the importer from actually writing the data into the database and changes the behavior if values that need to be mapped to values in controlled lists at museum-digital are encountered. Usually an import stops the moment such data is to be imported and not yet mapped. During a dry run, the error is collected and the import proceeds. All unmapped entries are listed together at the end of the import, allowing for a simpler mapping (possibly aided by <a href="https://concordance.museum-digital.org/">concordance.museum-digital.org</a>).</p>



<h3 class="wp-block-heading">Dashboard</h3>



<p class="wp-block-paragraph">The first page of the dashboard, which for almost all users also means the start page of musdb right after the login process, was significantly reworked during the last month. The almost entirely unused notetaking features and discourse integration were removed in favor of a feed of recent blog posts. See also the section <a href="https://blog.museum-digital.org/2025/12/29/trimming/">&#8220;Communications&#8221;</a> in the respective blog post.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" src="https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb-1024x576.webp" alt="Screenshot of the dashboard in musdb, as of 2025-12-29." class="wp-image-4594" srcset="https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb-1024x576.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb-300x169.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb-1536x864.webp 1536w, https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb.webp 1920w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">The dashboard in musdb now features a feed of recent news relevant to the development of museum-digital and whatever is going on regionally. The posts are sorted chronologically.</figcaption></figure>



<h3 class="wp-block-heading">Annotations for the Vocabulary Editing Team</h3>



<p class="wp-block-paragraph">Each event, displayed as a tile on object editing pages, featured speech bubble icons behind each time / actor / place to provide additional comments and hints for the central vocabulary editing team. This positioning of the annotation feature led to confusion over the years, with some users using the feature to comment on the relationship between the entity and the object (for which the event notes should be used). We hence repositioned the links and moved them to the respective entity&#8217;s page (e.g. a place page for giving hints and comments on a place entry). The hinting / commenting feature for times has been altogether removed, as providing comments to clarify the meaning of e.g. a year never made much sense.</p>



<h3 class="wp-block-heading">Smaller Updates and Bugfixes</h3>



<ul class="wp-block-list">
<li>Fixed a bug in the HTML generated for listing other objects linked to an object. Links to the other object were broken and are not anymore.</li>



<li>Image editing pages now embed the image directly instead of using the IIIF API. This reduces resource usage and increases stability at no cost.</li>



<li>Removed option to manually trigger the rewriting of EXIF and IPTC metadata of object images. Rewriting takes place in the background whenever an image or a linked object is updated, making user-triggered updates obsolete.</li>



<li>Re-introduce option to repeat linking to the last used linked object</li>



<li>Updated <a href="https://swagger.io/">Swagger UI</a> to version 5.30.3</li>
</ul>



<h2 class="wp-block-heading">Frontend</h2>



<p class="wp-block-paragraph">As stated above and lengthily described in the previous blog posts (<a href="https://blog.museum-digital.org/2025/12/09/updates-ai-scrapers-and-resilience/">here</a>, <a href="https://blog.museum-digital.org/2025/12/22/cleaning-out-our-closet/">here</a>, and <a href="https://blog.museum-digital.org/2025/12/29/trimming/">here</a>) we struggled with stability over the last month. This means that most changes in the frontend are aimed at improving stability.</p>



<h3 class="wp-block-heading">Reworked Default Image Page</h3>



<p class="wp-block-paragraph">Thoroughly described in <a href="https://blog.museum-digital.org/2025/12/09/updates-ai-scrapers-and-resilience/">Updates, AI scrapers, and Resilience</a>, we replaced the default view for single object image pages. While the default view was previously built around the IIIF viewer Mirador, the new default view uses OpenLayers and the unmediated image file for capabilities such as zooming. The new view also brings with it some new features, such as an option to reference specific sections of an image.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="672" src="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-page-1024x672.webp" alt="" class="wp-image-4615" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-page-1024x672.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-page-300x197.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-page-1536x1007.webp 1536w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-page-2048x1343.webp 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">The reworked default image page.</figcaption></figure>



<h3 class="wp-block-heading">Serving Resource-Intensive Pages / Functionalities Only When Resources Are Available</h3>



<p class="wp-block-paragraph">PDF generation, the IIIF Image API, and the suggestions for alternative search queries on failed search pages are now limited to reduce their impact on the overall system stability. This follows two strategies:</p>



<ul class="wp-block-list">
<li>Suggestions on failed search pages and PDF generation will only appear if the overall load on the system is low. The threshold for when or when they are not provided is influenced by the user&#8217;s browser language: If a user uses a browser set to the primary language of a given instance of museum-digital (e.g. German in Hesse, Hungarian in Budapest), the threshold is much higher, meaning users will be able to access the pages at a medium server load. In the case of PDFs, high server load will forward users to the print dialogue for object pages instead of receiving a PDF generated on the server side.</li>



<li>PDF generation and the IIIF Image API are served with a different PHP configuration and set of processes than the rest of the frontend. This configuration significantly reduces available resources for these two functionalities.</li>



<li>The option to generate PDFs featuring all images of an object with between 10 and 40 images has been entirely removed. Given its constraints, the feature was hard to explain and rarely accessible anyway. The primary &#8220;users&#8221; were noticeably AI scrapers.</li>
</ul>



<h3 class="wp-block-heading">Image Search</h3>



<p class="wp-block-paragraph">The image search feature was refactored and reduced to further separate it from the primary object search. The number of available search options has been reduced to be more easily explainable and reduce possibilities for very resource-intensive queries.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="602" src="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-search-1024x602.webp" alt="" class="wp-image-4614" srcset="https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-search-1024x602.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-search-300x176.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-search-1536x903.webp 1536w, https://blog.museum-digital.org/wp-content/uploads/2026/01/20260112_frontend-image-search-2048x1204.webp 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">The reworked image search settings overlay.</figcaption></figure>



<h3 class="wp-block-heading">Batch Export of Object Metadata / OAI</h3>



<p class="wp-block-paragraph">Updated the LIDO API to almost entirely match the LIDO as generated during exports from musdb</p>



<h3 class="wp-block-heading">Smaller Updates and Bugfixes</h3>



<ul class="wp-block-list">
<li>Improved performance of object search by tags and places by filtering searched entities to those who are actually linked in the given instance of museum-digital.</li>



<li>Object groups with only one object are henceforth not displayed and linked on object pages anymore</li>



<li>Fixed link in footer: Clicking on &#8220;museum-digital&#8221; should lead to the home / start page of the given instance of musdb.</li>



<li>Updated <a href="https://swagger.io/">Swagger UI</a> to version 5.30.3</li>
</ul>



<h2 class="wp-block-heading">nodac</h2>



<ul class="wp-block-list">
<li>User-provided comments / hints have been removed for times (see above)</li>



<li>Tooltips for linked objects now display which institution an object belongs to
<ul class="wp-block-list">
<li>This is particularly important for vocabulary editors who do not have access to the museums&#8217; data. This way they get a limited preview with the required information for unpublished objects despite their otherwise lacking permissions.</li>
</ul>
</li>
</ul>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img loading="lazy" decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
					<wfw:commentRss>https://blog.museum-digital.org/2026/01/12/state-of-development-december-2025/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Trimming.</title>
		<link>https://blog.museum-digital.org/2025/12/29/trimming/</link>
					<comments>https://blog.museum-digital.org/2025/12/29/trimming/#respond</comments>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Mon, 29 Dec 2025 01:10:16 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[Feature Removal]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[System administration]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4592</guid>

					<description><![CDATA[The recent issues with server instability have been solved. To do so, we had to significantly reduce resources available to the IIIF API. And in learning from the whole situation, a feed of the most recent relevant blog posts are now displayed to users directly in musdb.]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">In the last weeks we struggled with server stability. As <a href="https://blog.museum-digital.org/2025/12/09/updates-ai-scrapers-and-resilience/">written before</a>, the critical, resource-heavy and publicly available tasks have for a long time been the generation of timelines (and thus complicated search queries) and on the other hand those involving the processing or generation of large files; namely the <a href="https://iiif.io/">IIIF</a> API and PDF generation.</p>



<p class="wp-block-paragraph">In the <a href="https://blog.museum-digital.org/2025/12/22/cleaning-out-our-closet/">last post</a>, I detailed how we severely restricted the availability of the public PDF generation functionalities in museum-digital according to available system resources. That, as it turned out, was not enough to bring reliable stability to our systems. After the server fell over on December 26th once more, we hence moved the IIIF Image API into the same PHP setup used for PDF generation &#8211; meaning that any user/IP can only request the API 10 times a minute and that for any instance of museum-digital, only one PHP worker serves it. This allowed us to severely reduce the maximum available resources per worker for the frontend outside of those two use cases (where the IIIF Image API may use up to 80 MB of RAM, no other part of the frontend will go beyond 5). Since then, the system runs as smoothly as if AI scraping had never become an issue.</p>



<h2 class="wp-block-heading">A Limited Goodbye to IIIF &amp; Server-Side Image Manipulation</h2>



<p class="wp-block-paragraph">Now, what does that mean in practice? On the one hand, we have not fully removed the IIIF image API. All links generated using it remain valid and will be served, even if comparatively slowly.</p>



<p class="wp-block-paragraph">On the other hand the user experience with viewing the images in a IIIF viewer will be significantly worse, even though this strongly depends on the IIIF viewer. The most popular &#8220;full&#8221; IIIF viewers being <a href="https://projectmirador.org/">Mirador</a> and <a href="https://universalviewer.io/">Universal Viewer</a>, significant problems (or a complete inability to use an object&#8217;s images) are to be expected with Mirador. Mirador in its default configuration loads multiple segments of an image separately to then assemble the displayed image from those &#8211; with the creation of the segments happening on the server, thus consuming resources centrally. It also seems to set extremely low limits on accepted response times, which museum-digital&#8217;s IIIF Image API now regularly exceeds due to the aggressive rate limiting. Simply looking at the demo installation of Universal Viewer, the software seems to be much more targeted in its API calls and might still work well despite the restrictions.</p>



<p class="wp-block-paragraph">As far as I know, there are no published numbers on the market share of the different IIIF image viewers. And about whether IIIF viewers external to whoever provides the API are actually regularly used or not. The most jaded &#8211; and likely true &#8211; assumption would be that the share of users who use IIIF without a viewer hosted next to the API is miniscule and that most users will use one of the abovementioned. Our experience, once again, seems to support that hypothesis: We released our implementation of IIIF 2 in 2020, but essentially nobody noticed before we also started hosting a IIIF viewer.</p>



<p class="wp-block-paragraph">As we do use Mirador as a viewer, assume the &#8220;visible&#8221; IIIF image API at museum-digital to be more or less broken. Developers and those making direct use of the API without our installation of Mirador can still benefit from the API. But those are comparatively few.</p>



<p class="wp-block-paragraph">The radical restriction of resources provided to the IIIF Image API is thus likely indeed a goodbye to IIIF, if a limited one. The basic idea is great &#8211; to create a unified way to reference parts of an image (or later a wider media file) and annotate it. In times of significantly increased bot activity, reduced funds, and foreseeably rising hosting costs, our example may be an early sign that the decision to realize that aim by specifying an API to be implemented by the data providers restricts the ability to fully support IIIF to very well resourced institutions. And as funds are shrinking, that is less and less institutions. Let&#8217;s hope that the most basic need IIIF wished to fulfill can be achieved in a different way in the future; one that is accessible to anybody. Realistically this means that computing would need to happen on the client PCs, not on a server.</p>



<p class="wp-block-paragraph">To end the saga on a more positive note: Since we limited the IIIF Image API, our systems run wonderfully smoothly again and we were able to reduce the overall rate limiting on the rest of museum-digital&#8217;s portals. We will monitor the situation and increase the limit slowly to allow more simultaneous API requests without risking stability.</p>



<h2 class="wp-block-heading">Communications</h2>



<p class="wp-block-paragraph">Second, the whole ordeal posed a challenge to our communication channels. If any significant error occurs anywhere on museum-digital, I personally am sent an encrypted error message via mail. Usually. In this case, the primary component falling over was the PHP server, which is also responsible for managing the sending of mails. If a service fell over, the primary way to learn of it was receiving mails about that instead. Reaction times were thus worse than they needed to be. This means that we need to improve our monitoring.</p>



<p class="wp-block-paragraph">On the other hand there was the issue of explaining what was going on. We had a thread about it in the <a href="https://forum.museum-digital.info/d/69-uploading-images-in-musdb-are-slow-and-buggy">forum</a>, which few people read. We had the blog posts. Which few people read. We lack (or lacked) a unified source of information about current events that we can assume people to read. The blog could and should be exactly that.</p>



<p class="wp-block-paragraph">At the top right of the login screen of musdb, the two most recent blog posts from the respective region as well as from the &#8220;development&#8221; category of the blog have been shown for years. Then we turned on the &#8220;remember me&#8221; feature by default, which means that people only very rarely see the login page at all anymore.</p>



<p class="wp-block-paragraph">The first page most users see upon logging in or opening musdb while logged in is the dashboard, the default subsection of which previously offered a summary of the database contents a user has access to, a tile for writing personal notes to oneself, a tile with messages from the respective regional administrators, a tile for the integration of a <a href="https://www.discourse.org">discourse</a> forum, and links to the museum elsewhere on the web.</p>



<p class="wp-block-paragraph">The summary of database contents and the links to the museum elsewhere are certainly useful. The other features not so much. Checking their actual use revealed that barely anybody used any of the note-taking features (likely also because musdb itself offers better alternatives elsewhere), while the discourse integration has not been in use for years. The very first features one sees when opening musdb were thus largely unused, wasting space that could be filled with a feed of relevant blog entries.</p>



<p class="wp-block-paragraph">And so we removed the unused features and replaced them with a more prettily designed feed. This feed now contains the two newest blog posts from the development feed in the user&#8217;s language, the regional or national feed (again in the user&#8217;s language) as well as &#8211; importantly &#8211; the English-language development feed. None of the most recent development-related posts were translated to any language other than the original English, mainly because the time was better spent trying to alleviate or fix the issues than describing them in yet another language. Besides, most people know enough English to grasp the posts. And for those who do not: Community contributions to the blog &#8211; also translations for those who do not &#8211; are always welcome.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" src="https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb-1024x576.webp" alt="Screenshot of the dashboard in musdb, as of 2025-12-29." class="wp-image-4594" srcset="https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb-1024x576.webp 1024w, https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb-300x169.webp 300w, https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb-1536x864.webp 1536w, https://blog.museum-digital.org/wp-content/uploads/2025/12/20251229_screenshot-musdb.webp 1920w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">The dashboard in musdb now features a feed of recent news relevant to the development of museum-digital and whatever is going on regionally. The posts are sorted chronologically.</figcaption></figure>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img loading="lazy" decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
					<wfw:commentRss>https://blog.museum-digital.org/2025/12/29/trimming/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Cleaning Out Our Closet</title>
		<link>https://blog.museum-digital.org/2025/12/22/cleaning-out-our-closet/</link>
					<comments>https://blog.museum-digital.org/2025/12/22/cleaning-out-our-closet/#respond</comments>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Mon, 22 Dec 2025 17:20:40 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[Minor Improvements]]></category>
		<category><![CDATA[System administration]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4586</guid>

					<description><![CDATA[Since the last post (i.e. the update to PHP 8.5 amid an onslaught of AI scrapers) and the later introduction of much stricter per-IP rate limiting, the stability issues around md are better &#8211; but they are not yet completely resolved. As such, we have expanded our efforts in rewriting and reformulating key resource-intensive functionalities <a href="https://blog.museum-digital.org/2025/12/22/cleaning-out-our-closet/" class="more-link">...</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Since the last post (i.e. the update to PHP 8.5 amid an onslaught of AI scrapers) and the later introduction of much stricter per-IP rate limiting, the stability issues around md are better &#8211; but they are not yet completely resolved.</p>



<p class="wp-block-paragraph">As such, we have expanded our efforts in rewriting and reformulating key resource-intensive functionalities for increased stability. Different from before, we have also started to fully remove or disable functionalities that are simply not tenable anymore under the current conditions.</p>



<h2 class="wp-block-heading">PDF Generation</h2>



<p class="wp-block-paragraph">Thus far, there were two basic types of PDFs that were generated (on the server side) in museum-digital&#8217;s portals: PDF representations of object pages (&#8220;data sheet&#8221;) on the one hand and PDFs encapsulating all images of an object in one document for easy printing.</p>



<p class="wp-block-paragraph">The latter was &#8211; simply by nature of its envisioned task &#8211; extremely resource-intensive. All image files had to be loaded from disk, embedded into the PDF, compressed and served. The option had thus been available for fewer and fewer objects. Where it was originally available in case of any object with more than three images, it was later limited to objects of less than 40 images. As such, its availability was increasingly hard to communicate clearly, while its usefulness was relatively reduced with the introduction of a new download option for all images of an object. Its natural resource-intensiveness remained a problem however, and as scrapers will click any link they can find, this type of PDF generation continued to be used quite regularly (every few seconds before the recent surge in bot activity). As of last week, the functionality has been entirely removed.</p>



<p class="wp-block-paragraph">The &#8220;data sheet&#8221; PDF generation has been further limited as well. As stated in the previous blog post, its usefulness is significantly reduced with the introduction of a print stylesheet (you will get better results simply pressing CTRL + P on an object page and printing the page to PDF). Nevertheless, it remained rather popular and has not been removed entirely. To reduce its impact on server stability, we however further limited its availability: If the server load is any higher than comfortable, the PDF will not be generated and an error message will appear. If the load is high (up from around 70% of <em>comfortable</em>) and the user&#8217;s browser language is not the default language of an instance of museum-digital, the same error message will appear.</p>



<h2 class="wp-block-heading">Failed Search Pages</h2>



<p class="wp-block-paragraph">If a search query for objects fails, users are forwarded to a failed search page, on suggestions for alternative search queries are made. This is essentially the same as Google automatically suggesting corrections when search terms contain typos. Identifying the alternatives and offering previews for each is not free. As it is simply suggestions, the benefit or general accuracy of the suggestions fluctuates from case to case.</p>



<p class="wp-block-paragraph">Now, looking at the logs, we had a large number of queries for non-existing entities &#8211; obviously scrapers who were trying out different IDs after analyzing the URL scheme. Each of those queries was executed and then forwarded to the failed search page, triggering the loading of suggestions and previews and thus further using resources on the server for little benefit (besides getting more links to scrape). We have now introduced a similar logic to the limitations on the data sheet PDF generation. Suggestions and previews are only generated when server load is comparatively low, with non-primary language users being slightly disadvantaged vis-a-vis primary-language users in an instance.</p>



<h2 class="wp-block-heading">Timelines</h2>



<p class="wp-block-paragraph">Timelines remain popular &#8211; and a problem. A very common type of query we would see in our logs would combine timelines with searches by start and end date. This was likely due to another possible loop of endless URL generation for scrapers &#8211; specify a timeline until it forwards to search pages for a given timespan, then open the timeline for that timespan. Exactly that behavior has now been made impossible. If a search by a timeline (&#8220;start after&#8221;, &#8220;end before&#8221;) has been set, timelines will not be offered in the sidebar anymore. Trying to generate them for such a search using URL manipulation or the API will return an error page.</p>



<h2 class="wp-block-heading">Search: Cleanup, Image Search &amp; Checking Entity Existence Early</h2>



<p class="wp-block-paragraph">A more messy way of optimizations hit the core of the object search. In around 2021, we introduced a new search logic. Almost all pages relying on the core search logic &#8211; search overview pages, maps for objects, timelines, were adjusted to work with the new logic. The only exception from this was the image search. Still, as the new search logic re-used some of the old search logic&#8217;s functions, we kept both as separate classes, which grew over time. Simply loading the new search logic took about one ms (without OPCache enabled, measured through <a href="https://phpbench.readthedocs.io/en/latest/">PHPBench</a>). This sounds like little, but hints at a lack of modularization of the code and gains relevance with many unpredictable requests with servers automatically spinning up and down.</p>



<p class="wp-block-paragraph">And indeed, in writing the new search logic, we did not modularize thoroughly HTML generation, query building and database querying. With last weeks updates, there are now separate classes for each of these and functionalities relevant only to the old search functions have been moved to class managing the image search logic. This reduces startup time for only the new / main search logic by about half (ca. 0.6 ms).</p>



<p class="wp-block-paragraph">Second, we reduced the available search options for image searches. The remaining search parameters are either those actually relevant to the images or those linked to the controlled vocabularies. As a positive side effect, this also solves some issues in communication: Making it legible what the difference between searching images by their own license and by the license of (unrelated) metadata of objects the images are linked to is, is complicated.</p>



<p class="wp-block-paragraph">Finally, as stated above, the logs revealed a lot of queries for objects linked to e.g. either entirely non-existent places or places that are not linked to any object in the instance of museum-digital altogether. When a place or tag is queried, we hence check whether there exists any public mention of the entity in the current instance of museum-digital during query building. If there is no link at all, it is clear early on that a more detailed (i.e. costly) query combining the search by that entity with other parameters will not return any results.</p>



<h2 class="wp-block-heading">The Current Situation</h2>



<p class="wp-block-paragraph">All these improvements help, but a look at the current real-world numbers is warranted. On the one hand, the database server now often falls down to half or even less of the expected server load. This is a positive sign for system stability outside of peak times.</p>



<p class="wp-block-paragraph">On the other hand, there are noticably spikes in the morning (around 10:20 in Germany) and in the afternoon (starting around 5 p.m.). The spike in the morning is likely related to the start of workdays and has led to the server falling over multiple times last week. This can likely be fixed only with a further tuning of the PHP-FPM settings. The spikes in the afternoon and early evening on the other hand remain hard to explain, but are altogether much less critical.</p>



<p class="wp-block-paragraph">We&#8217;re on it.</p>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img loading="lazy" decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
					<wfw:commentRss>https://blog.museum-digital.org/2025/12/22/cleaning-out-our-closet/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Updates, AI scrapers, and Resilience</title>
		<link>https://blog.museum-digital.org/2025/12/09/updates-ai-scrapers-and-resilience/</link>
					<comments>https://blog.museum-digital.org/2025/12/09/updates-ai-scrapers-and-resilience/#respond</comments>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Tue, 09 Dec 2025 00:11:30 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[System administration]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4580</guid>

					<description><![CDATA[Between Thursday last week (November 27th) and yesterday (December 6th), museum-digital has seen its most instable week in about four years. Now that the dust has settled a bit, there's finally some time to discuss what happened and how we managed to tackle the multiple issues leading to the (very noticeable) instability.]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Between Thursday last week (November 27th) and yesterday (December 6th), museum-digital has seen its most instable week in about four years. Now that the dust has settled a bit, there&#8217;s finally some time to discuss what happened and how we managed to tackle the multiple issues leading to the (very noticeable) instability.</p>



<h2 class="wp-block-heading">Background</h2>



<h3 class="wp-block-heading">Scrapers</h3>



<p class="wp-block-paragraph">There were (or are) two factors simultaneously pushing our servers to their limits and requiring changes. On the one hand, scraping of museum-digital has gotten even more aggressive. Where we usually has something around 10-30 requests per second across all of museum-digital a year ago, we had around 300 two weeks ago. Right now it&#8217;s often between 500 and 700. This number excludes any access to static files.</p>



<p class="wp-block-paragraph">As I&#8217;ve written elsewhere, the scrapers are mostly noticable by coming from IP ranges in Asia or (to a lesser extent) the US. On the other hand, the IPs change constantly and user-agents etc. resemble regular users. Likely they simply use an actual chrome browser for scraping. Which is to say, attempting to block them is futile. Worse yet, attempts to block scrapers would likely also impact some real users.</p>



<p class="wp-block-paragraph">Fortunately museum-digital is run on dedicates servers paid by time rather than by compute. The onslaught of scrapers thus has no financial impact on us. But the scrapers still use resources, and as they try to scrape as many different pages as possible, it is much harder to optimize for them than it is to optimize for actual human users (see this article on a similar issue at <a href="https://arstechnica.com/information-technology/2025/04/ai-bots-strain-wikimedia-as-bandwidth-surges-50/">Wikimedia</a>).</p>



<p class="wp-block-paragraph">Either way, AI scrapers can result in improvements. Viewed positively, they essentially act as a free stress test on a service and enforce efficiency in all aspects. If most pages are optimized for performance already, scrapers will find the unoptimized ones and bring down a service by overusing those. Which is to say, they help to identify yet unoptimized scripts/pages/classes and enforce that necessary changes are made. At museum-digital, there are three main weak spots that are hard to optimize: timelines, image manipulation (including the IIIF API), and PDF generation.</p>



<h3 class="wp-block-heading">PHP</h3>



<p class="wp-block-paragraph">On November 20th PHP 8.5 was released. Thus far, museum-digital had been running on PHP 8.3 for web hosting and PHP 8.4 on the command line. When we attempted to update to 8.4 last year, the server fell over. This was mainly caused by the IIIF API (and thus, image manipulation via <a href="https://www.libvips.org/">libvips</a>).</p>



<p class="wp-block-paragraph">Dependencies at museum-digital are (like pretty much universal with PHP) handled using the package manager <code>composer</code>. Setting up a new instance of museum-digital, composer (managed on version 8.4) required PHP 8.4 or later to run &#8211; the new instance was thus unable, being stuck on version 8.3 for hosting.</p>



<p class="wp-block-paragraph">That leaves two options: Either to set up composer using PHP 8.3 again, or to simply update everything to the current version. While PHP 8.3 will be <a href="https://www.php.net/supported-versions.php">supported until 2027</a>, it is generally advisable to update when possible. So updating it was.</p>



<p class="wp-block-paragraph">Importantly, PHP at museum-digital is run via <a href="https://www.php.net/manual/de/install.fpm.php">PHP-FPM</a>. Before the update, we had one socket running per subdomain. This means, that if a PHP process serving the frontend stopped working for any reason, users in musdb were impacted as well.</p>



<h2 class="wp-block-heading">Upgrading PHP to version 8.5</h2>



<p class="wp-block-paragraph">Once we upgraded the PHP version to 8.5 on Thursday, the same problems we faced with PHP 8.4 appeared again. The server would run rather smoothly for some hours, then more and more PHP processes would die and PHP-FPM would fall over for a given subdomain, and users would get a 504 gateway timeout error. Again, the IIIF API and image manipulation were the main causes of PHP-FPM getting stuck. Of course, the number of AI scrappers continuing to use the site did not help.</p>



<h3 class="wp-block-heading">PHP-FPM settings</h3>



<p class="wp-block-paragraph">A natural first point to consider was the configuration of PHP-FPM. PHP-FPM knows three basic modes for running an application:</p>



<ul class="wp-block-list">
<li><code>ondemand</code> You define a maximum number of processes the application may use. When a new request is made, idle processes get used. If there is no idle process, PHP-FPM starts a new one. After a specified number of requests or a given number of seconds, an old process is closed. This is primarily aimed at being able to scale way down &#8211; if there is no requests, there will be no processes (which is to say, less resources used). On the other hand, starting new processes takes time.</li>



<li><code>static</code> You define a number of processes that should always be running for the application. This means that there should always be processes already started and ready for usage, but it also means that those processes take up resources even when they are little used. Which is to say, this is useful if one has a high and constant stream of users.</li>



<li><code>dynamic</code> You define a maximum number of processes, as well as how many processes should be always running for immediate use, and a (minimum and maximum) number of spare processes to keep running. PHP-FPM then manages if more processes should be started or if one of the already running ones shall be used. This, in theory, is useful if one wants to reliably and quickly serve users, expects some use all the time, but wants the server to dynamically scale up and down as needed.</li>
</ul>



<p class="wp-block-paragraph">With museum-digital spread out over around 80 subdomains, we had thus far used the <code>ondemand</code> mode for most subdomains. Only the largest and most used instances / subdomains of museum-digital were run using <code>dynamic</code> mode. With the update to PHP 8.4 and then 8.5, the behavior of the <code>ondemand</code> mode seems to have changed. If one process dies, the whole subdomain goes seems to go down with it (I have not found a documentation on this, but it&#8217;s evident from the last two weeks).</p>



<p class="wp-block-paragraph">We hence moved critical subdomains impacted by the errors (which is to say, any &#8220;regular&#8221; instance of museum-digital) to dynamic mode. As dynamic mode enforces stricter limits on how many processes can be run respective to the available hardware (which is to say, dynamic mode requires a better-written configuration), this also meant that we needed to adjust the specified numbers of processes per subdomain according to their use.</p>



<p class="wp-block-paragraph">To actually grasp <em>real</em> use of a subdomain including bots, we turned to the logs we keep for about a week (and then rotate out). In server logs, usually one line corresponds to a single request. With a small script, we loop all the different subdomains and check how many requests were made. To be really sure that only requests to relevant PHP scripts are processed, we filter them by the presence of the substring &#8220;php&#8221; before counting. The result for today between 1 a.m. and 4 p.m. looks as follows:</p>



<pre class="EnlighterJSRAW" data-enlighter-language="raw" data-enlighter-theme="" data-enlighter-highlight="" data-enlighter-linenumbers="" data-enlighter-lineoffset="" data-enlighter-title="" data-enlighter-group="">| Requests count in instance                         |      Total |      musdb |        PDF | 
| -----                                              |      ----- |      ----- |      ----- | 
| agrargeschichte.museum-digital.de                  |     341508 |       1245 |        719 | 
| bawue.museum-digital.de                            |     454228 |      12559 |       6819 | 
| bayern.museum-digital.de                           |     176291 |          0 |        158 | 
| berlin.museum-digital.de                           |     223280 |      14917 |       6814 | 
| brandenburg.museum-digital.de                      |      63286 |       6927 |       3873 | 
| bremen.museum-digital.de                           |     221208 |          0 |       2026 | 
| bund.museum-digital.de                             |        261 |        167 |          5 | 
| collectors.museum-digital.de                       |     108398 |        449 |        648 | 
| hamburg.museum-digital.de                          |      35489 |          0 |         11 | 
| hessen.museum-digital.de                           |      50932 |       7962 |       2486 | 
| meckpomm.museum-digital.de                         |      94177 |         11 |        139 | 
| nds.museum-digital.de                              |     137703 |       4105 |       4134 | 
| owl.museum-digital.de                              |     427667 |       1258 |       2412 | 
| rheinland.museum-digital.de                        |      64838 |       1753 |       1276 | 
| rlp.museum-digital.de                              |     207944 |       7405 |       7532 | 
| sachsen.museum-digital.de                          |     120931 |      16117 |       6034 | 
| saarland.museum-digital.de                         |        210 |          0 |          1 | 
| smb.museum-digital.de                              |     228542 |          0 |      11517 | 
| sh.museum-digital.de                               |      21098 |          0 |         48 | 
| st.museum-digital.de                               |     317913 |       6243 |       6217 | 
| thue.museum-digital.de                             |     117893 |          0 |        495 | 
| westfalen.museum-digital.de                        |     101584 |       2033 |       3310 | 
| br.museum-digital.org                              |      43413 |          0 |         16 | 
| jateng.id.museum-digital.org                       |        211 |          0 |          0 | 
| jatim.id.museum-digital.org                        |      23410 |          0 |        159 | 
| lazio.it.museum-digital.org                        |        295 |          0 |          0 | 
| ma.pl.museum-digital.org                           |        385 |          0 |          0 | 
| noe.at.museum-digital.org                          |     906386 |          0 |        369 | 
| tirol.at.museum-digital.org                        |        537 |          0 |          7 | 
| vbg.at.museum-digital.org                          |         96 |          0 |          0 | 
| wien.at.museum-digital.org                         |     472305 |        586 |       3243 | 
| ulster.ie.museum-digital.org                       |      28869 |          0 |          2 | 
| connacht.ie.museum-digital.org                     |        392 |          0 |          0 | 
| va.srb.museum-digital.org                          |       5599 |          0 |         22 | 
| ko.rou.museum-digital.org                          |       9036 |        635 |        567 | 
| mm.rou.museum-digital.org                          |        235 |          0 |          0 | 
| ca.usa.museum-digital.org                          |       3946 |          0 |          0 | 
| ma.usa.museum-digital.org                          |        357 |          0 |          0 | 
| ny.usa.museum-digital.org                          |      19576 |          0 |        294 | 
| syddanmark.dk.museum-digital.org                   |        675 |          0 |          9 | 
| de.pt.museum-digital.org                           |       1241 |          0 |         29 | 
| zh.ch.museum-digital.org                           |     233280 |        512 |        650 | 
| ba.hu.museum-digital.org                           |      99927 |       1901 |         72 | 
| be.hu.museum-digital.org                           |     100830 |        244 |       3005 | 
| bk.hu.museum-digital.org                           |     489446 |         55 |       3985 | 
| bu.hu.museum-digital.org                           |     213616 |       6206 |       5753 | 
| bz.hu.museum-digital.org                           |     598550 |        680 |       1788 | 
| cs.hu.museum-digital.org                           |      88585 |          0 |       1054 | 
| fe.hu.museum-digital.org                           |     199812 |          7 |        215 | 
| gs.hu.museum-digital.org                           |     216680 |       4215 |        912 | 
| hb.hu.museum-digital.org                           |      61250 |          0 |         65 | 
| he.hu.museum-digital.org                           |      26312 |          7 |         26 | 
| jn.hu.museum-digital.org                           |      11970 |          0 |        131 | 
| ke.hu.museum-digital.org                           |     370219 |       2959 |       1680 | 
| no.hu.museum-digital.org                           |     119487 |          0 |       1545 | 
| pe.hu.museum-digital.org                           |     603846 |       2957 |       1446 | 
| so.hu.museum-digital.org                           |     308116 |       6151 |       6698 | 
| sz.hu.museum-digital.org                           |        116 |          0 |          0 | 
| to.hu.museum-digital.org                           |      52406 |          0 |       1229 | 
| va.hu.museum-digital.org                           |     184231 |       2839 |       1666 | 
| ve.hu.museum-digital.org                           |    1015509 |       3672 |        296 | 
| za.hu.museum-digital.org                           |        199 |          0 |          6 | 
| ce.cz.museum-digital.org                           |          3 |          0 |          0 | 
| ccc.cz.museum-digital.org                          |         17 |          0 |          0 | 
| academia.hu.museum-digital.org                     |       9158 |          0 |         13 | 
| cherkasy.ua.museum-digital.org                     |      25567 |          0 |         26 | 
| chernihiv.ua.museum-digital.org                    |       3258 |         99 |        156 | 
| dnipro.ua.museum-digital.org                       |      26725 |          0 |        109 | 
| donetsk.ua.museum-digital.org                      |         17 |          0 |          0 | 
| ivfr.ua.museum-digital.org                         |        722 |          0 |          9 | 
| kharkiv.ua.museum-digital.org                      |      12932 |          0 |         39 | 
| kyiv.ua.museum-digital.org                         |     436482 |       5967 |       1351 | 
| kyivska.ua.museum-digital.org                      |       2159 |          0 |         79 | 
| lviv.ua.museum-digital.org                         |     163358 |        188 |        274 | 
| poltava.ua.museum-digital.org                      |       7657 |        284 |          3 | 
| odesa.ua.museum-digital.org                        |         93 |          0 |          1 | 
| rivne.ua.museum-digital.org                        |      59510 |         65 |        156 | 
| sumy.ua.museum-digital.org                         |      35890 |        303 |          3 | 
| ternopil.ua.museum-digital.org                     |     150700 |         37 |        184 | 
| zhytomyr.ua.museum-digital.org                     |          3 |          0 |          0 | 
| vinnytsia.ua.museum-digital.org                    |      14229 |          0 |          0 | 
| volyn.ua.museum-digital.org                        |      16705 |          0 |        485 | 
| zakarpattia.ua.museum-digital.org                  |       2865 |          0 |         30 | 
| zaporizhzhia.ua.museum-digital.org                 |      24348 |        338 |         56 | 
| scotland.museum-digital.org                        |          0 |          0 |          0 | 
| md.museum-digital.org                              |          0 |          0 |          0 | 
| demo.museum-digital.org                            |         12 |          2 |          0 | 
| goethehaus.museum-digital.de                       |     260072 |          0 |         85 | 
| lmw.museum-digital.de                              |     326724 |          0 |         65 | 
| gedenkstaetten.museum-digital.de                   |       3474 |          0 |          0 | 
| turcica.museum-digital.de                          |      75533 |          0 |          1 | 
| nat.museum-digital.de                              |    1238860 |          0 |       4657 | 
| at.museum-digital.org                              |     631578 |          0 |         89 | 
| cz.museum-digital.org                              |          2 |          0 |          0 | 
| dk.museum-digital.org                              |       5415 |          0 |          4 | 
| hu.museum-digital.org                              |     359619 |          0 |       2827 | 
| id.museum-digital.org                              |       8030 |          0 |          0 | 
| ie.museum-digital.org                              |       2073 |          0 |          0 | 
| it.museum-digital.org                              |         78 |          0 |          0 | 
| rou.museum-digital.org                             |       8277 |          0 |        466 | 
| pl.museum-digital.org                              |        142 |          0 |          0 | 
| pt.museum-digital.org                              |          0 |          0 |          0 | 
| srb.museum-digital.org                             |        565 |          0 |          0 | 
| ua.museum-digital.org                              |     232115 |          0 |        805 | 
| usa.museum-digital.org                             |       3752 |          0 |         34 | 
| ch.museum-digital.org                              |      53417 |          0 |          1 | 
| global.museum-digital.org                          |     727690 |          0 |       2199 |</pre>



<p class="wp-block-paragraph">Note that the number of requests obviously is also impacted by bots changing attention &#8211; once a scraper is done with one subdomain, they turn to the next. The elevated number of requests in ve.hu.museum-digital.org is normal, but still starkly exaggerated when compared to other days. The Germany-wide instance is persistently the most frequented one, usually the global one is second at around 80% of requests.</p>



<p class="wp-block-paragraph">Now equipped with actual numbers, we could scale the PHP-FPM to a much more suitable configuration than before (we had thus far never bothered counting actual requests, instead relying on the number of objects).</p>



<p class="wp-block-paragraph">A second step in the PHP-FPM configuration was to reduce the impact the problems had. Previously there was one shared configuration and socket per subdomain. On the one hand, this meant that stuck processes in the frontend impacted users in musdb (and vice-versa). On the other hand, some constraints on resource usage cannot be set on a per-directory level but must be set per PHP-FPM socket / server (see the PHP documentation on <a href="https://www.php.net/manual/en/configuration.file.per-user.php">user.ini</a> and the list of <a href="https://www.php.net/manual/en/ini.list.php">php.ini directives</a>). As the frontend and musdb have different requirements (frontend: low maximum memory use, short timeouts, no file uploads, generally strict settings; musdb: long timeouts for uploads, generally more lenient), being able to configure them independent of each other is useful in general.</p>



<p class="wp-block-paragraph">We thus separated the configuration for the frontend, musdb, and PDF generation in the frontend; providing dedicated sockets for each. The frontend has a reduced <a href="https://de.wikipedia.org/wiki/Nice_(Unix)">priority</a> on the system overall, strict constraints on how it may be used, etc. The settings are stricter than they were before. musdb has an elevated priority and more lenient settings (file uploads, longer timeouts), in fact more lenient than before. Finally, PDF generation is a special case as it offers no real benefit over the browser&#8217;s print tool (see MDN on <a href="https://developer.mozilla.org/en-US/docs/Web/CSS/Guides/Media_queries/Printing">print CSS</a>), while being resource-intensive. As such, it has a far reduced priority and very strict settings.</p>



<p class="wp-block-paragraph">With the separated configuration and sockets, we can now better tailor the configuration to each application&#8217;s needs and have the added benefit of problems in one application not impacting the other.</p>



<h3 class="wp-block-heading">Code</h3>



<p class="wp-block-paragraph">As we had already prepared the codebase for PHP 8.4 awaiting an eventual upgrade, the upgrade to PHP 8.5 only required minimal changes. Aside from the deprecation of the functions <code>finfo_close()</code> and <code>curl_close()</code>, references to which were accordingly removed from the code, the update necessitated no further work.</p>



<h2 class="wp-block-heading">Scaling in Software</h2>



<p class="wp-block-paragraph">Improving the PHP configuration was not enough to fix the issues, especially with the now increased number of requests from bots. To get some breathing room, we adjusted the most resource-intensive pages.</p>



<h3 class="wp-block-heading">Frontend</h3>



<p class="wp-block-paragraph">In the frontend these are, again, the IIIF API, PDF generation, and timelines. Finally, we made changes to the pages for failed searches to better handle high load situations.</p>



<h4 class="wp-block-heading">Image pages</h4>



<p class="wp-block-paragraph">The IIIF API was used for the main image pages in the frontend. We used (and use) <a href="https://projectmirador.org/">Mirador</a> as a IIIF viewer. Simply opening an image page thus meant three requests to fetch different regions of an image. Zooming into the image triggered further requests to fetch the relevant parts of the image. Cropping the image to the requested region with IIIF happens on the server (which is no problem if there are few users, but is turning into a problem when you have hundreds of requests per second).</p>



<p class="wp-block-paragraph">We thus changed the default of image pages: The new default image page is the old, non-IIIF one. As features like zooming into images, that Mirador comes with, are popular and useful and the old image page did not support those, we worked to improve the page. To do so, we rely on <a href="https://openlayers.org/">OpenLayers</a>, a library we already use for maps. Besides including maps from tile servers, OpenLayers also supports loading simple image files &#8211; which we do here. The image is hence loaded once in full size and zooming etc. happen entirely in the browser.</p>



<p class="wp-block-paragraph">Taking the opportunity, we improved the page overall. An often noticed problem of image pages thus far was, that users who opened image pages coming from external services (think Google Images) had problems identifying that the image was an object image and that there is further object data to be found on object pages. The updated image pages now come with a header stating reflecting the name of the image, the name of the object and the name of the institution. Note that many images do not feature a dedicated title, musdb uses the object name as a default image title in that case, which is why the object title will often appear twice in the header. Maybe this can be used as an encouragement for the colleagues working in musdb to more consistently set expressive image titles in the future.</p>



<p class="wp-block-paragraph">Also new is a mini map at the bottom left, displaying where in the wider context of the image one has currently zoomed in, as well as the ability to link exactly the region one has currently zoomed into. To enable the latter, the URL updates as one zooms or navigates around the image. Somebody else opening the same URL will then open exactly the same image region the linking person was viewing when copying the URL. Finally, we finally set specific <a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/CSP">Content Security Policies</a> relevant to the currently opened media. If the displayed media entry is an internally stored image, no external images need to be allowed to load. If the displayed media entry is an audio file stored on archive.org, archive.org needs to be whitelisted as a source for audio files &#8211; but only archive.org and no other page. Previously, embedding images from anywhere on the net was allowed, increasing the potential damage a potential attacker may cause.</p>



<p class="wp-block-paragraph">Making the use of Mirador a secondary, non-default option reduced the need for server-side image manipulation and the corresponding resource use significantly. The IIIF remains largely unchanged, but its use must now be requested explicitly.</p>



<h4 class="wp-block-heading">PDF generation</h4>



<p class="wp-block-paragraph">As stated above, PDF generation brings little advantages to the browser&#8217;s print functionality in combination with object pages. On the contrary, the PDFs generated using the frontend&#8217;s templates feature less information. But they come with the file ending &#8220;.pdf&#8221; and seem to be extremely popular with bots. On the other hand, PDF generation means, among others, loading whatever images are to be embedded into the PDF and manipulating them fit into the PDF. The resulting files are significantly larger than the corresponding HTML files and thus also use more of the available bandwidth.</p>



<p class="wp-block-paragraph">The update to handle PDF generation respective to resource usage was already introduced in the last months: publicly linked PDFs are now only generated if overall load on the server is low, if a user has set their <a href="https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Accept-Language">browser language</a> to any language different from a museum-digital instance&#8217;s default language. As most scrapers do not bother to change their browser language (which means they come with either none, English or Chinese), this means they will mostly be unable to trigger the generation of PDFs. They see an error page instead.</p>



<h4 class="wp-block-heading">Failed Search Pages</h4>



<p class="wp-block-paragraph">If a user tries to execute a search query without any results, they will get suggestions for similar search terms &#8211; similar to how Google will ask one searching for &#8220;Berrlin&#8221;, if they meant &#8220;Berlin&#8221;. Trying to identify suitable suggestions obviously costs resources and whether the suggestions are actually what a user wanted is by nature hit or miss &#8211; it&#8217;s suggestions after all. In the case of scrapers, suggesting alternative search queries offers them a never-ending stream of possible search queries to run and keep scraping the subdomain with &#8211; to nobody&#8217;s benefit (not even the scrapers&#8217;, as they likely got the same content with other search queries already).</p>



<p class="wp-block-paragraph">We thus now use the same function used to identify whether PDFs should be generated for a user to check if search suggestions should be provided. It a user comes with a non-default browser language and resource use is high, no suggestions will be provided.</p>



<h4 class="wp-block-heading">Timelines</h4>



<p class="wp-block-paragraph">Timeline pages as implemented in museum-digital&#8217;s frontend offer another source of endless links and search queries, as they link to further and further specifications of the time searched by. Again, an improvement already introduced months ago, was to better parse queries by time: If a user searches for objects that are linked to times &#8220;after 1920&#8221; and &#8220;after 1930&#8221;, the latter already includes the former. &#8220;After 1920 and after 1930&#8221; means exactly the same as &#8220;after 1930&#8221;. Which is one <a href="https://en.wikipedia.org/wiki/Join_(SQL)">join</a> instead of two &#8211; half the resource usage.</p>



<p class="wp-block-paragraph">A minor improvement we noticed on the side was impact of automatic redirects in the timelines. Say, a user searches objects by their link to a given tag and then generates a timeline for said objects. If all objects were created in the 20th century, the timeline will automatically redirect so as to &#8220;zoom&#8221; into a more appropriate time scale than from the big bang to now. Until the last weekend, script execution was not stopped when that redirect happened &#8211; which means that all database queries for time time from the big bang to now were still executed even though the user never got to see them. That is now fixed.</p>



<h2 class="wp-block-heading">The Anti-Climactical Solution</h2>



<p class="wp-block-paragraph">All of those changes got the frontend more or less stable. Problems with uploading images remained however. Finally, the only thing that helped was uninstalling libvips (which we use for image manipulation) and reinstalling it. That seems to have fixed the issues.</p>



<p class="wp-block-paragraph">Especially as the number of requests from scrapers continues to increase, the current strategy outlined above seems to be fruitful. By reducing the use (and sometimes the availability altogether) of especially resource-intensive and &#8211; depending on the context &#8211; little useful functionalities, much stability and can be gained.</p>



<p class="wp-block-paragraph">The update seems to finally be largely completed (aside from maybe some further fine-tuning of the PHP-FPM configuration) and museum-digital is stable despite the bot problem, while we haven&#8217;t had to take more drastic or costly actions yet &#8211; such as blocking or adding additional servers.</p>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img loading="lazy" decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
					<wfw:commentRss>https://blog.museum-digital.org/2025/12/09/updates-ai-scrapers-and-resilience/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>State of Development, November 2025</title>
		<link>https://blog.museum-digital.org/2025/12/03/state-of-development-november-2025/</link>
					<comments>https://blog.museum-digital.org/2025/12/03/state-of-development-november-2025/#respond</comments>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Wed, 03 Dec 2025 01:53:58 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Importer]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[API]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[OAI-PMH]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4575</guid>

					<description><![CDATA[Frontend musdb Importer Core Parser]]></description>
										<content:encoded><![CDATA[
<h2 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/frontend/">Frontend</a></h2>



<ul class="wp-block-list">
<li>On source / reference pages, linked objects are now sorted by the position within the source work on which they are referenced or which they do themselves reference</li>



<li>The target URL of the regular / unspecified search bar for objects now follows the new, prettier URL schema</li>



<li>Support for an <a href="https://www.openarchives.org/pmh/">OAI-PMH</a> API for object metadata
<ul class="wp-block-list">
<li>Standardized endpoint for aggregators seeking to retrieve data in batch</li>



<li>Metadata formats thus far supported:
<ul class="wp-block-list">
<li>LIDO</li>



<li>OAI-DC (mandatory)</li>
</ul>
</li>



<li>See also: <a href="https://blog.museum-digital.org/2025/11/24/making-interoperability-easy/">Blog</a></li>
</ul>
</li>



<li>PDFs are only generated for users with a browser set to a non-default language if load on the server is low
<ul class="wp-block-list">
<li>The resource use caused by AI bots scraping museum-digital has been growing and growing. Generally, we see bots included in our mission to enable access to cultural heritage. On the other hand, nobody is served if the service is bogged down by bots. One functionality that is commonly used among bots and resource intensive is the generation of PDFs for object pages. The same information can be loaded from the object page itself and printed to a PDF using the browser&#8217;s print option. There are thus rather few downsides to limiting access to PDF generation to timmes, when server load is low. So that&#8217;s what we did.</li>
</ul>
</li>



<li>Collection-specific ISIL identifiers are now also used in the LIDO API</li>



<li>Alternative numbers of an object can now be displayed on object pages
<ul class="wp-block-list">
<li>This includes tooltips for types of alternative numbers, that can be set by the museum on the institution-wide settings pages of musdb</li>
</ul>
</li>
</ul>



<h2 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/musdb/">musdb</a></h2>



<ul class="wp-block-list">
<li>Search for objects
<ul class="wp-block-list">
<li>Type-ahead search for languages (of the object&#8217;s content)</li>



<li>Search by object&#8217;s revision status (open, read-ony, archived, etc.)</li>
</ul>
</li>



<li>Batch editing of objects&#8217; revision status</li>



<li>Parameters of the full text search index were updated to improve the search of word compounds in German</li>
</ul>



<h3 class="wp-block-heading">Importer</h3>



<h4 class="wp-block-heading">Core</h4>



<ul class="wp-block-list">
<li>The dry-run mode now does not abort an import anymore, if an unmapped value is encountered. Unmapped entries are collected and displayed together afterwards.
<ul class="wp-block-list">
<li>This means, that unmapped entries can now much more easily be copied to and mapped in <a href="https://concordance.museum-digital.org/">concordance.museum-digital.org</a></li>
</ul>
</li>



<li>Support for the import of alternative numbers (of objects)</li>



<li>Support for the import of space hierarchies</li>
</ul>



<h4 class="wp-block-heading">Parser</h4>



<ul class="wp-block-list">
<li><code>AdlibXml</code>
<ul class="wp-block-list">
<li>Support for importing objects&#8217; alternative numbers</li>
</ul>
</li>



<li><code>CsvXml</code>
<ul class="wp-block-list">
<li>Support for importing objects&#8217; alternative numbers</li>
</ul>
</li>



<li><code>CsvLocations</code>
<ul class="wp-block-list">
<li>New parser for csv-based imports of space hierarchies</li>
</ul>
</li>



<li><code>ImageByInvno</code>
<ul class="wp-block-list">
<li>New setting: append_chars (Adds suffixes, that exist in the inventory number, but not in file names)</li>
</ul>
</li>
</ul>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img loading="lazy" decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
					<wfw:commentRss>https://blog.museum-digital.org/2025/12/03/state-of-development-november-2025/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>State of Development, October 2025</title>
		<link>https://blog.museum-digital.org/2025/11/25/state-of-development-october-2025/</link>
					<comments>https://blog.museum-digital.org/2025/11/25/state-of-development-october-2025/#respond</comments>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Tue, 25 Nov 2025 16:55:09 +0000</pubDate>
				<category><![CDATA[Community]]></category>
		<category><![CDATA[Development]]></category>
		<category><![CDATA[Dissemination]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[Presentations]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[Object search (frontend)]]></category>
		<category><![CDATA[User interface]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4564</guid>

					<description><![CDATA[A summary of recent updates and development around museum-digital in October 2025.]]></description>
										<content:encoded><![CDATA[
<h2 class="wp-block-heading">Development</h2>



<h3 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/frontend/">Frontend</a></h3>



<ul class="wp-block-list">
<li>Significantly reworked the display of transcriptions on object pages
<ul class="wp-block-list">
<li>Titles of transcriptions are now displayed
<ul class="wp-block-list">
<li>If none is set, the type of the transcription (original or translation) is used as a replacement</li>
</ul>
</li>



<li>Transcriptions are sorted by their titles</li>



<li>Improved the display of transcriptions in tiles
<ul class="wp-block-list">
<li>Problems with vertical scrolling are now solved</li>



<li>If only one transcription has been recorded, it will be displayed on the full width of the page</li>



<li>If there are more than two transcriptions for an object, they are folded in by default and can be unfolded later on</li>
</ul>
</li>
</ul>
</li>



<li>Batch export of object metadata via the API
<ul class="wp-block-list">
<li>Thus far available in JSON &amp; LIDO</li>



<li><a href="https://nat.museum-digital.de/swagger/#/object/jsonExportObjects">API documentation</a></li>



<li><a href="https://blog.museum-digital.org/2025/11/24/making-interoperability-easy/">See also</a></li>
</ul>
</li>



<li>Dots as a separator in floating point numbers for object measurements are replaced with a comma in languages that require that</li>



<li>Collection-specific ISIL IDs are used in the LIDO API</li>
</ul>



<h3 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/musdb/">musdb</a></h3>



<ul class="wp-block-list">
<li>Added a field for recording titles / names of transcriptions</li>



<li>Added the option to set collection-specific ISIL IDs</li>



<li>Setting object type tags via the improvement suggestions now correctly classifies the thus created link between object and tag</li>



<li>Additional shapes are now available
<ul class="wp-block-list">
<li>E.g.: round, square</li>
</ul>
</li>



<li>Object groups can now be filtered by whether they have a superordinate one or not</li>
</ul>



<h3 class="wp-block-heading">Dissemination</h3>



<ul class="wp-block-list">
<li>2025-10-08: <a href="https://www.jrenslin.de/talks/interoperabilitaet-schaffen-geschichten-aus-1001-importen-herbsttagung/">Presentation</a> at the Autumn Conference of the Working Group Documentation of the German Museum Association: &#8220;Interoperabilität schaffen &#8211; Geschichten aus 1001 Importen&#8221;
<ul class="wp-block-list">
<li><a href="https://files.museum-digital.org/de/Praesentationen/2025-10-08_1001-Importe_Herbsttagung-FG-Doku_JRE.pdf">PDF</a></li>



<li><a href="https://files.museum-digital.org/de/Praesentationen/2025-10-08_1001-Importe_Herbsttagung-FG-Doku_JRE.odp">ODP</a></li>
</ul>
</li>



<li>2025-10-14: <a href="https://www.jrenslin.de/talks/civers-2025/">Talk</a> on a workshop of the project <a href="https://www.dainst.org/forschung/projekte/citation-of-versioned-web-pages-by-pid-civers/5926">CiVers (Citation of Versioned Web Pages by PID)</a>
<ul class="wp-block-list">
<li><a href="https://files.museum-digital.org/de/Praesentationen/2025-10-14_museum-digital_Civers_JRE.pdf">PDF</a></li>



<li><a href="https://files.museum-digital.org/de/Praesentationen/2025-10-14_museum-digital_Civers_JRE.odp">ODP</a></li>
</ul>
</li>



<li>2025-10-17: <a href="https://verein.museum-digital.de/museum-digital-usertagung-2025/">museum-digital Usertagung 2025</a></li>
</ul>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img loading="lazy" decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
					<wfw:commentRss>https://blog.museum-digital.org/2025/11/25/state-of-development-october-2025/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
