<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Artificial Intelligence | museum-digital: blog</title>
	<atom:link href="https://blog.museum-digital.org/tag/artificial-intelligence/feed/" rel="self" type="application/rss+xml" />
	<link>https://blog.museum-digital.org</link>
	<description>A blog on museum-digital and the broader digitization of museum work.</description>
	<lastBuildDate>Thu, 01 Oct 2026 12:38:19 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	

<image>
	<url>https://blog.museum-digital.org/wp-content/uploads/2020/01/cropped-mdlogo-code-512px-32x32.png</url>
	<title>Artificial Intelligence | museum-digital: blog</title>
	<link>https://blog.museum-digital.org</link>
	<width>32</width>
	<height>32</height>
</image> 
<atom:link rel="search" type="application/opensearchdescription+xml" title="Search museum-digital: blog" href="https://blog.museum-digital.org/wp-json/opensearch/1.1/document" />
	<item>
		<title>Braving the Waves. Surviving in Times of AI.</title>
		<link>https://blog.museum-digital.org/2026/10/01/braving-the-waves-surviving-in-times-of-ai/</link>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 12:30:59 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Infrastructure]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Bots]]></category>
		<category><![CDATA[Performance]]></category>
		<category><![CDATA[Web Development]]></category>
		<category><![CDATA[Web Hosting]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4755</guid>

					<description><![CDATA[Now that the preview release of musdb&#8217;s new UI has finally been released, there&#8217;s time to write about the other sides of museum-digital. More usual struggles. Which is to say: New, larger waves of AI scrapers and bots that once again started to impair our services starting July this year. We maintain that hypocrisy is <a href="https://blog.museum-digital.org/2026/10/01/braving-the-waves-surviving-in-times-of-ai/" class="more-link">...</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Now that the <a href="https://blog.museum-digital.org/2026/09/30/pre-release-of-musdbs-new-ui/">preview release of musdb&#8217;s new UI has finally been released</a>, there&#8217;s time to write about the other sides of museum-digital. More usual struggles. Which is to say: New, larger waves of AI scrapers and bots that once again started to impair our services starting July this year.</p>



<p class="wp-block-paragraph">We maintain that hypocrisy is not for us. I personally have been involved in discussions about the <em>FAIR</em> principles from 2015 onwards. And in 2015 they were already running under that label for a while. As a formalized document, they&#8217;ve been around since 2016, but disregarding the label, the principles and surrounding discussions go back, I guess, at least to the 1990s.</p>



<p class="wp-block-paragraph">In the cultural sector, with many institutions being primarily tax-payer funded, we want to provide our data and services in a <em>F</em>air, <em>A</em>ccessible, <em>I</em>nteroperable, and <em>R</em>eusable manner. Relevant here are primarily <em>accessible</em>, <em>interoperable</em>, and _reusable</p>



<p class="wp-block-paragraph">Realistically, accessibility refers to two things: Immediate, universal barriers like a requirement for authentication should be avoided where possible. On the other hand, user-specific barriers should be reduced as much as possible as well: Human-readable pages should e.g. be designed with sufficient contrast.</p>



<p class="wp-block-paragraph">If one sees other services and machines as users, which may be reductionist but is logically much more coherent than doing otherwise, then <em>interoperable</em> refers simply to a special form of accessibility: As data should be provided in a way that as many human users can use as possible, it should also be served in formats that machines can read. Ideally in ways that machines can read without further adjustments &#8211; following open standards.</p>



<p class="wp-block-paragraph"><em>Reusable</em> then builds upon these. As the data is now freely accessible by others, setting appropriate licenses allows for the creative reuse of the data.</p>



<p class="wp-block-paragraph">Ten years after the FAIR principles were formalized, we finally have someone who reuses our data on a massive scale. We should rejoice. Unfortunately all our previous conceptions of how that reuse might take place turned out to be wrong.</p>



<h2 class="wp-block-heading">Celebrating AI Scrapers?</h2>



<p class="wp-block-paragraph">AI scrapers flooding web services, and especially larger, well-established infrastructures with requests has been an ongoing story for some two years now. Last year museum-digital was met with a first large wave of AI scrapers. See the previous blog posts (<a href="https://blog.museum-digital.org/2025/12/09/updates-ai-scrapers-and-resilience/">1</a>, <a href="https://blog.museum-digital.org/2025/12/22/cleaning-out-our-closet/">2</a>, <a href="https://blog.museum-digital.org/2025/12/29/trimming/">3</a>).</p>



<p class="wp-block-paragraph">On the one hand, using our data for training AI models is undoubtedly reuse. It is even productive reuse: It&#8217;s a small contribution to the training of technology that, at its current quality, sounded like science fiction 10 years ago.</p>



<p class="wp-block-paragraph">On the other hand, this type of reuse takes place without attribution. AI scrapers work based on scale; Machine-readable APIs are uninteresting to scrapers. They suck up as much data as possible, as quickly as possible. Any page-specific adjustments, such as finding and using the API, would reduce the speed of scraping. The requests are so numerous, that they routinely crash servers. Last but not least, those prominently benefitting from the AI hype are far from sympathetic people.</p>



<p class="wp-block-paragraph"><strong>But principles are principles.</strong></p>



<p class="wp-block-paragraph">It, again, cannot be denied that what AI scraping is eventually aimed at is reuse. Even though they use HTML rather than the APIs we lovingly carved for them, they interoperate with our services in the way most accessible to them.</p>



<p class="wp-block-paragraph">The problem is, that they are not playing fair. And if the previously unlikely number of requests leads to servers crashing and shutting down, then all principles were for naught. A service that is providing no data at all cannot provide them _FAIR_ly either.</p>



<p class="wp-block-paragraph">At museum-digital we have seen the new requests as an opportunity to improve our software. If it can withstand the onslaught of thousands of bots a second, it is confirmedly battle-tested and stable. Improving our publishing software on the other hand also meant cutting out overly resource-hungry functionalities. Last year we dropped most publicly available, server-side PDF generation capabilities, we moved the IIIF APIs to an separate configuration that is specced to be able to fall over without impairing the rest of our services. We improved the parsing of search terms and introduced rate-limiting (how many requests a single IP can perform for a given time).</p>



<p class="wp-block-paragraph">The new wave of AI scrapers since July is even bigger than last year&#8217;s. Those improvements were not sufficient to keep our primary server, which hosts the production databases, stable.</p>



<h2 class="wp-block-heading">From Each According to Their Ability: Load Shedding</h2>



<p class="wp-block-paragraph">We continued on last year&#8217;s course however, analyzing the scrapers&#8217; requests, their influence on server stability (which is to say, which requests were especially resource-intensive) and adjusted our setup accordingly. Note, that we did <strong>not</strong> scale up: museum-digital runs on exactly the same bare-metal servers today as it did two years ago.</p>



<p class="wp-block-paragraph">In analysis, we identified two types of especially resource-intensive browsing behaviors that are almost entirely limited to scrapers (as well as some very engaged researches, to whom we express our regret for now regularly banning them):</p>



<ul class="wp-block-list">
<li>Scrapers follow links on search pages, especially the facet search, and end up combining more and more search parameters. It is unlikely that a human would perform a search for objects that &#8220;Are related to Berlin, and related to Germany, and related to Institution X, but not related to Institution Y, and also related to a time between 1990 and 2000.&#8221; For an automated scraper it is normal behavior.</li>



<li>If a published object page has been updated, a snapshot of its current state is automatically saved to provide an archived version of the page at a given time. This is thought mainly for researches to be able to cite a given state of the page. In everyday operation, archive pages should be barely used. They are also delisted from search machines. But saving snapshots again and again leads to the existence of a large number of links and separate pages scrapers can scrape to no benefit to anybody.</li>
</ul>



<p class="wp-block-paragraph">Both functionalities are legitimately useful but scale badly.</p>



<p class="wp-block-paragraph">We hence introduced load shedding: Whenever a search is performed or an archive page is accessed, the current server load is evaluated. If it is above a certain threshold, the server does not perform complicated search queries (depending on how high load is, it is restricted to a minimum of three search parameters) or does not load the archive page. Instead a warning is presented, that the given action is currently unavailable and a custom HTTP error code is sent. If the same IP causes that error code to be sent twice, the IP is blocked for a while.</p>



<p class="wp-block-paragraph">As of the time of writing writing, we have thus blocked a total of 131,014,530 IPs in two month, with 5,802,771 being currently blocked.</p>



<p class="wp-block-paragraph">This approach follows our principles: We try to be fair with everybody. But if users are not fair and do not follow advise, we don&#8217;t feel obliged to continue playing fair either.<br>Unfortunately this approach works on the level of blocking whole IPs. As &#8220;residential proxies&#8221; &#8211; back in the days we called them botnets infecting home routers &#8211; have become an ever larger problem, it is not unlikely that actually interested, normal users also unknowingly host an AI scraper at home. I fear I&#8217;ve already encountered one such case, were a colleague could not access museum-digital for seemingly no good reason from her home network. Otherwise, it is surprisingly effective &#8211; the number of requests has barely been reduced, but our services are essentially stable.</p>



<p class="wp-block-paragraph">The only crash since happened the day before yesterday and was caused by the combination of bots and an error in our controlled vocabularies &#8211; the latter of which is entirely our fault.</p>



<h2 class="wp-block-heading">Numbers Games: Logging Strategy</h2>



<p class="wp-block-paragraph">After the above-mentioned actions had effectively returned stability to our publicly accessible portals, one issue remained for internal services: File uploads were incredibly slow. It turned out that the large number of requests being permanently logged overburdened the SSD. We have hence disabled the general access logging, restricting ourselves to logging requests that either cause errors or trigger slow responses. This reduces the ability to react well-informed to new waves of requests and actual attacks. Thankfully those also commonly trigger actual errors we can still identify.</p>



<h2 class="wp-block-heading">Fairness, FAIRness</h2>



<p class="wp-block-paragraph">With these actions we have again averted the need to use more drastic strategies that would limit accessibility to users &#8211; human and machine alike &#8211; like blanket bans on whole regions of the world or the installation of software such as <a href="https://github.com/TecharoHQ/anubis/">Anubis</a>. And, again, we have not yet needed to improve our hardware (and pay the price for that), even though that might be in order eventually. It might be wise for other reasons as well.</p>



<p class="wp-block-paragraph">A holistic picture of one&#8217;s setup &#8211; software, hardware, systems administration &#8211; and the capability to think and act upon these together helps a lot.<br>But, if anything, hosting in a <em>FAIR</em>, principled way still works in these trying times.</p>



<p class="wp-block-paragraph"><em>(Post image generated using Anima Base v1)</em></p>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>State of Dev, June &#038; July 2025</title>
		<link>https://blog.museum-digital.org/2025/08/23/state-of-dev-june-july-2025/</link>
					<comments>https://blog.museum-digital.org/2025/08/23/state-of-dev-june-july-2025/#respond</comments>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Sat, 23 Aug 2025 21:12:41 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Importer]]></category>
		<category><![CDATA[musdb]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Minor Improvements]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[Object search (musdb)]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4519</guid>

					<description><![CDATA[June and especially July were at first glance once again rather slow months in terms of development at museum-digital. Generally, the pace and type of development seems to have changed this year. Rather than doing many small improvements all over the place, there is less but larger and more labor intensive changes and new features. <a href="https://blog.museum-digital.org/2025/08/23/state-of-dev-june-july-2025/" class="more-link">...</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">June and especially July were at first glance once again rather slow months in terms of development at museum-digital.</p>



<p class="wp-block-paragraph">Generally, the pace and type of development seems to have changed this year. Rather than doing many small improvements all over the place, there is less but larger and more labor intensive changes and new features. See for example the <a href="https://blog.museum-digital.org/2025/01/13/version-control-batch-transfer-between-data-fields-of-object-records/">versioning</a> in musdb (January), the tool for <a href="https://blog.museum-digital.org/de/2025/03/08/das-importieren-automatisieren/">automating imports</a> based on what others call a &#8220;hot folder&#8221; and the sort option to <a href="https://blog.museum-digital.org/2025/03/06/sort-by-beauty/">sort objects by their images&#8217; aesthetics score</a> (both presented here in March), and the new tool for suggesting formulations for object descriptions using large language models (June).</p>



<h2 class="wp-block-heading">July</h2>



<h3 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/frontend/">Frontend</a></h3>



<ul class="wp-block-list">
<li>Translation of the software to <a href="https://blog.museum-digital.org/2025/07/13/hindi/">Hindi</a> and <a href="https://blog.museum-digital.org/2025/07/02/browse-museum-digital-in-telugu/">Telugu</a></li>



<li>Grouping of tags by their relation to a given object<br><em>If an object is linked to more than 10 tags, the tag list of object pages quickly starts looking unorganized and messy. In such cases, the tags will now be displayed grouped by their relationship to the object (thus far: general tag, material, technique, object type, display subject)</em></li>
</ul>



<h3 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/musdb/">musdb</a></h3>



<h4 class="wp-block-heading">New Features</h4>



<ul class="wp-block-list">
<li>Export option for the specific LIDO as expected by the <a href="http://Koloniale Kontexte-Portal der Deutschen Digitalen Bibliothek">German Digital Library&#8217;s &#8220;colonial contexts&#8221; portal</a></li>
</ul>



<h4 class="wp-block-heading">Improvements &amp; Changes</h4>



<ul class="wp-block-list">
<li>The minimum length of fulltext search terms for object search parameters is now visibly enforced in the extended search user interface<br><em>To not overly burden the fulltext search server, any term in a full text search in musdb needs to be a minimum of two characters long. Thus far, shorter full text search parameters were simply ignored. This certainly was confusing at times. Since June, attempting to perform search queries with shorter search parameters is made impossible by a check in the extended search overlay.</em></li>



<li>Reception history of objects: Statements of the relevant position within a source can now be 40 characters long</li>



<li>Transcriptions
<ul class="wp-block-list">
<li>May now be up to 4000000 long</li>



<li>New fields: Notes on the transcription, status, aims</li>
</ul>
</li>
</ul>



<h4 class="wp-block-heading">Bugfixes</h4>



<ul class="wp-block-list">
<li>Fixed a bug in the batch editing of specific visibility flags for data fields on the addendum tab</li>
</ul>



<h2 class="wp-block-heading">Juni</h2>



<h3 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/frontend/">Frontend</a></h3>



<ul class="wp-block-list">
<li>Performance improvements
<ul class="wp-block-list">
<li>Object search now runs without a connection to the full text search server if no full text search parameter is set</li>



<li>If multiple search parameters for an earliest / latest time have been set, they are parsed and combined before being forwarded to the database</li>
</ul>
</li>



<li>Improvements in the deletion of temporarily created PDF files (PDF export)</li>



<li>Navigation has been translated to <a href="https://blog.museum-digital.org/de/2025/06/23/tamil/">Tamil</a></li>
</ul>



<h3 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/musdb/">musdb</a></h3>



<h4 class="wp-block-heading">New Features</h4>



<ul class="wp-block-list">
<li>Recipient of deaccessed objects can now be linked from within the address book</li>



<li><a href="https://blog.museum-digital.org/de/2025/06/19/ki-objektbeschreibungen/">New tool for AI-aided formulation o object descriptions (based on existing other metadata)</a></li>
</ul>



<h4 class="wp-block-heading">Improvements</h4>



<ul class="wp-block-list">
<li>Object search now runs without a connection to the full text search server if no full text search parameter is set</li>
</ul>



<h3 class="wp-block-heading"><a href="https://blog.museum-digital.org/category/development/importer-en-en/">Importer</a></h3>



<ul class="wp-block-list">
<li>The CSVXML parser has been extended to cover new event types and markings</li>



<li>Object groups automatically generated to group all objects of an import can now receive a description as set within the import configuration</li>
</ul>



<h3 class="wp-block-heading"><a href="https://en.about.museum-digital.org/software/nodac/">nodac</a></h3>



<p class="wp-block-paragraph">The list of selectable languages for the navigation of nodac has now been restricted to those in which there is actually a complete translation</p>
]]></content:encoded>
					
					<wfw:commentRss>https://blog.museum-digital.org/2025/08/23/state-of-dev-june-july-2025/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Sort by Beauty</title>
		<link>https://blog.museum-digital.org/2025/03/06/sort-by-beauty/</link>
		
		<dc:creator><![CDATA[Joshua Ramon Enslin]]></dc:creator>
		<pubDate>Wed, 05 Mar 2025 23:25:18 +0000</pubDate>
				<category><![CDATA[Development]]></category>
		<category><![CDATA[Frontend]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[New Features]]></category>
		<category><![CDATA[Object search (frontend)]]></category>
		<category><![CDATA[Search]]></category>
		<guid isPermaLink="false">https://blog.museum-digital.org/?p=4333</guid>

					<description><![CDATA[Last month a new sort option appeared on museum-digital: "Aesthetics prediction". Thoughts on AI, beauty, and the discriminating nature of sorting.]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"><a href="https://blog.museum-digital.org/2025/02/14/state-of-dev-december-2024-january-2025/">Last month</a> a new sort option appeared on museum-digital: &#8220;Aesthetics prediction&#8221;. Based on the <a href="https://github.com/LAION-AI/aesthetic-predictor">LAION aesthetics predictor</a>, each published object&#8217;s main thumbnail aesthetics are scored. The objects can then be sorted according to this score.</p>



<h2 class="wp-block-heading">AI, Aesthetics and Discrimination</h2>



<p class="wp-block-paragraph">When working with AI it is common to criticize the inherent biases of the models used. Already underrepresented entities (people, viewpoints, etc.) are excluded, because they are not sufficiently represented in the data the model was trained on. In turn, AI repeats what it learned &#8211; pre-existing biases in society are replayed and reinforced by AI.</p>



<p class="wp-block-paragraph">A second, well-founded criticism against AI applications is their impreciseness, their tendency to &#8220;hallucinate&#8221; entirely wrong results, and the lack of reproducibility of results.</p>



<p class="wp-block-paragraph">Both criticisms are entirely valid. And both hint at why AI might be a useful tool exactly for generating a sort order by aesthetics. Like AI, aesthetics are imprecise: Ask 10 people to rank some pictures by their aesthetic appeal and you will get 10 different answers. But the general thrust of the answers will likely be more or less the same. There are societal rules &#8211; biases &#8211; for what pictures are more or less beautiful. But they are fuzzy and interpreted differently by each and every observer. A sense of aesthetics is about learning and reproducing those biases &#8211; learned by the inputs we learn in everyday life -, just as AI will be biased based on the data it is trained on.</p>



<p class="wp-block-paragraph">Sorting then is, in essence, an exercise of discrimination. Sorting entries by ID / age in a database system will favor the newest entries and discrimate against older ones. Sorting entries alphabetically in an ascending order will favor entries whose title starts with &#8220;A&#8221; while it will discrimate against those whose title starts with &#8220;Z&#8221;. Sorting based on an AI-generated rank is discriminating. Sorting by <em>beauty</em> or <em>aesthetics</em> is as well. AI allows sorting by beauty.</p>



<h3 class="wp-block-heading">What is discriminated against?</h3>



<p class="wp-block-paragraph">Now, if everything in sorting is discriminating, it is all the more important to consider what is actually evaluated. In other words, on what basis the discrimination takes place when ranking digitized museum objects by the beauty of their digital reproductions. As stated above, the objects are scored based on the aesthetics score established for their main thumbnails.</p>



<p class="wp-block-paragraph">Images then have a range of aspects that influence their aesthetics. Image composition, lighting, contrast, the motive, and many more (an art historian could likely list hundreds). For the common critique, it is essentially only the motive that matters: Favoring images of <em>white</em> women over images of Asian men reproduces a range of societal problems that should not be reproduced. Transferred to object photography, the motive may be further differenciated between object type and the actual motive (e.g. of a painting). And such a discrimination is noticable in the results: Paintings seem to be generally slightly better ranked than pictures of tools. A bias based on the displayed subjects of e.g. different paintings is not something anybody from the team would have noticed, but is to be assumed that one will exists.</p>



<p class="wp-block-paragraph">On the other hand, the influence of motive-focused biases is many times weaker than the actually useful discrimination based on the technical aspects of the images. Objects images taken by a professional photographer with up-to-date equipment are ranked much better than images taken using a digital camera from the 1990s. Images without a timestamp or a watermark are ranked better than ones that do feature one. Images taken with proper lighting and contrast settings are ranked much better than images taken in a dark room, presenting the objects in gray on black. Similarly, one of the most important facets contributing to the aesthetics score the predictor returns seems to be the composition: If an object is centered in the photo taken, it will score way above images featuring multiple objects at different corners of the image (a classic example would be images that show both the actual object and a photo bar).</p>



<p class="wp-block-paragraph">To reiterate, these technical aspects are visibly the main contributors to the score. And discriminating based on those images is actually useful: It allows museum-digital to present newer users who are just browsing the published collections with objects recorded with a more consistently high visual quality.</p>



<h3 class="wp-block-heading">What are the alternatives, <em>or</em> who is the audience?</h3>



<p class="wp-block-paragraph">museum-digital suffers the old problem of a lack of a specified target audience. A common user may be a hobbyist looking for other versions of the model train they just bought. They might be a person interested in what art of the 16th century looked like. Or which museum to visit next. Or they might be a specialist, with a much better idea of what they are actually looking for. It is common for users to leave the page after less than a minute. But there is also an astonishing number of users who stay for hours. Accordingly, it is all the more important to offer (sort) options catering to different needs.</p>



<p class="wp-block-paragraph">A new user who chanced upon the platform and randomly browses it with no further background on museums will benefit from the new sorting method. Sort orders like sorting by ID or title are linear and follow a consistent logic &#8211; but that logic may in essence just reflect its own type of randomness. Sorting newer objects over older ones is essentially a random sort order, if the museum does &#8211; as is usual &#8211; record objects as they are needed in an exhibition, newly enter the museum or are simply the next object in the shelf &#8211; all of which follow a very particular, not publicly comprehensible logic. Similarly, sorting objects by their names or titles is linear. In contrast to the object entry&#8217;s age in the database, it is also immediately comprehensible to users. But museums are free to determine the object name freely &#8211; and often this is necessary. As per the best practices around publication on museum-digital, an object name should be descriptive and usable to distinguish between different objects. As most objects simply are nameless by themselves, this often means that a colleague at the responsible museum <em>invented</em> a name. Seeing a green vase, there may be a rather short list of names that come to mind for most people &#8211; &#8220;green vase&#8221;, &#8220;vase, green&#8221;, &#8220;vase&#8221;, &#8220;green-ish vase&#8221;, &#8220;vase in pale green&#8221; -, making the selection of the actual name comprehensible. But &#8220;green vase&#8221; is in a very different position from &#8220;vase, green&#8221; when sorting objects alphabetically. One might actually be involuntarily be sorting the objects by curator.</p>



<p class="wp-block-paragraph">Other than an object&#8217;s title/name and institution, the one other facet users see when listing objects on an overview or search page is the object&#8217;s thumbnail. And as a sense of aesthetics &#8211; fuzzy as it certainly is &#8211; is roughly shared among most people (with globalization, even actually <em>most</em> people), sorting objects of diverse origins by the aesthetics of their thumbnails suddenly starts to look like a comparatively relatable and reasonable sort order. Nobody will agree 100%, but the rough order instinctively makes sense. That&#8217;s more than can be said about the alternatives, unless one actually looks at the sort settings.</p>



<p class="wp-block-paragraph">Presenting new users with a more consistently appealing and streamlined set of search results (at the first glance at least) on the other hand might encourage them to stay for longer. And if they stay longer, they might end up actually chancing upon more objects &#8211; including ones with less well-taken pictures &#8211; eventually.</p>



<p class="wp-block-paragraph">For those users who are researchers and/or specialists, who need a linear, logical sorting, the aesthetics prediction sort option is obviously of much less merit. But it is safe to assume that is this group of users are overrepresented in the above-mentioned group of people who browse the site for hours. And it is also rather safe to assume, that they are generally more used to online databases and the existence of different sort options. Consequently, it can be assumed that they are able to change the sort settings to their needs &#8211; or at least have a higher likelihood of being able to do so.</p>



<p class="wp-block-paragraph">Being very useful for the general public while less so for specialists who can be assumed to be more skilled anyway, it is sensible to make aesthetics-based sorting the default. There are thus three different default sort orders, depending on what objects one lists:</p>



<ul class="wp-block-list">
<li>If a user lists all objects of a given instance, the default sort order remains sorting by the age of the entries</li>



<li>If a fulltext search was performed, the results are sorted by how well the query string matched the entry by default</li>



<li>Otherwise &#8211; in case of ID-based search queries like a query for all objects created in Berlin &#8211; the search results are sorted using the new aesthetics prediction by default</li>
</ul>



<h2 class="wp-block-heading">Operationalizing the Prediction</h2>



<p class="wp-block-paragraph">On a side note, operationalizing the prediction proved a challenge in itself. All of museum-digital&#8217;s servers are built for traditional web hosting. Hence, they feature quite powerful CPUs, lots of RAM and no GPU. In other words: They are least suitable for AI applications. And they are all the more unsuitable for scoring over a million thumbnails in bulk while remaining otherwise performant.</p>



<p class="wp-block-paragraph">If the existing servers cannot be used, there are three reasonable alternatives. First, we could have simply rented another server. As a mostly volunteer-run project, this was not an option. Second, we could have used browser-based AI to calculate the score on users&#8217; machines &#8211; we might have e.g. let an uploading user&#8217;s machine calculate the score of a thumbnail whenever the user uploads an image. But the users&#8217; PCs vary widely, laptops and tablets (usually again without GPUs) have become more and more popular when compared to workstation, and thus such an approach would have made uploading images unbearably slow. It would have also not helped with calculating scores for the already pre-existing million of objects and their main thumbnails. Again, this was not really an option.</p>



<p class="wp-block-paragraph">Finally, we could use our private machines (specifically: mine). And that is the approach we chose. Based on a new search API parameter aesthetics_score for querying objects whose thumbnails have not been scored yet (aesthetics_score:10001), my PC downloads the yet unevaluated object thumbnails and evaluates a score for each. These scores are then locally stored in one SQLite database per instance of museum-digital and exported into a CSV file. The CSV file is then uploaded to the server and ingested back into the database to set the aesthetics score for the relevant images. The search index is automatically updated with the new scores following the upload.</p>



<p class="wp-block-paragraph">This structure unfortunately also means that the scoring does not occur in real-time. Depending on when I turn on or shut down my PC, the most recently published objects may remain uncategorized for a while. To be able to fairly represent such objects, they are assigned an impossibly high score that serves to both sort them above the already-scored objects while also marking them as yet unscored. Specifically, the aesthetics predictor returns a score between 0 and 10. As calculating with full integers is generally much more simple than working with floating point numbers, the score is multiplied by a thousand and then rounded, leaving one with a score between 0 and 10000. An object yet unscored will hence be assigned a default score of 10001.</p>



<div class="wp-block-cgb-cc-by message-body" style="background-color:white;color:black"><img decoding="async" src="https://blog.museum-digital.org/wp-content/plugins/creative-commons/includes/images/by.png" alt="CC" width="88" height="31"/><p><span class="cc-cgb-name">This content</span> is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International license.</a> <span class="cc-cgb-text"></span></p></div>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
