The Ghost of Artworks and the Memory of Algorithms
The history of a work of art never ends with the signature on the canvas. It continues through a trail of transactions, inheritances, auctions, and, at times, institutionalized theft. During the Third Reich, the Nazi regime organized the systematic looting of tens of thousands of items belonging to Jewish families. Returning these objects to their rightful owners or their descendants requires a titanic effort of documentation. This discipline—provenance research—often resembles an archaeological dig through fragmented archives. At Santa Clara University, a group of researchers is experimenting with a radically different method to accelerate these restitutions, using artificial intelligence to decipher the silences of history.
The starting point for this initiative is not a computer science lab, but the courtroom. Michael Santoro, a professor of management at the California university, closely followed the journey of his friend Timothy Reif. A judge on the U.S. Court of International Trade, Reif had embarked on a long battle to recover the heritage stolen from his ancestors. Although the judge ultimately prevailed, securing the return of artifacts estimated to be worth several million dollars, the victory left a bitter taste. A case that seemed straightforward on paper became bogged down in years of complex legal disputes. Faced with this administrative and judicial sluggishness, Michael Santoro wondered whether technology could break through the deadlocks that discourage so many families.
Turning Everyday Language into Historical Evidence
To turn this idea into a tangible tool, the management professor turned to his colleagues specializing in information systems and data analysis, Haibing Lu and Michele Samorani. Together, they developed a specialized chatbot, officially unveiled earlier this month and named the AI Provenance Assistant. The specifications for this interface are based on a clear goal: to break down the barrier of technical expertise.
The system allows users to formulate their queries in everyday English. It is no longer necessary to master the jargon of art historians or the complex syntax of institutional databases. A family searching for traces of a looted artwork can query the interface by combining a surname, an artistic medium, or a few descriptive details. The algorithm interprets this query and generates a list of matching works. The results display the titles, identified artists, and all information associated with the piece throughout its history.
The Challenge of Parisian Archives and Corrupted Data
The effectiveness of artificial intelligence depends entirely on the data it is fed. The team had to grapple with the reality of the existing records. The researchers settled on an extremely well-documented corpus: the database cataloging the cultural looting carried out by the Einsatzstab Reichsleiter Rosenberg, compiled by the Jeu de Paume in Paris. This registry centralizes information on 40,000 looted works.
Integrating this massive dataset revealed the inherent complexity of the art market during wartime. Michael Santoro explains that the descriptive records for the works are rarely straightforward or perfectly accurate. The trio discovered that the recorded information contains numerous gaps, approximations, and even deliberate falsifications entered at the time to cover their tracks. To overcome these obstacles and make the information usable, Haibing Lu conducted extensive research, downloading the entire Paris registry in order to classify and structure it for machine learning.
Academic Caution in the Face of Digital Illusions
The use of algorithms to process looted art archives has sparked keen interest among specialists in the field, while also raising methodological questions. Carla Shapreau studies these mechanisms in detail. A lawyer and professor of law at UC Berkeley, she focuses her provenance research on a specialized field: musical instruments, manuscripts, and documents related to music. While the prospect of using artificial intelligence to accelerate data cross-referencing fills her with genuine optimism, she expresses reservations about the tool’s autonomy.
The inherent risk of current artificial intelligence systems lies in their ability to produce erroneous associations or factual hallucinations. For the law professor, no result generated by the AI Provenance Assistant can replace human validation. Every lead suggested by the program requires rigorous verification. To strengthen the project’s credibility over time, she advocates for a major technical advancement: a direct link between the algorithm’s results and digitized versions of primary sources. This direct link to historical documents would allow researchers and legal professionals to examine the visual evidence without blindly relying on computer-generated summaries.
The Financial Equation of an Unfinished Quest
The program’s developers share this vision and view their current interface as a foundation that can be improved. Haibing Lu expresses a desire to involve researchers in the tool’s development, in order to calibrate the algorithms to their actual investigative methods and day-to-day needs. The goal is to refine the accuracy of responses to increasingly complex queries.
This ambition, however, faces a practical reality. The development of the AI Provenance Assistant currently relies exclusively on the personal commitment and personal funds of its three creators. The project lacks the financial support needed to scale up. An injection of capital would make it possible to expand the robot’s search capabilities by connecting it to other international databases, thereby increasing the chances of locating works scattered across the globe. Technical feasibility has been proven; only the funding is lacking to deploy the necessary digital architecture.
Behind the lines of code and data analysis, the driving force behind the project remains deeply human. The financial investment and hours of programming are not measured by conventional profitability standards. For Michael Santoro, the success of this endeavor will not be measured by the volume of requests processed. The identification and return of a single piece taken from a family would, on its own, justify all the efforts made by the California-based team.


