Digital resources in the Social Sciences and Humanities OpenEdition Our platforms OpenEdition Books OpenEdition Journals Hypotheses Calenda Libraries OpenEdition Freemium Follow us

Designing an Open-Data Workflow for the Humanities: Notes from a Hands-On Workshop

When the paper ‘Towards a Contextualised Spatial-Diachronic History of Literature: Mapping Emotional Representations of the City and the Country in Polish Fiction from 1864 to 1939’, introduced the (meta)corpus creation workflow, I had the feeling that it could serve a greater and more general purpose than for only one type of resource. Since then, I have been trying to use it as a model for data-driven research within the humanities.

1. Curiosity as the origin of a research workflow

Designing a meaningful workflow for open data in the humanities often feels like navigating between two worlds. On one side, there is the deeply interpretive, context-driven nature of research in the humanities; on the other, the idea that you can explain the world of cultural phenomena with data, algorithms, tools, and code. During the workshop ‘Open Data in Art Research: Practices and Innovations in Museum Collaboration’, which took place in October 2025 and was held at the Polish Academy of Sciences’ Scientific Centre in Vienna, we set out to show that these two worlds are not in conflict. Instead, they can be incorporated into a coherent research process in which conceptual studies and data-based modelling reinforce each other. Thus, with a significant contribution provided by open science, we applied the workflow for data-driven research to art history.

The starting point was deliberately simple: curiosity. Our workflow does not begin with code or data cleaning, nor with the well-known wondering about ‘what digital humanities can I use in my work?’, but with questions. The research question should always be step no. 1. The rest is follow-up. In our case, the invitation came from the ArtEmis dataset, a large corpus of artworks accompanied by written descriptions of emotions provided by human annotators. It is a compelling resource because it captures something central to humanistic inquiry – subjective experience – and shapes it into a structured form. Confronted with this material, it is only natural to form questions that are interpretive before they are analytical: asking why certain paintings evoke specific emotional patterns, how viewers justify the feelings they express, and whether artistic styles can be characterised through the emotions they most commonly elicit. These questions arise as a result of our encounter with cultural artefacts; the role of open data is simply to make them suitable for analysis.

2. Metadata as Conceptual Architecture

From these research questions, we moved to the design of the metadata. This step, although unglamorous, is where the intellectual architecture of the project takes shape. The task was not to create a detailed catalogue of attributes, but to identify the elements necessary for presenting our interest and the multi-dimentional characteristics of the artworks. We had the research questions; now we needed a structure. In practice, this meant thinking carefully about what information is required for comparison and interpretation. Metadata is not a neutral container: it encodes theoretical commitments and sets the boundaries for what the workflow can reveal. Designing metadata is simply a scholarly act of modelling that forces us to articulate which categories matter – we focused on people, places, and paintings, and the relations between them.

The next step; creating the data collection was rather straightforward. We used the existing ArtEmis database, and connected it with the content of the wikiart.org service. From there we connected the data with Wikipedia and Wikidata. The workshop led us from spreadsheets containing information about museum collections, through to Linked Open Data and Google Colab notebooks. This experience highlighted a reality that is often overlooked in discussions about open science: cultural data is messy. Names are inconsistent, artworks appear in different forms, there is a lack of connections between data sources, licensing varies widely, and the use of persistent identifiers, if there are any, is usually inconsistent. Yet, this difficult situation is pedagogically valuable, because it reveals that datasets are not found objects but constructed research artefacts shaped by institutional histories as well as human interests and decisions. Data construction and data cleaning are scientific activities.

3. Persistent identifiers and knowledge graphs AKA How to increase the visibility of research and manage uncertainty?

The creators of the knowledge are those that are specialists in the subject. Let’s not shift the burden of clarity onto the recipients. To stabilise the message we need persistent identifiers (PIDs). Assigning VIAF IDs to artists in the ArtEmis database transformed a closed list of names into a node within a broader informational ecosystem. This step demonstrated that interoperability is not just a technical requirement but a scholarly responsibility that helps us to put our data into context and helps with the uncertainty. The semi-automated tool used during the workshop proposed matches for each artist’s name, but human verification afterwards was crucial. Authority control reveals connections across collections while also exposing the gaps, ambiguities, biases in open science data representation, and occasional errors. Persistent identifiers do not eliminate interpretive labour; they simply make it more transparent.

Once the artists and artworks were stabilised, the workflow shifted toward building a relational environment – an RDF-based knowledge graph. Converting the rows of a spreadsheet into triple structures that have subjects, predicates, and objects, opened a conceptual space that does not exist in flat tables. Instead of thinking in terms of columns and rows, participants began to think in entities and the relationships between them. Loading the data into a graph database and posing SPARQL queries made these relationships visible. Patterns that had been invisible emerged as constellations of networks involving art movements, art institutions, genres, personal and influential connections, painting schools, pupils, and teachers, as well as clusters of emotional language. At this point, the dataset ceased to be a static resource and became a dynamic landscape open for interpretive discoveries.

4. Verification and exploration – let’s not forget about the human-in-the-loop

There are no unbiased data. All datasets are the product of research activities and the interests of specific individuals, projects, and institutions. We needed to remain a bit skeptical. The moment at which the network appeared persuasive was the moment to question it. Participants examined automatically generated links, searching for incorrect matches, inconsistent metadata, or conceptual clusters that risked oversimplifications or anecdotal observations. This semi-manual verification stage played a key epistemological role because it reminded us that although computational tools are of great help in seeing more, they can produce many mirages, hallucinations, and afterimages. 

After the graph’s refinement, the workflow entered a semantic exploration phase, where the dataset became a navigable environment through which participants moved, tracing how shared stylistic terms, biographical similarities, and emotional categories generated valuable scientific meaning. Exploring these networks with the help of open data made visible the biases embedded in cultural metadata: some traditions carried richer descriptive histories than others, vocabularies often privileged Western categories, and certain interpretations suffered from repetitiveness. Closing data doesn’t erase these drawbacks, it just makes it more difficult to observe and understand.

Two Colab notebooks were used to translate the conceptual workflow into practical tools to support the exploration. The goal of one of the notebooks was to guide participants through loading the data into a graph, executing queries, and generating visualisations; while the other enabled artist networks to be investigated, allowing for the selection of particular figures and the examination of their structural positions. This solution has a clear aim – to enhance forth-comming research by allowing results to be downloaded for further analysis. These tools showed that sophisticated open-data workflows can be carried out using lightweight solutions and at little expense, which is an interesting addition that balances the use of costly infrastructure. 

5. Interpretation and the epistemic role of openness

We had the questions, then the data, then the analysis. The final stage of the workflow returned to interpretation. With the help of a carefully crafted prompt, participants sent their analytical outputs to a language model instructed to simulate expert reasoning. The goal was to produce contextual, interpretative commentary that had the combined expertise of an art historian, a curator, and a network-analysis specialist, with the intention of drafting ideas, testing hypotheses, and producing food for thought. This is not to replace researchers with computers, algorithms, and GPUs. AI as an interpretive partner is not a definitive authority, it’s simply an extension of humanistic inquiry.

Reflecting on the entire process reveals a broader lesson about open data in the humanities. Openness is not merely a commitment to accessibility or transparency; it is a methodological approach that reshapes how we organise research materials, articulate questions, and construct scientific arguments. The workflow demonstrates that a thoughtful sequence of steps – from curiosity to metadata design, data creation, PID assignment, graph construction, verification, exploration, and AI-supported interpretation – makes it possible to move confidently between theory and computation without sacrificing the nuance that defines humanistic scholarship.

Open data is, ultimately, an epistemic invitation to rethink categories, patterns, vocabularies, and the whole language of research, as well as to collaborate across disciplines and experiment with new analytical forms. Most importantly, it enables us to construct transparent and shareable research environments, allowing others to follow the same conceptual pathway and make humanities more reproducible and verifiable. A well-designed workflow becomes more than a technical sequence, it’s rather a way of thinking – open, critical, and generative, expanding the possibilities for scholarly work in a rapidly evolving landscape.

Cezary Rosiński
Cezary Rosiński

OpenEdition suggests that you cite this post as follows:
Cezary Rosiński (April 21, 2026). Designing an Open-Data Workflow for the Humanities: Notes from a Hands-On Workshop. SCIROS. Retrieved August 17, 2026 from https://sciros.hypotheses.org/2688


Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.