Speaker 1
Welcome to the BigOpenScience podcast. Today we are in Munich where we are hosted by Professor Sabine Leonelli and her team at the Technical University of Munich. Today we will be talking about the Ethical Data Initiative, an inspiring project that aims to foster open discussions about ethics of data. As a non-partisan platform, EDI is coordinated by the University of Exeter and TUM Think Tank. Hello and welcome to the Big Open Science Podcast. Today our guests are Paul Trautmannsdorff and Silvia Milano from Ethical Data Initiative. Good afternoon.
Speaker 2
Good afternoon.
Speaker 3
Good afternoon.
Speaker 1
Today we will be discussing about your project, Ethical Data Initiative. So, but before we start, if you could shortly introduce yourselves and what is your role in this project.
Speaker 3
My name is Paul Krautmannsdorff. I am a research fellow in the Chair of Philosophy and History of Science and Technology and in the Ethical Data Initiative here in Munich. And my role here is I am responsible for the education pillar of the Ethical Data Initiative together with Kim Hayek, a colleague also in the chair and in the initiative.
Speaker 2
And hi, I’m Silvia Milano. I’m the head of research for the Ethical Data Initiative here at the TUM Munich, and I also teach on ethics of AI and social epistemology and various topics related to AI in the chair of history and science of philosophy and history of science and technology.
Speaker 1
Perfect, thank you. Okay, so let’s, let’s go for the Ethical Data Initiative. How did it start? What inspired it? And how it operates today, maybe the most important thing, including the size of the team, the types of activities you focus on. Just, you know, a brief introduction to the project.
Speaker 3
Yeah, the Ethical Data Initiative is relatively young. It was launched last year, spring last year, in here in Munich at the TUM Think Tank. It’s spearheaded by Sabina Leonelli, professor and chair in philosophy and history of science and technology. And it is now jointly coordinated by the TUM Think Tank and the University of Exeter. So part of our team is in Exeter. Very generally, we are a non-partisan platform that fosters open discussions on data ethics, but is also significantly informed by philosophical, historical, and social studies of science, with a particular focus on equity and engagement across different domains of data work for public interest. And I’d say this sort of focus on equity and the engagement with various stakeholders, data learners, data communities, of a variety of different domains and fields is sort of the focus of the initiative. So we are particularly trying to engage with people who have not sort of immediate access or also expertise in data ethics and governance and try to collaborate as much as possible.
Speaker 1
Celia, would you like to add something?
Speaker 2
I think this is a great introduction to what we do. I maybe would highlight a few of the activities that have been really central to, to the approach of the EDI. So one is around the production of data stories that are used as a teaching tool and also as a research as a way to conduct research into interesting topics and to understand how data works and data journeys across different use cases. And also I wanted to mention that we are very much dedicated to education, outreach, and engagement, but all of this is underpinned by our commitment to research. And so we are really We really try to be a community where research is informed by these outreach activities and is also geared towards producing good teaching materials, and this in turn informs the way that we also approach our research.
Speaker 1
Do you have any specific case that you would like to share with us, the case that you recently covered as the initiative?
Speaker 3
I can speak maybe on one particular activity we are, you know, conducting in the education pillar. This is the format of data clinics. It’s a particular collaborative, interactive teaching format also, and we try to invite partners, partner organizations of the Ethical Data Initiative, to to come and bring a kind of real-world challenge drawn from their actual data work, data environments and issues, and collaborate with students here at TUM in order to find throughout this clinic solutions, reflections, and yeah, find outcomes really to the challenge that a partner brings to this to this clinic format. So I think that’s one of the recent clinics that we had was on a challenge around responsible fintech in Western African countries. So we had a partner from the Rwandan Charter Foundation, the Center of Law and Innovation, who brought this challenge to our our clinic, Data Clinic, and collaborated with students on finding, identifying, first of all, obviously ethical sensitive issues around data governance, on identifying risks, also mapping, trying to understand better the data journey from sort of, for example, data to particular score profiles in financial technologies and also of trying to reflect on opportunities to not solve, but at least find ways to think about financial technologies in Western African countries in a more responsible and sustainable way.
Speaker 1
These days, it is extremely hard to speak about data governance practices and research data management without referring to AI and the rise of AI technologies? And how do you approach this problem in Ethical Data Initiative?
Speaker 2
Okay. So this obviously is a multifaceted question and there’s so many different contexts in which AI is currently being used. And so Ethical Data Initiative cannot have a single answer for all of these domains. For instance, the way in which AI is deployed in agriculture will be very different from how it is deployed in the kind of systems that I am currently more focused on, which are recommender systems, where it’s more about personal data and collection of behavioral data from users and inferences about people. Whereas I know that many other researchers in our group focus, for instance, on use of AI in agriculture, where the kind of problems and issues of data access can be very different. So maybe I can speak a bit about the area that I am currently working on. So as I said, these recommender systems is a type of algorithmic systems that are used to structure the information that we see online. So for instance, whenever you log into your social media, there will be some algorithmic feed that is powered by AI, you could say, trained on the personal data and, yeah, millions of users. And it’s nominally trying to predict what might be relevant for you, what, for instance, in the case of social media, what post you might like to see, or what kind of reactions you might have two different pieces of content and then trying to serve those pieces of content that are predicted to have the right reaction from you.
Speaker 2
So in these kind of applications, they’re obviously very pervasive. So everywhere in our online life, we will encounter many different recommendation systems. Trained on various types of data.
Speaker 1
The recommender system use data that are sensitive and these are mostly behavioral data. And in which disciplines is this specifically dangerous? I mean, in which disciplines you encounter this problem very often?
Speaker 2
So I’m thinking of commercial applications that, as I said, this could be social media or it could be e-commerce websites. So in all these cases, one of the issues is that the research and the development of these systems that is being done within companies is very disconnected from academic research that has been going on in the fields of computer science or other disciplines that could be relevant to these kind of decision scenarios. That is actually one of the main issues that we are identifying as a focus for our work in the next months, in the upcoming semester. Because we are observing that there is this disconnect where academic research into recommender systems, especially research into recommender systems that might be termed as for social good. So for promoting socially good outcomes, for instance, for steering people towards more environmentally friendly behaviors or ensuring that there’s fair access to information, for instance, about news and things like that, is very disconnected from what is going on in the real world in applications that are more commercially that are developed by commercial actors. This shows on various levels. So on the level of the data that is used for developing the currently more advanced commercial applications, that is completely or almost completely out of reach from academic researchers.
Speaker 2
There’s very little that researchers can do when they might have an interest in accessing this type of data, for instance, to conduct social science experiments, it’s usually not possible because of proprietary issues. And therefore, there’s a limited scope for advancing academic research using cutting-edge datasets which are outside of the reach. And so what happens is that a lot of the interesting things that are happening that would have potentially a lot of relevance for also the social sciences, understanding how our behavior is influenced by algorithmic architectures, is actually very difficult to set up. And we can only guess or set up experiments that don’t have access to to the datasets that the commercial applications have. They can only often be observational studies that track people’s behavior, keeping the AI recommender systems as black boxes in the background. So this is a big issue, and as EDI, we would like to put a spotlight on this and try to raise awareness of the fact that there needs to be more research done into recommend systems for social good in the first place. And also to spread maybe awareness that we need to demand more access to more quality data.
Speaker 2
And academic research needs to, in this sense, shift from relying on some datasets that have become sort of canonical. For instance, in the case of computer science work on Recommender systems. This could be MovieLens datasets about how people’s tastes about movies or some datasets that have been really overexploited. And instead try to pool resources together and share information about alternative datasets that may be more useful to train different applications and to explore different ways of using recommender systems to both have applications for social good and also understand or use them in a social science perspective to understand how people behave in interaction with these systems.
Speaker 1
When you speak about Ethical Data Initiative, I see it as a very diverse project, because you, Paul, spoke about teaching, which is very important. Like, not only, I know, and about reaching out to stakeholders, different stakeholders, and educating them, or at least exchanging experience and knowledge regarding data with them. And you, you speak about very specific problem that is, that is caused by— well, let’s not be, let’s not be afraid of saying that it’s caused by commercial companies that actually develop this type of systems. Right now we can use them for social good, but it hasn’t been the case so far. So when you, when you speak about it, what I wonder is, as we came here to Munich to actually discuss this trust in science and how we can, how we can basically fight this informational misinformation that is related to science, And I wonder whether you, uh, you think about it, whether you, whether you have— you, you pose yourselves this type of questions, how you as Ethical Data Initiative can counter, um, counter-fight disinformation and misinformation caused by, yeah, by also by scientific act, by actors that are related, or they are or simply scientists or scholars?
Speaker 3
Yeah, it’s a very big question. The issue of trust and distrust is, of course, a huge issue today. And obviously, we encounter that also in the Ethical Data Initiative, in our activities, in the engagements we have with our partners, with students, etc. I think, I mean, there’s no sort of easy solution to this. I think obviously also one of our aims in the Ethical Data Initiative is to help understand where this distrust comes from. For example, caused by AI, the huge amounts of data gathered that originate from sources that are known, that are unreliable, etc. At the same time, I think we are sort of trying to approach data ethics from a sort of practice-based and situational approach, which means that we really try to reflect and implement ethics, so to say, together in collaboration with partners and stakeholders, and try to think of something that is not separate from the everyday science and everyday data work, but an integral practice. And so that’s our approach to ethics. That’s how we try to infuse or to try to think about ethics and data governance in every stage of the data work. And I think that— and also assess what kind of implications this has for forms of knowledge, the accessibility to knowledge for various stakeholders, various actors, but also the various effects that has, for example, for planetary health and issues like trust and distrust.
Speaker 3
But there’s no one-size-fits-all solution. I think that’s very clear. And yeah, I think with our engagement with various stakeholders, actors, partners, the network, we are trying to find various types of strategies, interventions at multiple levels. Yeah, I think that’s what I would say.
Speaker 2
Yeah, I totally agree with what both said so far. I maybe can add some perspective with regards to the topic of recommender systems that we are now putting a spotlight on. Obviously, if you are thinking of misinformation spread online and similar phenomena, recommender systems are really instrumental to that phenomenon because it’s the algorithmic infrastructure that permits this kind of viral information spreading. So it It is very important that we understand these phenomena, so these algorithmic systems, how they are set up. What we are currently observing is that in the academic research, as I said, in the research that comes out of computer science departments, there’s a lot of focus on designing accurate algorithms, but this is really disconnected with what is relevant in the real world when, for instance, we are approaching issues like misinformation. Because misinformation has little to do with how accurately you can predict whether someone will click on a certain piece of content. It has more to do with a deeper understanding of the meaning of the content and what its transmission between different nodes, different people in the network, will affect in terms of how belief spreads and how they relate to other beliefs and maybe to political perspectives and so on and so forth.
Speaker 2
And all these analyzes cannot be really done in isolation when you’re just defining a target variable, say a behavioral variable, and defining a very simple evaluation metrics. To evaluate your algorithm, which is not really what is going on in the real world. So one thing that we would like to do as EDI is try to be a catalyst for also the community, the research community in this specific case, to be able to share different approaches. To evaluating recommender systems. As I said, sharing information and possibility to access datasets to train and research recommender systems with different properties and share also best practices about how to evaluate these recommender systems. And this also speaks to the point of equity because currently we are either relying on very simplified, simplistic evaluation metrics. Recommender systems are examples, but I think this probably applies across the board to many AI systems because they are convenient for research and the way in which research publication works. On the other side, we have commercial actors that may have very complex evaluation metrics with complex business KPIs that are not well understood and not transparent outside of their organization. And it’s difficult to make a connection.
Speaker 2
So what we want to do instead is develop a community where best practices for evaluation can be discussed. And this also means that different voices can bring different ideas to the table and stakeholders that are currently not taken into account when designing, for instance, a recommendation algorithm for say news or social media posts, will, may actually have a voice in defining what is relevant and what should be measured. So these are two ways in which we want EDI to have a positive contribution.
Speaker 1
Really, thank you very much for, for these explanations, because to be honest, they are very much in line with what we discussed yesterday during the workshop. That we did with Sabina Lonelli in the Technical University of Munich. So, so yes, I also believe in grounded and situation-oriented and stakeholders-oriented solutions. So let’s hope for the best and that in the future we will have some means at least to, to, if not fight, to at least solve some of the problems that disinformation may cause in our social space and our social communication. Thank you once again for the conversation.
Speaker 2
Thank you. Thank you. Thank you. Bye. Stay connected and follow Skyros for more insight. Find us at our blog skyros.hypothesis.org. SKYROS project is made possible thanks to the support of the Polish National Agency for Academic Exchange under the Strategic Partnership Program. And remember, science is best when it’s open. See you next time on Big Open Science Podcast.
OpenEdition suggests that you cite this post as follows:
Gabriela Manista (May 11, 2026). Big Open Science Podcast | S01E12: Inside the Ethical Data Initiative: research, data and trust. SCIROS. Retrieved August 16, 2026 from https://sciros.hypotheses.org/2775

