Digital resources in the Social Sciences and Humanities OpenEdition Our platforms OpenEdition Books OpenEdition Journals Hypotheses Calenda Libraries OpenEdition Freemium Follow us

Can we measure the impact of research infrastructures with the automated detection of software mentions?

Measuring the impact of Open Research Infrastructures (ORI) helps assess the effectiveness of infrastructures, justify investments in their development, and facilitate informed decision-making regarding the funding, development, and maintenance of services.

Current ORIs seek effective and objective ways of capturing their use. Due to the nature of open infrastructures, providing usage data on their services is particularly challenging. This phenomenon has been recently discussed in analyses and publications (see for example Cost-Benefit Analysis of Open Science: Case Study Synthesis and UniProt and RCAAP case studies) prepared by CSIL, one of the SCIROS project partners, as part of the project “Pathos: Open Science Impact Pathways”.

Ilustration 1. PathOS project’s logotype

As business models of ORIs depend on the wide public availability of resources, the collection of data that captures their impact is particularly challenging. It is important to develop and enhance capacities to better understand the use of ORIs and to facilitate empirical analysis of their use that could be communicated to funders, policymakers etc. This is of utmost importance today as ORIs face pressures to showcase their contribution to European competitiveness, with the European Union’s focus on knowledge valorisation (one of the recent initiatives in that field included the summit Driving EU Prosperity: The Future of Knowledge Valorisation, where SCIROS was present).

Automated detection of software mentions


One of the ways in which the impact of ORIs can be measured are software mentions in scientific literature (and other publications). The current reality is that software is often an “invisible” component in publications. Howison & Bullard found that 63% of software mentions are informal (mentions without references), with the remaining 37% being software names associated with paper references, and almost never with software references or PIDs. Automated detection of software mentions addresses this gap by transforming research software into a first-class bibliographic record that can be identified and registered using Persistent Identifiers (PIDs), thus making it

  • Findable,
  • Accessible,
  • Interoperable,
  • and Reusable (FAIR).

Key progress in this field has been achieved thanks to the creation of relevant datasets of software mentions. Schindler et al. developed a supervised information extraction pipeline, which was trained on the comprehensive SOMESCI gold standard corpus. This approach was applied to over 3.2 million scientific articles from PubMed Central (PMC), resulting in the creation of the Software Knowledge Graph (SoftwareKG). This dataset comprises 11.8 million software mentions. SoftwareKG was then used to provide extensive insights into the evolution of software usage and citation patterns over time across various scientific fields, journal ranks, and publication impact.

González-Guardia et al. introduced Softalias-KG, a novel knowledge graph explicitly focused on structuring software aliases. This graph was constructed by cleaning and transforming the clustered alias data generated during a recent large-scale analysis by the Chan-Zuckerberg Initiative (CZI), which extracted software mentions from over 3.8 million papers in PubMed Central. Softalias-KG organizes over 50,000 unique aliases into more than 34,000 unique software application groups, and is enriched with additional software entities and metadata imported from Wikidata.

A recent initiative advancing this area is the SoFAIR project, which aims to make progress in developing infrastructures for software mentions, especially through the integration of GROBID (GeneRation Of BIbliographic Data) and SoftcCite tools. As described by Du et al., the Softcite dataset was developed as the “gold-standard” corpus of software. This dataset comprises 4,093 manually annotated software mentions, extracted from the full text of 4,971 open access, academic articles in biomedicine and economics, published between 2000 and 2010. Softcite leverages GROBID (GeneRation Of BIbliographic Data), which employs deep learning algorithms, statistical models, and rule-based methods to process PDF documents and extract structured metadata, including titles, authors, abstracts, affiliations, and citations. The SoFAIR project has developed another gold standard dataset, through a task led by IBL PAN, and in collaboration with CLARIN-PL. This dataset contains more than 9,000 software mentions, divided into 10 categories, as well as more than 2,000 relationships between mentions, which come from almost 500 texts belonging to 18 scientific disciplines.

Ilustration 2. SoFAIR overall workflow

Are tools for automated detections of software ready for implementation by research infrastructure?


For various tasks existing tools claim precision scores ranging from around 0.75 to 0.9, but there is an agreement amongst experts that for reliable implementation manual curation is still required. The tools make frequent mistakes by incorrectly tagging parts of the text as software, partly due to the lack of comprehensive registries of software that could serve as reference databases. What these tools are good at is detecting trends over time, and general patterns in the data, e.g. comparing software-mention saturation between publication categories.


At the same time, a software producer could benefit from running the tools over publications and then developing a workflow for validating results for its own benefit, as with some technical expertise, valuable data can already be obtained through partially automated workflows, which may serve as a competitive advantage for research infrastructures.

References

Du C, Cohoon J, Lopez P, Howison J. Softcite dataset: A dataset of software mentions in biomedical and economic research publications. J Assoc Inf Sci Technol. 2021; 72: 870–884. https://doi.org/10.1002/asi.24454

Pride, David; Cancellieri, Matteo; Knoth, Petr; Rosinski, Cezary; Umerle, Tomasz; Rudnicka, Ewa; et al. (2025). SoFAIR Project Dataset – An annotated dataset of software mentions in full text scientific articles.. The Open University. Dataset. https://doi.org/10.21954/ou.rd.30374830.v2

Schindler D, Bensmann F, Dietze S, Krüger F. 2022. The role of software in science: a knowledge graph-based analysis of software mentions in PubMed Central. PeerJ Computer Science 8:e835 https://doi.org/10.7717/peerj-cs.835

Howison, J. and Bullard, J. (2016), Software in the scientific literature: Problems with seeing, finding, and using software mentioned in the biology literature. J Assn Inf Sci Tec, 67: 2137-2155. https://doi.org/10.1002/asi.23538

tomaszumerle
tomaszumerle

OpenEdition suggests that you cite this post as follows:
tomaszumerle (November 28, 2025). Can we measure the impact of research infrastructures with the automated detection of software mentions? SCIROS. Retrieved September 12, 2026 from https://sciros.hypotheses.org/1604


Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.