Open research information infrastructures, scientific databases, data repositories, and community-driven research platforms have become essential components of modern scientific discovery. They enable researchers to access, validate, share, and reuse scientific data while supporting reproducibility and collaboration across disciplines. However, the impact of such infrastructures is still commonly evaluated through conventional bibliometric indicators such as citation counts, publication trends, co-authorship networks, and keyword statistics. Although these measures capture scholarly visibility and influence, they provide limited insight into a more fundamental question: How are research infrastructures, datasets, databases, and computational tools actually used within scientific workflows?
To address this gap, we introduce a domain-informed, Natural Language Processing (NLP)-driven framework for full-text scientometric analysis. The framework moves beyond citation-based assessment by analysing the scientific content of publications and extracting direct evidence of infrastructure usage. Rather than treating a research article solely as a bibliographic record, the proposed approach treats its full text as a rich source of information about the datasets researchers rely on, the computational tools they adopt, the scientific entities they study, the workflows they integrate, and the research themes that evolve over time.
As illustrated in Figure above, the proposed framework establishes two complementary analytical perspectives. The first represents the conventional bibliometric approach, in which bibliographic metadata are used to derive indicators such as citations, co-authorship patterns, keywords, and publication trends. These measures remain valuable for understanding the visibility and broad scholarly influence of a research infrastructure.
The second perspective introduces the proposed NLP-driven full-text framework. Here, complete scientific publications are processed through information-extraction methods and interpreted using a dedicated domain knowledge layer. This layer incorporates discipline-specific entities, datasets, repositories, software tools, workflow relationships, semantic context, and underlying research themes. The extracted information is subsequently transformed into infrastructure-centric scientometric indicators, capturing dimensions such as dataset reuse, software adoption, workflow integration, entity–database relationships, and thematic evolution.
Together, these complementary layers provide a richer and more integrated understanding of scientific impact by combining who cites an infrastructure with how that infrastructure is actually embedded within research practice.
The complete analysis pipeline is released as open-source software, supporting reproducibility, community validation, and future adaptation across scientific domains.
The proposed methodology provides a modular and transferable framework for usage-aware scientometrics. Generic components-including document processing, NLP preprocessing, co-occurrence analysis, topic modelling, and visualization-can be retained across scientific fields, while the domain knowledge layer can be adapted to different entities, repositories, datasets, software tools, and scientific vocabularies.