Physical Sciences › Computer Science › Information Systems
Research Data Management Practices
25 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Derniers papiers
- Critical Data Studies in the Anthropocene
Ana Valdivia · 22 septembre 2026
This chapter introduces the concept of the Anthropocene into critical data studies, a field that has, for the past decade, explored the entanglements between datafication and politics. With the scaling up of contemporary datafication alongside generative computing, it is essential to expand these de…
- Participant-Mediated Collection of Sensitive Digital Trace Data: The CANDOR Research Infrastructure
Andrew Zhao, Rijul Magu, Ekta Raj, Teresa Elinjikkal, Munmun De Choudhury · 17 septembre 2026
Digital trace data provide rich measures of behavior in everyday settings, but the research ecosystem supporting their collection is constrained by declining platform API access and a historical reliance on publicly observable data. Participant-mediated data donation offers a complementary approach …
- Measuring the Installed Base: Nordic Health Dataset Catalogues Against HealthDCAT-AP Release 7
Fabio Rovai · 31 août 2026
The European Health Data Space requires member states to publish machine readable descriptions of the health datasets available for secondary use, and the European Commission publishes HealthDCAT-AP as the metadata profile those descriptions are meant to satisfy. The profile has been designed and va…
- Turning interest into institutional change: teaching advocacy for sustainable research
Kirsty Pringle, Lorna Smith, Erinma Ochu, Greg Wilson · 20 août 2026
Systemic change across the digital research landscape is required to reduce the environmental impact of digital research, but while many researchers and technical professionals are motivated to act, they often lack the skills required to translate motivation into lasting organisational change. We pr…
- Data Findability, Governance, and Community Engagement for M\=aori Research Data Sovereignty
Paul T. Brown, Kiri West, Maree Sheehan, Hana Rapata, Te Taka Keegan · 11 août 2026
M\=aori data sovereignty (MDSov) has established important principles for recognising M\=aori rights and interests in data. Relatively little attention has been given to how these principles can be operationalised within research institutions who, as a function of their operations, collect and use M…
- From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
Muhsen Hammoud · 24 juillet 2026
Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, technology, and innovation (STI) analytics, yet suffer from temporal lag, semantic shallowness, and an inability to capture the non-linear dynamics of contemporary knowledge ecosystems. Dynamic…
- AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
Ming Chen, Pranav Pai · 20 juillet 2026
Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult. Existing evaluators use different rubrics and evidence sources and may fail on JavaScript-rendered pages or repository-specific identifiers. For 50 datasets fr…
- FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis
Marlena Flüh, Soo-Yon Kim, Carolin Victoria Schneider, Sandra Geisler · 14 juillet 2026
Retrieval-Augmented Generation (RAG) addresses the limitations of Large Language Models (LLMs) when providing responses to domain-specific questions. Graph-based RAG approaches, such as GraphRAG, enhance retrieval by capturing semantic relationships within knowledge graphs (KGs). While the FAIR prin…
- Causal evidence of racial and institutional biases in accessing paywalled articles and scientific data
Hazem Ibrahim, Fengyuan Liu, Khalid Mengal, Wisam Alshaibi, Aaron R. Kaufman, Yasir Zaki, Talal Rahwan · 9 juillet 2026
Scientific progress depends on researchers' ability to access and build upon the work of others. Yet, much published work remains behind expensive paywalls, and even accessible articles often rest on datasets shared only "upon reasonable request" to the authors. Researchers can try to overcome these…
- FAIR+S: A validation study of a framework for sustainable research data and software
Danila Valko, Jan S\"oren Schwarz, Jorge Marx G\'omez, Ralf Isenmann · 1 juillet 2026
The FAIR principles (Findable, Accessible, Interoperable, Reusable) have transformed research data management, but they do not address the environmental impact of creating and using research software and data, such as energy consumption, carbon emissions, and life-cycle impacts that become central t…
- Development and Design of FLKit: A Structured Onboarding Toolkit for Federated Learning in Health and Life Sciences
Ashkan Pirmani, Ilse Vermeulen, Goran Vinterhalter, Lotte Geys, Axel Faes, Muhammad Quamber Ali, Nishkala Sattanathan, Geert Vandeweyer, Yves Moreau, Liesbet M. Peeters · 23 juin 2026
Federated learning lets institutions train shared models without moving their data, which makes it a natural fit for health and life sciences research under strict privacy regulation. The methods are maturing fast, but the practical barrier now comes earlier: a team starting a federated project meet…
- The Dataset Friction Framework: measuring user-facing friction as a complement to FAIR
Emma Pidduck, Umberto Modigliani · 23 juin 2026
Open research data services have matured to the point where the cost of sustaining them at scale has become a primary design constraint, driving providers to make deliberate choices that may reduce user convenience to keep the service viable. The FAIR (Findable, Accessible, Interoperable, Reuseable)…
- AI for Monitoring and Classifying Data Used in Research Literature
Rafael Macalaba, Aivin V. Solatorio · 1 juin 2026
While platforms like Google Scholar and Semantic Scholar track citations for academic papers, no comparable infrastructure exists for monitoring dataset usage in research literature, leaving the landscape of data use largely opaque. Addressing this gap is critical for transparency, reproducibility, …
- Do Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
Shiyu Chen, Tarfah Alrashed, Alon Halevy, Natasha Noy · 28 mai 2026
In the era of autonomous agents, machine-actionable data is critical for data-driven workflows. For more than a decade, semantic metadata like schema.org has anchored the FAIR principles (Findable, Accessible, Interoperable, and Reusable) for machine-actionable data and enabled discovery tools like …
- Evaluating Structured Documentation as a Tool for Reflexivity in Dataset Development
Eshta Bhardwaj, Ciara Zogheib, Christoph Becker · 13 mai 2026
It is prominently recognized that dataset development in machine learning is a value-laden process from problem formulation to data processing, use, and reuse. Structured documentation frameworks such as datasheets, data statements, and dataset nutrition labels have been created to aid developers in…
- Autonomous FAIR Digital Objects: From Passive Assertions to Active Knowledge
Zeyd Boukhers, Oya Beyan, Cong Yang, Christoph Lange · 12 mai 2026
Scientific knowledge on the Web is published as passive assertions and cannot decide when to validate evidence, reconcile contradictions, or update confidence as findings accumulate. Curation depends on centralised middleware and institutional continuity, but when registries close, active stewardshi…
- Measuring research data reuse in scholarly publications using generative artificial intelligence: Open Science Indicator development and preliminary results
Lauren Cadwallader, Iain Hrynaszkiewicz, parth sarin, Tim Vines · 1 mai 2026
Numerous metascience studies and other initiatives have begun to monitor the prevalence of open science practices when it is more important to understand the 'downstream' effects or impacts of open science. PLOS and DataSeer have developed a new LLM-based indicator to measure an important effect of …
- A Proposed Biomedical Data Policy Framework to Reduce Fragmentation, Improve Quality, and Incentivize Sharing in Indian Healthcare in the era of Artificial Intelligence and Digital Health
Nikhil Mehta, Sachin Gupta, Gouri RP Anand · 14 avril 2026
India generates vast biomedical data through postgraduate research, government hospital services and audits, government schemes, private hospitals and their electronic medical record (EMR) systems, insurance programs and standalone clinics. Unfortunately, these resources remain fragmented across ins…
- Doctoral Theses in France (1985-2025): A Linked Dataset of PhDs, Academic Networks, and Institutions
William Aboucaya, Dastan Jasim · 13 avril 2026
This paper presents a comprehensive dataset of doctoral theses defended in France between 1985 and 2025, constructed from multiple national academic metadata sources. The dataset is primarily based on data from the French national thesis platform and is enriched using additional authority and biblio…
- Demystifying Funding: Reconstructing a Unified Dataset of the UK Funding Lifecycle
William Thorne, Rupert Shepherd, Diana Maynard · 30 mars 2026
We present a reconstruction of UKRI's Gateway to Research (GtR) database that links funding opportunities to their resulting project proposals through panel meeting outcomes. Unlike existing work that focuses primarily on funded projects and their outcomes, we close the complete funding lifecycle by…
- Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik · 18 mars 2026
Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queries, but lack understanding of their users. We address this with MyScholarQA (MySQA), a personalized DR agent that: 1) infers a profile with a user's res…
- Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
Kai Ruan, Xuan Wang, Jixiang Hong, Peng Wang, Yang Liu, Hao Sun · 24 février 2026
While Large Language Models (LLMs) demonstrate remarkable capabilities in scientific tasks such as literature analysis and experimental design (e.g., accurately extracting key findings from papers or generating coherent experimental procedures), existing evaluation benchmarks primarily assess perfor…
- The Climate Change Knowledge Graph: Supporting Climate Services
Miguel Ceriani, Fiorela Ciroku, Alessandro Russo, Massimiliano Schembri, Fai Fung, Neha Mittal, Vito Trianni, Andrea Giovanni Nuzzolese · 24 février 2026
Climate change impacts a broad spectrum of human resources and activities, necessitating the use of climate models to project long-term effects and inform mitigation and adaptation strategies. These models generate multiple datasets by running simulations across various scenarios and configurations,…
- Intelligent Knowledge Mining Framework: Bridging AI Analysis and Trustworthy Preservation
Binh Vu · 22 décembre 2025
The unprecedented proliferation of digital data presents significant challenges in access, integration, and value creation across all data-intensive sectors. Valuable information is frequently encapsulated within disparate systems, unstructured documents, and heterogeneous formats, creating silos th…
- SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
Luca Foppiano, Sotaro Takeshita, Pedro Ortiz Suarez, Ekaterina Borisova, Raia Abu Ahmad, Malte Ostendorff, Fabio Barth, Julian Moreno-Schneider, Georg Rehm · 15 décembre 2025
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more tha…
Autres sujets du thème Systèmes d'information
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Software Engineering Research853 papiers / 12 mois+218 %
- Information Retrieval and Search Behavior727 papiers / 12 mois+506 %
- Recommender Systems and Techniques600 papiers / 12 mois+154 %
- Expert finding and Q&A systems205 papiers / 12 mois+1175 %
- Information and Cyber Security169 papiers / 12 mois+1500 %
- Blockchain Technology Applications and Security146 papiers / 12 mois+650 %
