From FAIR Principles to AI-Ready Science
Sept 18, 2026
By Kimberly Mann Bruch, Lynne Schreiber and Christine Kirkpatrick
The scientific community continues to implement the FAIR principles, which aim to make data, software and models findable, accessible, interoperable and reusable. Where practices exist, most have not been widely adopted and they vary by discipline. Now, generative AI and large language models are introducing new challenges for how research data and infrastructure are prepared, documented and used.
A new special issue of AI Magazine, FAIR Principles and Machine Learning, AI Readiness, and AI Reproducibility, features research from the FAIR Principles and Machine Learning, AI Readiness and AI Reproducibility (FARR) community. Its five papers examine the intersection and continuing relevance of the FAIR principles in the age of AI, moving beyond aspirational tenets toward concrete, operational practices for making data, models and workflows ready for AI-driven science.
“Most of the AI readiness conversation has focused on algorithms and computation, without enough focus on the data that underpins frontier models and scientific foundation models. While the advances in deep learning have simplified text, image, and video analysis and production, well described and structured data is still critical for context. While specificity may not matter when an LLM suggests a grocery list for a given set of recipes, or a travel itinerary, it absolutely matters for integrating biomedical data, oceanographic observations, or energy traces - the sort of work being done at HSDSC and by our partners in FARR. This means that for AI-enabled research to produce results that can be understood, evaluated, reproduced and built upon, scientists need access to data that are findable, interoperable and accompanied by sufficient documentation about their origins, structure, quality and appropriate use,” said Christine Kirkpatrick, who directs Research Data Services at the San Diego Supercomputer Center and is an Associate Editor at AI Magazine, and co-edited the special issue and co-authored the introduction with Daniel S. Katz, Yuhan Rao and Lynne Schreiber. “We are proud to have worked with AAAI and AI Magazine to enrich the scholarship around AI readiness for data with researchers working at the forefront on this topic.”
From FAIR principles to AI-Ready data
Brewer et al., in “Data Readiness Pipeline Patterns for Scientific AI at Scale,” draw on work in the U.S. Department of Energy-funded laboratory across four demanding domains: climate, fusion, life sciences, and materials. Their framework pairs pipeline stages with a five-level maturity scale, enabling researchers to map a generic AI workflow onto data readiness and data processing and assess how far a dataset has to go. Cross-domain case studies demonstrate what the different levels look like in practice.
Majithia et al. approach the challenge from a different angle. “An Actionable Framework for AI-Ready Data” examines workflows across disciplines to identify common preprocessing steps, data quality challenges, and domain-specific constraints, then proposes a triad of practices: data must be technically suitable, self-describing, and programmatically accessible. Readiness, they argue, is not something that simply happens to a dataset; it is a matter of deliberate design.
Making metadata work for AI
Musen et al., in “Knowledge Engineering for Open Science,” describe CEDAR, which turns community reporting guidelines into reusable metadata templates linked to ontologies and controlled vocabularies. The templates can be deployed through web forms, embedded editors, and spreadsheets, demonstrating how knowledge engineering can help operationalize standards across research communities and repositories.
Stewart et al. extend this idea to physical sensors in “Datasheets for Machine Learning Sensors.” AI-enabled sensor systems support real-time data collection and decision-making in applications ranging from industrial anomaly detection to wildlife tracking, creating new challenges for documentation, verification, validation, and reproducibility. Their formalized datasheet framework provides a way to capture important information about these systems.
Greenberg et al., in “The Metadata Ecosystem and AI,” describe a symbiotic relationship between metadata and AI: structured metadata makes data interpretable and usable by AI systems, while AI can help generate, validate, and refine metadata at scale. The paper emphasizes metadata as an essential computational asset, rather than simply supporting documentation.
Collectively, these contributions reflect FARR’s focus on the interconnected challenges of FAIR data, AI readiness, and AI reproducibility.
Preparing for the future of AI-driven science
Together, the papers mark a shift from principle to practice. Moving forward will require interoperable infrastructure connecting data, models, and metadata across domains; standards that can evolve with rapidly changing technologies; and sustained collaboration among domain scientists, data stewards, information scientists, and AI researchers.
The goal is not simply to prepare today’s data for today’s AI, but to build adaptable foundations for the future of scientific discovery. As AI becomes increasingly embedded throughout the research lifecycle, strong data stewardship, transparent workflows, and reproducible practices will be essential to ensuring that AI-enabled science remains understandable, reusable, and trustworthy.
This work is supported by the FARR Research Coordination Network, funded by the National Science Foundation (award no. 2226453).





Comments