Onedata Publications
Scientific papers and peer-reviewed articles from the Onedata research team.
Smart Data Management and ML-Based Workflow Prediction for Cognitive Compute Continuum in SPICE
Integrated Data, Metadata, and Paradata Management System for 3D Digital Cultural Heritage Objects: Workflow Automation, Federated Authentication, and Publication
The complexity of high-quality 3D digitised cultural heritage objects creates challenges for existing data management systems as they need to develop metadata management and processing capabilities to provide semantic insight into the interconnectivity of data that constitutes cultural heritage objects. To address these challenges, we propose a global federated authentication and authorisation mechanism, a data and metadata management system, and an integrated engine for designing and executing...
Full details & cite →Automated Management of 3D Digital Cultural Heritage Objects in Eureka3D with Onedata
Cultural Heritage 3D Object Management with Integrated Automation Workflows
The complexity of high-quality 3D digitised cultural heritage objects creates challenges for existing data management systems as they need to develop metadata management and processing capabilities to provide semantic insight into the interconnectivity of data that constitute cultural heritage objects. To address these challenges, we propose a data and metadata management system, together with the federated authentication and authorisation mechanism, and an integrated system for designing and...
Full details & cite →Distributed Management and Processing of ALICE Monitoring Data with Onedata
Achieving Decentralized Authority for Collaborative Data Sharing with Consensus
In the age of globalization and international collaboration, collaborative data sharing is crucial for computer-aided research. There is a growing demand for spontaneous, short-lived, and fine-grained collaboration acts that could be satisfied by a decentralized platform for the management of entities (users, groups, datasets). Our research is based on a pre-existing concept of such a decentralized platform, called the decentralized entity persistence layer (DEPL). We describe a novel...
Full details & cite →Indexing Legacy Data-Sets for Global Access and Processing in Multi-Cloud Environments
Global data access and processing for multi-cloud applications proves challenging where there is a need for efficient access to large pre-existing, legacy data-sets. To address this problem, we created an indexing subsystem, allowing us to index and gather the metadata of legacy data-sets. The gathered information is later incorporated into a multi-cloud data management system to achieve global integration of legacy storage systems containing large data-sets. The solution is based on metadata...
Full details & cite →Towards Open Science with Multi-Cloud Computing Using Onedata
Global Access to Legacy Data-Sets in Multi-Cloud Applications with Onedata
Data access and management for multi-cloud applications proves challenging where there is a need for efficient access to large pre-existing, legacy data-sets. To address this problem, we created an indexing subsystem incorporated into Onedata data management system achieving a global multi-cloud integration of legacy storage systems containing large data-sets. The solution is based on metadata management, organization and periodic monitoring of legacy data-sets, what makes possible scheme-less...
Full details & cite →Group Membership Management Framework for Decentralized Collaborative Systems
Scientific and commercial endeavors could benefit from cross-organizational, decentralized collaboration, which becomes the key to innovation. This work addresses one of its challenges, namely efficient access control to assets for distributed data processing among autonomous data centers. We propose a group membership management framework dedicated for realizing access control in decentralized environments. Its novelty lies in a synergy of two concepts: a decentralized knowledge base and an...
Full details & cite →Increasing Data Availability and Fault Tolerance for Decentralized Collaborative Data-Sharing Systems
In order to realize collaboration on a global scale, academic research requires often large quantities of data to be shared between geographically dispersed organizations. The requirement to protect and govern data in a network of loosely coupled, autonomous institutions is an incentive for decentralized solutions, where the participants are in full control of their data without trusting a third-party provider to store and process the data. In order to increase data availability and fault...
Full details & cite →New Approach to Global Data Access in Computational Infrastructures
Data access and management on a large, global scale is currently at the center of scientific interest. This follows from the need for data access from multi- and hybrid-cloud applications. In most cases existing solutions provide sufficient functionality to scale computing resources but scaling resources in terms of efficient data access e.g. for data-intensive applications is still not comprehensively resolved. In this paper, we present a new approach to global data access that supports the...
Full details & cite →Scientific Workflow Management on Hybrid Clouds with Cloud Bursting and Transparent Data Access
Cloud bursting is an application deployment model wherein additional computing resources are provisioned from public clouds in cases where local resources are not sufficient, e.g. during peak demand periods. We propose and experimentally evaluate a cloud-bursting solution for scientific workflows. Our solution is portable thanks to using Kubernetes for deployment of the workflow management system and computing clusters in multiple clouds. We also introduce transparent data access by employing a...
Full details & cite →The eXtreme-DataCloud Project - Solutions for Data Management Services in Distributed e-Infrastructures
The eXtreme DataCloud (XDC) project is aimed at developing data management services capable to cope with very large data resources allowing the future e-infrastructures to address the needs of the next generation extreme scale scientific experiments. Started in November 2017, XDC is combining the expertise of 8 large European research organisations. The project aims at developing scalable technologies for federating storage resources and managing data in highly distributed computing...
Full details & cite →Advancements in Data Management Services for Distributed e-Infrastructures: The eXtreme-DataCloud Project
The development of data management services capable to cope with very large data resources is a key challenge to allow the future e-infrastructures to address the needs of the next generation extreme scale scientific experiments. To face this challenge, in November 2017 the H2020 eXtreme DataCloud - XDC project has been launched. Lasting for 27 months and combining the expertise of eight large European research organisations, the project aims at developing scalable technologies for federating...
Full details & cite →Transparent Data Access for Scientific Workflows across Clouds
We present a scientific workflow data management solution that combines global data access with a block-level optimization of data transfer, wherein only the data blocks that are used by a remote job are transferred over the network, significantly reducing data movement for specific common data access patterns. We propose the implementation of the solution based on the HyperFlow workflow management system and the Onedata data management platform. Preliminary results confirm the advantages of...
Full details & cite →Harmonizing Sequential and Random Access to Datasets in Organizationally Distributed Environments
Computational science is rapidly developing, which pushes the boundaries in data management concerning the size and structure of datasets, data processing patterns, geographical distribution of data and performance expectations. In this paper we present a solution for harmonizing data access performance, i.e. finding a compromise between local and remote read/write efficiency that would fit those evolving requirements. It is based on variable-size logical data-chunks (in contrast to fixed-size...
Full details & cite →Scientific Computing with Application Containers: Onedata and HyperFlow Use Cases
Trust-Driven, Decentralized Data Access Control for Open Network of Autonomous Data Providers
The observation of current trends in data access, especially in the field of scientific computations, shows that global data access that crosses federation boundaries is highly desirable. However, administrative constraints require that data centers remain autonomous, which effectively eliminates the possibility of cooperation. To overcome this, we plan to establish an open network of cooperating data providers. In this paper, we address the issue of data access control for such network. Our...
Full details & cite →Metadata Synchronization Protocol for a Decentralized Network of Data Providers
Two-Layer Load Balancing for Onedata System
The recent years have significantly changed the perception of web services and data storages, as clouds became a big part of IT market. New challenges appear in the field of scalable web systems, which become bigger and more complex. One of them is designing load balancing algorithms that could allow for optimal utilization of servers' resources in large, distributed systems. This paper presents an algorithm called Two-Level Load Balancing, which has been implemented and evaluated in onedata -...
Full details & cite →Onedata Virtual Filesystem for Hybrid Clouds
Towards Transparent Data Access with Context Awareness
Open-data research is an important factor accelerating the production and analysis of scientific results as well as worldwide collaboration; still, very little data is being shared at scale. The aim of this article is to analyze existing data-access solutions along with their usage limitations. After analyzing the existing solutions and data-access stakeholder needs, the authors propose their own vision of a data-access model.
Full details & cite →Concept of Decentralized Access Control for Open Network of Autonomous Data Providers
Consistency Models for Global Scalable Data Access Services
Developing and deploying a global and scalable data access service is a challenging task. We assume that the globalization is achieved by creating and maintaining appropriate metadata while the scalability is achieved by limiting the number of entities taking part in keeping the metadata consistency. In this paper, we present different consistency and synchronization models for various metadata types chosen for implementation of global and scalable data access service.
Full details & cite →Effective and Scalable Data Access Control in Onedata Large-Scale Distributed Virtual File System
Nowadays, as large amounts of data are generated, either from experiments, satellite imagery or via simulations, access to this data becomes challenging for users who need to further process them, since existing data management makes it difficult to effectively access and share large data sets. In this paper we present an approach to enabling easy and secure collaborations based on the state of the art authentication and authorization mechanisms, advanced group/role mechanism for flexible...
Full details & cite →Using Onedata for Global Sharing and Processing of Legacy Large Data Sets
Kademlia with Consistency Checks as a Foundation of Borderless Collaboration in Open Science Services
The concept of Open Science emerges as a powerful new trend, allowing researchers to exchange and reuse valuable knowledge, data and analyses. Innovative tools are needed to facilitate such global scientific collaboration, which is the main objective of the Onedatasystem. It aspires to provide a Open Science platform based on openness and decentralization. To achieve this, a distributed location service must be introduced that will allow to locate resources in this vast environment. This paper...
Full details & cite →Onedata – A Data Management Platform Supporting Global Open Science
Implementation of Open Data at Global Scale in Low-Trust Environment
Metadata Organization and Management for Globalization of Data Access with Onedata
The Big Data revolution means that large amounts of data have not only to be stored, but also to be processed to unlock the potential of access to information and knowledge for scientific research. As a result, scientific communities require simple and convenient global access to data which is effective, secure and shareable. In this article we analyze how researchers use their data working in large scientific projects and show how their requirements may be satisfied with our solution called...
Full details & cite →Delegation of Authority in a Distributed Data Access System
Efficient Storing of Metadata for Distributed Data Management
Onedata – A Step Forward towards Globalization of Data Access for Computing Infrastructures
To satisfy requirements of data globalization and high performance access in particular, we introduce the originally created onedata system which virtualizes storage systems provided by storage resource providers distributed globally. Onedata introduces new data organization concepts together with providers' cooperation procedures that involve use of Global Registry as a mediator. The most significant features include metadata synchronization and on-demand file transfer.
Full details & cite →Two-Level Load Balancing for Onedata
Globalization of Data Access for Computing Infrastructures
Accounting and Monitoring in Distributed Storage Services
Uniform and Efficient Access to Data in Organizationally Distributed Environments
Harnessing Organizationally Distributed Data with VeilFS
Storage Management Systems for Organizationally Distributed Environments PLGrid PLUS Case Study
With the increasing amount of data the research community is facing problems with methods of effectively accessing, storing, and processing data in large scale and geographically distributed environments. This paper addresses major data management issues, in particular use cases and scenarios (on the basis of Polish research community organized around the PLGrid PLUS Project) and discusses architectures of data storage management systems available in both PL-Grid and other similar federated...
Full details & cite →