Friday, 23 October 2009

EIDCSR technical analysis: from soft to hard

After having conducted the EIDCSR audit and requirements analysis exercise, we have started converting the high level requirements gathered into technical requirements. The idea is to produce a systems design document for a Systems Developer to start with the implementation. Howard Noble, from Computing Services, is leading this exercise for the next two months.

To start with the technical analysis, Howard and I have had a very fruitful meeting this morning. We have brainstormed ideas for a high level system design trying to identify the practical things that can be done to support the data management workflows of the research groups taking part in EIDCSR.


Using a board to produce a "rich picture" recording the processes we have encountered and our thoughts was extremely useful. We will now produce a "cleaner" version of this picture and bring it forward to key people in the research groups in a workshop. This will hopefully helps us to communicate what the project aims to achieve as well as getting feedback on the design so that researchers requirements drive any development .




Thursday, 15 October 2009

First EIDCSR workshop and executive board meeting

Yesterday was a busy day for the EIDCSR Project.

In the morning, the first project event took place at Rewley House in Oxford with an exciting group of speakers brought together under the theme of "Data curation: from lab to reuse". Their presentations are now available on the project website and a report will be produced shortly.

The afternoon served to held the first EIDCSR Executive Board meeting where progress and next steps for the project
were discussed with the extraordinary helpful and encouraging members of the board.

Overall, a great day providing loads of food for thought.

Monday, 12 October 2009

"Science these days has basically turned into a data-management problem"

The New York Times has an article about future scientists' ability to manage the large amounts of digital data being generated and how the likes of IBM or Google are trying to help, "Training to Climb an Everest of Digital Data", http://www.nytimes.com/2009/10/12/technology/12data.html. IBM and Google are contributing tools, computational power and access to large-scale datasets. It was actually two years ago this month that Google and IBM announced their partnership to provide universities with dedicated cluster computing resources, open source software, a dedicated website for collaboration, and a Creative Commons-licensed curriculum. In April this year the NSF funded projects at 14 US universities to take advantage of the IBM/Google Cloud Computing University Initiative. The New York Times article highlights some of these projects. The emphasis is certainly on the massive -- big compute clusters, big datasets -- and on data analysis. Not much though on the ongoing management of, access to, and preservation of data, even if Professor Jimmy Lin (University of Maryland) is quoted as saying, “Science these days has basically turned into a data-management problem”.

Wednesday, 23 September 2009

EIDCSR workshop on 14 October

The first EIDCSR project workshop is taking place on 14 October, more details below:


Date and location

14 October at Rewley House, 1 Wellington Square, Oxford OX1 2JA

The event will start at 10.30 and will finish with lunch at 13.00


Description

This workshop is organized as part of the dissemination activities of the JISC-funded EIDCSR Project. The aim of the workshop is to hear about proven practice in selected data management areas identified as challenging for researchers through the EIDCSR audit and requirements analysis exercise. Whilst the EIDCSR Project is addressing the requirements of researchers working within medical and life sciences, the event is likely to be of interest to those working in, or supporting, other disciplinary areas.

The expected audience includes researchers who generate data in labs and computing simulations and staff from service units with an interest in research data management and curation issues.


Outcomes

Participants in the workshop will have the opportunity to learn about, and contribute to discussion of, the different approaches to the ensuring the flow of data between laboratory and in silico experimentation. In particular, the workshop will discuss:

* methods for the capture, storage and reuse of metadata in the laboratory;

* lifecycles integrating wet lab and in silico experimental data;

* for delivery and visualisation of large-scale data.

Programme

Some of the speakers will include:


Alan Garny, Oxford Department of Physiology Anatomy and Genetics - Alan will discuss his research group data management workflow and challenges.

Brian Brooks, Unilever Cambridge Centre for Molecular Informatics - Brian will talk about their Chemical Laboratory Repository In/Organic Notebooks (CLARION) Project.

Angus Whyte, Digital Curation Centre - Angus will share the experiences from the DCC SCARP Project on data management best practice.

Booking

To book a place please email eidcsr@oucs.ox.ac.uk

Wednesday, 9 September 2009

Data audit and requirements analysis

One of the initial exercises to be conducted as part of the EIDCSR project was the audit and requirements analysis based on DAF to document the data practices and assets as well as to capture the requirements for tools and services of the research groups participating in the project. This exercise took place throughout the summer and the report describing the results will be available soon.

As I explained on a previous post, these research groups collaborate as part of a BBSRC grant to conduct research on ventricular architecture by using novel techniques such as Magnetic Resonance Imaging (MRI) and Diffusion Tensor MRI (DTMRI) and combine them with traditional histological techniques as well as with image processing with data registration and computational models for bio-mathematical simulation.

Their research workflow is well described by Gernot et. all (2009)* in the diagram below. It starts with the generation of complementary images stacks that are then processed in different ways to generate meshes that can be used for computational modelling of the heart.


The result of this complex process produces the following data outputs:
  • Histology data: large high resolution images produced by microscopes in the lab representing sections of a heart.
  • MRI and DTMRI data: stack of tiff images resulting from the raw data produced by the magnet in a lab.
  • Segmentation data: outputs resulting from applying image segmentation techniques to the histology and MRI data.
  • Mesh data: volumetric model produced from segmented data in a mesh generator.
  • Simulations: electrophysiological simulation using the mesh data and other input files that define the models and the parameters.
  • 3D heart atlas: representing an average representation of a heart ventricles obtained from the histology and MRI data.
And the research group requirements can be grouped under three themes:
  • Secure storage: all the data outputs presented above are stored on a combination of desktop computers and a project NAS system and researchers realize the need to keep the data safe by having appropriate and resilient back-up procedures.
  • Data transfer: the histology data are large and needs to be accessed by researchers within the groups and others.
  • Metadata: currently the provenance metadata for some of the data presented above is recorded in printed lab-books. This information is crucial when making the data available to others and it is required when publishing articles based on the data. In addition to this, it may be helpful to improve searching within the NAS system.
*Gernot Plank, Rebecca A.B. Burton, Patrick Hales, Martin Bishop, Tahir Mansoori, Miguel O. Bernabeu, Alan Garny, Anton J. Prassl, Christian Bollensdorff, Fleur Mason, Fahd Mahmood, Blanca Rodriguez, Vicente Grau, Jürgen E. Schneider, David Gavaghan, and Peter Kohl Generation of histo-anatomically representative models of the individual heart: tools and application Phil Trans R Soc A 2009 367: 2257-2292.

Tuesday, 28 July 2009

Provenance metadata: what and how to record it?

To effectively curate the research data produced by the two research groups participating in the EIDCSR project, it is crucial to capture provenance metadata that explains how the data was generated in the first place. This information enables validation and increases the value of the data.

The research groups in our case, collaborate as part of a BBSRC funded project and generate MRIs and histology data in laboratories using a variety of instruments and techniques, these datasets are then manipulated through process such segmentation to create 3D meshes, volumetric elements, that will serve to run computational simulations.

So what provenance metadata should be recorded and are there any subject specific metadata standards appropriate for these datasets?

Interviews with the researchers involved in the generation of data have shown that they well versed in recording information about their experiments on their lab-notebooks. When writing research articles they go back to these notebooks in order to document their methodologies. Therefore, I believe it is fair to assume that researchers know what information needs to be recorded about their experiments and simulations.


Discussions with the person responsible for the BBSRC Data Sharing Policy around metadata standards pointed us to the Minimum Information for Biological and Biomedical Investigations (MIBBI) portal. This resource provides minimum information guidelines for diverse bioscience domains and provides a registry of projects developing those guidelines.


A metadata standard used for experimental data widely used internationally is the Scientific Metadata Model developed by CCLRC (now STFC). The model includes information at the top level describing the study and the set of investigations i.e. experiment, measurement, simulation etc involved in this study. Then for each investigation it records specific information about the data:

Data holding - A logical hierarchy of the Data Collections and Atomic Data Objects and their directory style grouping. The Data Holding can be considered as the ‘root’ of the data file/object system.

§ Data description - A description of the data kept in this data holding from the data archive perspective. Including information like name, type, status, quality and software.

- Logical description - Reference to a set of logical description fields such as parameter [Name, id, class, units, value, facilities used, range], time period or facility used.

§ Data collection - Data Collections in the hierarchy of data organisation used in this Investigation; much like directories in a file system and they can be nested.

§ Atomic data object - Atomic Data Objects (files, blobs, named selects etc)

§ Related reference - Other Studies/Investigations related to this Data Holding and their type or relationship; e.g. derived from or used by

§ Data holding locator - A locator for addressing the overall Data Holding. (URI of top level directory or data)


How can this complex workflow process that involves several research groups with specialists skills and a variety of tools and techniques be recorded?

An answer to this question may be obtained by looking at the work of our colleagues in Southampton. Some weeks ago Simon Coles and Jeremy Frey visited the OeRC to tell us about their work on electronic lab notebooks. They have been involved in projects such as Smart Tea and CombeChem that deal with the management of laboratory information. Initially they had explored the idea of replicating printed lab-notebooks using tablet interfaces that would capture structured information. These have the benefits of good semantic information. In addition to this, they have experimented with the idea of laboratory blogs that allow recording step by step the process followed allowing discussing the data and providing flexibility and the power of web 2.0 technologies.




Monday, 13 July 2009

EIDCSR website launched











The EIDCSR website was launched last week at http://eidcsr.oucs.ox.ac.uk

This new site will contain information about this JISC funded project including background and methodology, it will aggregate posts from this blog, project bookmarks and will link to different reports, presentations and papers resulting from project activities.

ShareThis