Showing posts with label data management. Show all posts
Showing posts with label data management. Show all posts

Tuesday, 13 July 2010

Open Repositories 2010 in Madrid



This years the Open Repository Conference 2010 was held in Madrid organised by the he Spanish Foundation for Science and Technology FECYT and UNED, a Spanish public university that provides distance education.

Many of the talks discussed issues around research data and digital repositories. In the initial keynote, Prof. David de Roure emphasized the importance of capturing the research data but also the methods behind the data. In the future repositories will have a role in managing knowledge packs made of data, metadata, workflows, articles, presentations, results, etc.

The conference had a strong pressence from activities using the eSciDoc repository system based on Fedora. The BW eLab project uses this infrastructure to provide access to remote laboratory instruments as well as to manage the experimental data generated in the labs. During the workflow process eSync Daemon is used to monitor the file system of the computer connected to the instruments . The daemon replicates the new files and sends to a deposit where metadata is extracted to them deposit data and metadata in eSciDoc. A similar synchronization is used in the BRIL Project to monitor researcher's own desktop to capture as much data and metadata as possible.

Another interesting talk presented an open source repository for medical scientific research known as MIDAS. The system is used for the Insight Journal which provides open-access to articles, data, code, and reviews with an archive which hosts public collections of image datasets such as MRIs.

Other repository frameworks included Hydra, a collaboration between the Universities of Hull, Stanford and Virginia, that uses a technical architecture based on Fedora with a toolkit of reusable components that can assist with a range of content management, access and preservation. The University of Hull IR provides a Hydra use case.

Microsoft announced the release of v2.0 of their repository platform Zentity which makes use of the Open Data Protocol and uses Pivot for visualising and organising the data (see this example of pivot in action). The installation support services such as OAI-ORE and SWORD.

In the national approaches session the results of the Australian institutional research repository data readiness surveys 2010 were presented. Although repository managers are aware of ANDS and its services, there is little use of them and less than half of respondents were planning to incoorporate data in their repositories.

This has truly been a rewarding and stimulating conference.

Thursday, 27 May 2010

Digital Curation Centre Workshop at Oxford on the 16th June, 2010 – How to Manage Research Data

I am very pleased to announce that the Digital Curation Centre will be paying a visit to Oxford on the 16th June to present a workshop on managing research data. The workshop is aimed primarily at researchers interested in bidding for funding for projects with a data output, although it should also appeal to those who assist and support research activities and who would like to find out more about the challenges of data curation.

Although the workshop will obviously be of relevance to those interested in either the Sudamih or EIDCSR projects, it will not focus exclusively on a particular academic discipline but should be useful across the board. Sessions will include: the roles and responsibilities associated with conceptualising, creating and managing research data during the life of a project; the responsibilities associated with the longer-term management of research data after a project has ended; developing a data management plan; and preparing data for long-term curation and re-use.

The workshop is free for members of the University of Oxford, £50 for non-members.

Anyone interested in attending the workshop should register at http://www.dcc.ac.uk/training/digital-curation-101/digital-curation-101-lite-oxford

Friday, 23 April 2010

Report on the 'Institutional Policy and Guidance for Research Data' Workshop

'How to share expertise? Where to get advice'? Just two of the questions institutions need to address in their research data management policies according to Paul Taylor of Melbourne University. On the 29th March 2010, the place for advice and sharing expertise was the EIDCSR Institutional Policy Workshop in Oxford.

A significant part of the ‘Embedding Institutional Data Curation Services in Research’ Project has been to start developing an Institutional research data management policy for the University of Oxford, so this workshop offered us a chance both to say how things were going and find out the lessons learnt from others farther down the road.

The University of Melbourne has been grappling with the issues for some time already, and we were lucky enough to be joined by several of their representatives via videoconference. Indeed, given how close we were to not being joined by their representatives due to the videoconferencing equipment, ‘lucky’ is the operative word. Paul Taylor stressed that any effective policy needs to be implementable. This involves getting the researchers themselves involved in the development process and offering somewhere where people can go for information and advice. Compliance becomes easier the more central services exist, leaving researchers to do the research.

Another university which has already done a lot of work on data curation is Southampton, and Kenji Takeda introduced their long-term ambitions. The unfortunate incident a few years back when Southhampton’s Mountbatton building burnt down led to claims against the lost research data, so this has perhaps focussed minds more there than in other institutions. A cost-benefits analysis is now being undertaken which should help institutions better appreciate the value of their data outputs. Furthermore, they are looking to make data management courses compulsory. Herding academics into classrooms sounds ambitious, but there was a general sense from the workshop that without training there was little chance of persuading researchers to adopt best practices.

Jeff Hayward, from the University of Edinburgh emphasised that when it comes to data curation it is better to identify the opportunities than enumerate the problems, but then failed to ignore the various ‘inhibitors’ to good data management. “Researchers want data management, but don’t want to do it.” Quite. Nevertheless, Edinburgh are bravely forging ahead, setting up an experimental ‘DataShare’ service and adapting the Digital Curation Centre’s 101 training, with the intention of making it compulsory for doctoral students.

Finally, David McAllister of the BBSRC explained data management policies from a Research Council’s point of few – clearly a key driver for institutional policies.

Perhaps the last word should go to Jeff Hayward who concluded the panel questions session by indicating that the world would actually be a happier place if there were fewer data repositories. Individual universities should really act as repositories of last resort, but the onus must be on them to guarantee that research data is not lost or rendered inaccessible.

For a more complete report on the workshop, plus the various sets of slides used by the presenters, go to the EIDCSR Project website: http://eidcsr.oucs.ox.ac.uk/policy_workshop.xml

Friday, 18 December 2009

Scientific data repositories workshop in Barcelona


A couple of weeks ago I was invited to talk at an incredibly inspiring event organized by the Centre de Supercomputació de Catalunya titled "Repositorios de datos cientificos" under their Jornadas Catalanas de Supercomputació.

We had an extraordinary day with a fantastic group of speakers that discussed issues around supporting researchers with their data management as well as disciplinary perspectives provided by real researchers.

The whole event was filmed and is available
online (for those who speak spanish!) and I also got interviewed and filmed for online publication known as Global Talent, you can also see this video (again in spanish!).

Monday, 12 October 2009

"Science these days has basically turned into a data-management problem"

The New York Times has an article about future scientists' ability to manage the large amounts of digital data being generated and how the likes of IBM or Google are trying to help, "Training to Climb an Everest of Digital Data", http://www.nytimes.com/2009/10/12/technology/12data.html. IBM and Google are contributing tools, computational power and access to large-scale datasets. It was actually two years ago this month that Google and IBM announced their partnership to provide universities with dedicated cluster computing resources, open source software, a dedicated website for collaboration, and a Creative Commons-licensed curriculum. In April this year the NSF funded projects at 14 US universities to take advantage of the IBM/Google Cloud Computing University Initiative. The New York Times article highlights some of these projects. The emphasis is certainly on the massive -- big compute clusters, big datasets -- and on data analysis. Not much though on the ongoing management of, access to, and preservation of data, even if Professor Jimmy Lin (University of Maryland) is quoted as saying, “Science these days has basically turned into a data-management problem”.

Friday, 5 June 2009

Data imperative event

The data imperative event organized the RLUK/SCONUL Task Force on e-Research was held on Wednesday 3 June in Oxford with support from the Oxford e-Research Centre, RLUK, SCONUL and RIN.

This was an excellent opportunity to confirm the extraordinary interest of librarians in this area as well as the difficulty to clarify their role and where the necessary funding comes from to allow addressing the challenge. Chris Keene shares his notes of the event from his blog and RLUK will be shortly making the presentations available.

In the mean time you can access Prof. Paul Jeffrey's introduction to the workshop and my talk describing Oxford's recent work in this area.


ShareThis