Project title:

Scientific image processing within the NFFA-EUROPE Data Repository


Rossella Aversa

Defense Year: 2017-2018

This thesis is embedded within the NFFA-EUROPE project, that aims to setup
the first overarching Information and Data management Repository Platform
(IDRP) for the nanoscience community.

The main goal of the IDRP developement is to semiautomatize the harvest-
ing of scientific data coming from many different instruments among NFFA Eu-
ropean facilities, and identify the correct metadata to allow them to be search-
able accordingly to the fair principle.

For our analysis, we selected data from a specific instrument, the Scanning
Electron Microscope (SEM). The reasons for this choice are many:

  • the instrument is available at the CNR-IOM in Trieste, so we could discuss with scientists about their needs and have fast feedback;
  • a significant amount of SEM images (roughly 150,000) have been provided to us for testing purposes;
  • the plugin for the SEM has already been tested and is available for this instrument;
  • ten out of the twenty NFFA European partners have a SEM facility, so our work can be exported to a sizeable part of the community.

The specific goals of this thesis are the following:

  • explore machine learning algorithms to classify scientific images coming from the SEM instrument;
  • once the images have been categorized, enrich the plugin with the ingestion of new metadata. In this way, a search engine can be used to find the relevant images through a semantic search on the database;
  • setup a computational and storage infrastructure to process a massive amount of such images efficiently;

The achievements of our work are the following:

  • a supervised machine learning algorithm has been successfully implemented
  • and tested;
  • algorithms and tools for massive data processing have been identified, set
  • up, optimized, and finally benchmarked;
  • many tools for simplifying the setting up of complex computational and storage infrastructures have been extensively used. In particular, we em-
  • ployed Docker and OpenStack;
  • some estimates of the time needed to process the image data on different
  • environments have been provided.

The thesis is organized as follows: in Chapter 1 we will present the NFFA-
EUROPE project and we will frame our work within the NFFA-EUROPE; in Chapter 2 we will offer an overview of machine learning techniques and tools
used in thoughout our work; in Chapter 3 we will explain the main Spark
concepts and how we set the environment up to employ it; scientific and technical
results are discusses in Chapter 4 and Chapter 5, respectively; finally, in Chapter
6 we will present our conclusions.

Thesis not available.