Parallel extraction of Regions-of-Interest from social media data

Files in This Item:
Access to this item has been restricted by the copyright holder until:2021-01-02
File Description SizeFormat 
insight_publication.pdf8.78 MBAdobe PDF    Request a copy
Title: Parallel extraction of Regions-of-Interest from social media data
Authors: Belcastro, LorisKechadi, TaharMarozzo, FabrizioPastore, Lucaet al.
Permanent link: http://hdl.handle.net/10197/11699
Date: 2-Jan-2020
Online since: 2020-11-13T08:40:46Z
Abstract: Geotagged data gathered from social media can be used to discover places‐of‐interest (PoIs) that have attracted many visitors. Since a PoI is generally identified by geographical coordinates of a single point, it is hard to match it with people trajectories. Therefore, we define an area, called region‐of‐interest (RoI), represented by the boundaries of a PoI. The main goal of this study is to discover RoIs from PoIs using spatial data mining techniques. In this paper, we propose a new parallel method for extracting RoIs from social media datasets. It consists of two main steps: (i) automatic keyword extraction and data grouping and (ii) parallel RoI extraction. The first step extracts keywords identifying the PoIs; these keywords are used to group social media items according to the places they refer to. The second step uses a Parallel Clustering Approach (ParCA) of spatial dataset to identify RoIs. ParCA exploits a parallel execution of DBSCAN on subsets of data to generate subclusters on each processing node and then merge overlapping subclusters to form global clusters. ParCA was implemented using the MapReduce model. Experiments performed over a set of PoIs in the city of Rome using social media data show that our approach is highly scalable and reaches an accuracy of 79% in detecting RoIs. On a parallel computer with 50 cores, we obtained a speedup of 52 by processing large datasets divided into 32 splits, compared with the execution time registered when each dataset is not partitioned.
Funding Details: Science Foundation Ireland
metadata.dc.description.othersponsorship: Insight Research Centre
Type of material: Journal Article
Publisher: Wiley
Journal: Concurrency and Computation: Practice and Experience
Copyright (published version): 2020 Wiley
Keywords: Parallel clusteringRegions-of-interestRol miningScalabilitySocial media analysis
DOI: 10.1002/cpe.5638
Language: en
Status of Item: Peer reviewed
Appears in Collections:Computer Science Research Collection
Insight Research Collection

Show full item record

Page view(s)

153
Last Week
67
Last month
checked on Nov 30, 2020

Download(s)

36
checked on Nov 30, 2020

Google ScholarTM

Check

Altmetric


This item is available under the Attribution-NonCommercial-NoDerivs 3.0 Ireland. No item may be reproduced for commercial purposes. For other possible restrictions on use please refer to the publisher's URL where this is made available, or to notes contained in the item itself. Other terms may apply.