Discriminant analysis of interval data: an assessment of parametric and distance-based approaches

A. Pedro Duarte Silva*, Paula Brito

*Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

22 Citations (Scopus)

Abstract

Building on probabilistic models for interval-valued variables, parametric classification rules, based on Normal or Skew-Normal distributions, are derived for interval data. The performance of such rules is then compared with distancebased methods previously investigated. The results show that Gaussian parametric approaches outperform Skew-Normal parametric and distance-based ones in most conditions analyzed. In particular, with heterocedastic data a quadratic Gaussian rule always performs best. Moreover, restricted cases of the variance-covariance matrix lead to parsimonious rules which for small training samples in heterocedastic problems can outperform unrestricted quadratic rules, even in some cases where the model assumed by these rules is not true. These restrictions take into account the particular nature of interval data, where observations are defined by both MidPoints and Ranges, which may or may not be correlated. Under homocedastic conditions linear Gaussian rules are often the best rules, but distance-based methods may perform better in very specific conditions.
Original languageEnglish
Pages (from-to)516-541
Number of pages26
JournalJournal of Classification
Volume32
Issue number3
DOIs
Publication statusPublished - 1 Oct 2015

Keywords

  • Discriminant analysis
  • Interval data
  • Parametric modelling of interval data
  • Symbolic data analysis

Fingerprint

Dive into the research topics of 'Discriminant analysis of interval data: an assessment of parametric and distance-based approaches'. Together they form a unique fingerprint.

Cite this