Published August 26, 2021 | Version First version (26 Aug 2021)eng
Dataset Open

Vietic 116 item phylogenetic lexicon

  • 1. University of Sydney
  • 2. Montgomery College

Description

The file is a116 item lexicostatistical dataset for classification of the Vietic languages. The set includes 30 Vietic doculects, Proto-Vietic, plus Khmu and Jahai as out-groups. Included is a listing of the sources, and the NEXUS file with our cognate value assignments, which we created to run on SplitsTree to generate phylograms and NeighborNets. The 116-item list was the outcome of beginning with the Swadesh 100 and 200 lists and reconciling these with the available data with the aim of achieving at least 80% coverage for each lect in the analysis. Procedurally, sources were selected and lexicons aggregated in a spreadsheet, with rows identified with Swadesh 100 and 200 items, subject to semantic and phonological adjustments as we judged necessary. For most of the languages, full coverage of the Swadesh 100 categories was not possible, with 20 or more gaps being common. Some 40 additional categories were added from the Swadesh 200 list, based on the 40 best represented items in the aggregated data, seeking to achieve a 120-item list with at least 100 items coverage for all lects, ultimately settling on 116 items.

Notes

The dataset was created for a paper provisionally entitled "The Vietic Languages: A Phylogenetic Analysis". The paper is submitted for journal publication and a version submitted for presentation at ICAAL9, November 2021. We encourage sharing for the purpose of testing/reproducing results, and augmented or derived studies under Creative Commons Attribution licence.

Files

Files (199.8 kB)

Name Size Download all
md5:913cb924c3c879c9047c9ba67bc374e0
199.8 kB Download