Nearest-neighbor chain algorithm: Difference between revisions

Browse history interactively

← Previous edit

Content deleted Content added

VisualWikitext

Revision as of 06:09, 16 May 2022 edit Bruce1ee (talk \| contribs) Autopatrolled, Extended confirmed users, Pending changes reviewers, Rollbackers 302,074 edits m fixed lint errors – file options; size is ignored when using frame ← Previous edit		Latest revision as of 12:31, 2 July 2025 edit undo E992481 (talk \| contribs) 69 edits Link suggestions feature: 3 links added. Tags: Visual edit Newcomer task Suggested: add links
(5 intermediate revisions by 4 users not shown)
Line 9: ==Background== [[File:Hierarchical clustering diagram.png\|thumb\|upright=1.35\|A hierarchical clustering of six points. The points to be clustered are at the top of the diagram, and the nodes below them represent clusters.]] Many problems in [[data analysis]] concern [[Cluster analysis\|clustering]], grouping data items into clusters of closely related items. [[Hierarchical clustering]] is a version of cluster analysis in which the clusters form a hierarchy or tree-like structure rather than a strict partition of the data items. In some cases, this type of clustering may be performed as a way of performing cluster analysis at multiple different scales simultaneously. In others, the data to be analyzed naturally has an unknown tree structure and the goal is to recover that structure by performing the analysis. Both of these kinds of analysis can be seen, for instance, in the application of hierarchical clustering to [[Taxonomy (biology)\|biological taxonomy]]. In this application, different living things are grouped into clusters at different scales or levels of similarity ([[Taxonomic rank\|species, genus, family, etc]]). This analysis simultaneously gives a multi-scale grouping of the organisms of the present age, and aims to accurately reconstruct the [[branching process]] or [[Phylogenetic tree\|evolutionary tree]] that in past ages produced these organisms.<ref>{{citation \| last = Gordon \| first = Allan D. \| editor1-last = Arabie \| editor1-first = P. Line 36: \| arxiv = cs.DS/9912014 \| issue = 1 \| journal = J. ACM Journal of Experimental Algorithmics \| pages = 1–23 \| publisher = ACM Line 42: \| url = http://www.jea.acm.org/2000/EppsteinDynamic/ \| volume = 5 \| year = 2000\| doi = 10.1145/351827.351829 \| bibcode = 1999cs.......12014E \| s2cid = 1357701 }}.</ref><ref name="day-edels">{{citation \| last1 = Day \| first1 = William H. E. \| last2 = Edelsbrunner \| first2 = Herbert \| author2-link = Herbert Edelsbrunner Line 52: \| url = http://www.cs.duke.edu/~edels/Papers/1984-J-05-HierarchicalClustering.pdf \| volume = 1 \| year = 1984\| s2cid = 121201396 ~~\| year = 1984~~}}.</ref> The nearest-neighbor chain algorithm uses a smaller amount of time and space than the greedy algorithm by merging pairs of clusters in a different order. In this way, it avoids the problem of repeatedly finding closest pairs. Nevertheless, for many types of clustering problem, it can be guaranteed to come up with the same hierarchical clustering as the greedy algorithm despite the different merge order.<ref name="murtagh-tcj"/> ==The algorithm== Line 70 ⟶ 71: \| editor1-last = Abello \| editor1-first = James M. \| editor2-last = Pardalos \| editor2-first = Panos M. \| editor3-last = Resende \| editor3-first = Mauricio G. C. \| editor3-link = Mauricio Resende \| contribution = Clustering in massive data sets \| isbn = 978-1-4020-0489-6 Line 199 ⟶ 200: ===Centroid distance=== Another distance measure commonly used in agglomerative clustering is the distance between the centroids of pairs of clusters, also known as the weighted group method.<ref name="mirkin"/><ref name="lance-williams"/> It can be calculated easily in constant time per distance calculation. However, it is not reducible. For instance, if the input forms the set of three points of an [[equilateral triangle]], merging two of these points into a larger cluster causes the inter-cluster distance to decrease, a violation of reducibility. Therefore, the nearest-neighbor chain algorithm will not necessarily find the same clustering as the greedy algorithm. Nevertheless, {{harvtxt\|Murtagh\|1983}} writes that the nearest-neighbor chain algorithm provides "a good [[heuristic]]" for the centroid method.<ref name="murtagh-tcj"/> A different algorithm by {{harvtxt\|Day\|Edelsbrunner\|1984}} can be used to find the greedy clustering in {{math\|''O''(''n''<sup>2</sup>)}} time for this distance measure.<ref name="day-edels"/>