Deterministic coresets for k-Means of big sparse data

Citation DataAlgorithms, ISSN: 1999-4893, Vol: 13, Issue: 4

Publication Year2020

7
Citations
0
Usage
6
Captures
0
Mentions
30
Social Media

Metric Options: Counts1 Year3 Year

Metrics Details

Citations
7
- Citation Indexes
  7
Captures
6
- Readers
  6
Social Media
30
- Shares, Likes & Comments
  30

Article Description

Let P be a set of n points in R, k ≥ 1 be an integer and ɛ ∈ (0, 1) be a constant. An ɛ-coreset is a subset C ⊆ P with appropriate non-negative weights (scalars), that approximates any given set Q ⊆ R of k centers. That is, the sum of squared distances over every point in P to its closest point in Q is the same, up to a factor of 1 ±- ɛ to the weighted sum of C to the same k centers. If the coreset is small, we can solve problems such as k-means clustering or its variants (e.g., discrete k-means, where the centers are restricted to be in P, or other restricted zones) on the small coreset to get faster provable approximations. Moreover, it is known that such coreset support streaming, dynamic and distributed data using the classic merge-reduce trees. The fact that the coreset is a subset implies that it preserves the sparsity of the data. However, existing such coresets are randomized and their size has at least linear dependency on the dimension d. We suggest the first such coreset of size independent of d. This is also the first deterministic coreset construction whose resulting size is not exponential in d. Extensive experimental results and benchmarks are provided on public datasets, including the first coreset of the EnglishWikipedia using Amazon's cloud.

Bibliographic Details

DOI10.3390/a13040092

URL IDhttp://www.scopus.com/inward/record.url?partnerID=HzOxMe3b&scp=85084920672&origin=inward; http://dx.doi.org/10.3390/a13040092; https://www.mdpi.com/1999-4893/13/4/92; https://dx.doi.org/10.3390/a13040092

AUTHOR(S)

Artem Barger; Dan Feldman

PUBLISHER(S)

MDPI AG

TAG(S)

Mathematics; Computer Science

Provide Feedback

Have ideas for a new metric? Would you like to see something else here?Let us know