scikit-learn-contrib/hdbscan

Introduce sample weighting in HDBSCAN

Offen

#148 geöffnet am 05.12.2017

 (18 Kommentare) (10 Reaktionen) (0 zugewiesene Personen)Jupyter Notebook (535 Forks)github user discovery
help wantednew feature

Repository-Metriken

Stars
 (3.137 Sterne)
PR-Merge-Metriken
 (Keine gemergten PRs in 30 T)

Beschreibung

Like for DBSCAN it would be very useful to also have the possibility to use a "sample_weight" parameter when performing HDBSCAN clustering trough the fit method. This to take into account possible presence of data containing element duplicates.

DBSCAN is just using sample_weight in the following way:

n_neighbors = np.array([np.sum(sample_weight[neighbors]) for neighbors in neighborhoods])

do you know if this is already planned at same time or not?

Contributor Guide