scikit-learn-contrib/hdbscan

Introduce sample weighting in HDBSCAN

Ouverte

#148 ouverte le 5 déc. 2017

 (18 commentaires) (10 réactions) (0 personne assignée)Jupyter Notebook (535 forks)github user discovery
help wantednew feature

Métriques du dépôt

Stars
 (3 137 étoiles)
Métriques de merge PR
 (Aucune PR mergée en 30 j)

Description

Like for DBSCAN it would be very useful to also have the possibility to use a "sample_weight" parameter when performing HDBSCAN clustering trough the fit method. This to take into account possible presence of data containing element duplicates.

DBSCAN is just using sample_weight in the following way:

n_neighbors = np.array([np.sum(sample_weight[neighbors]) for neighbors in neighborhoods])

do you know if this is already planned at same time or not?

Guide contributeur