scikit-learn-contrib/hdbscan
Introduce sample weighting in HDBSCAN
Aberta
#148 aberto em 5 de dez. de 2017
help wantednew feature
Métricas do repositório
- Stars
- (3.138 estrelas)
- Métricas de merge de PR
- (Nenhuma PRs mesclada em 30d)
Description
Like for DBSCAN it would be very useful to also have the possibility to use a "sample_weight" parameter when performing HDBSCAN clustering trough the fit method. This to take into account possible presence of data containing element duplicates.
DBSCAN is just using sample_weight in the following way:
n_neighbors = np.array([np.sum(sample_weight[neighbors]) for neighbors in neighborhoods])
do you know if this is already planned at same time or not?