scikit-learn-contrib/hdbscan

Introduce sample weighting in HDBSCAN

开放

#148 创建于 2017年12月5日

 (18 条评论) (10 个反应) (0 位负责人)Jupyter Notebook (535 个派生)github user discovery
help wantednew feature

仓库指标

星标
 (3,137 个星标)
PR 合并指标
 (30 天内没有已合并 PR)

描述

Like for DBSCAN it would be very useful to also have the possibility to use a "sample_weight" parameter when performing HDBSCAN clustering trough the fit method. This to take into account possible presence of data containing element duplicates.

DBSCAN is just using sample_weight in the following way:

n_neighbors = np.array([np.sum(sample_weight[neighbors]) for neighbors in neighborhoods])

do you know if this is already planned at same time or not?

贡献者指南