alteryx/evalml

Data Health: Sparsity Refactor

Open

#1,801 opened on Feb 9, 2021

 (2 comments) (1 reaction) (1 assignee)Python (93 forks)auto 404
enhancementgood first issue

Repository metrics

Stars
 (852 stars)
PR merge metrics
 (PR metrics pending)

Description

Currently, the sparsity score is calculated using a discrete count of values as a threshold. It might make more sense to refactor the SparsityDataCheck to use a relative threshold instead of a fixed count threshold.

SparsityDataCheck.sparsity_score(col, relative_count_threshold=0.10):
    <implementation>

Here the new relative_count_threshold should be a percentage of the total length of the column.

Contributor guide