Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

SKLearn Pipeline calculate sensitivity of categorical features for Local DP

Aperta
#389 2 commenti 0 reazioni 1 assegnatario Vedi su GitHub

@grilhami ci sta già lavorando.

Dal 30/11/2021.

Valutazione

Questa issue non è ancora stata valutata.

Descrizione

Type: New Feature :heavy_plus_sign:

Feature Description

The current SKLearn Pipeline noise mechanism "operator" for Local DPassumes that the dataset given contains only all numerical features. This means that noise is calculated on top of the sensitivity calculation on numerical features.

However, most often, datasets also contain categorical features, which requires a different method to calculate the sensitivity. The "operator" should also support categorical features.

This applies to all the noise mechanisms: LaplaceMechanism, GaussianMechanism, and GeometricMechanism.

Note: as far as this issue was created, only LaplaceMechanism has been implemented, so it's a good starting point to start with LaplaceMechanism. Once GeometricMechanism and GaussianMechanism have been implemented, the specifications for categorical feature support are the same.

Additional Context

Preferably, the support for the categorial features would be in the form of parameters for the "operator" class.

For example, in the case of LaplaceMechanism, it would look something like this:

# Set a privacy budget accountant
accountant = BudgetAccountant(10000)

# Set sensitivity function for numerical data
sensitivity = lambda x: (max(x) - min(x))/ (len(x) + 1)

# Set sensitivity function for categorical data
sensitivity_cat = lambda x: ...

# Indecies of the categorical features in the dataset
cat_features = [0, 1, ...]

# Set laplace mechanism with epsilon, sensitivity, and accountant
laplace = LaplaceMechanism(
    epsilon=0.1, 
    sensitivity=sensitivity, 
    accountant=accountant,
    sensitivity_cat=sensitivity_cat,
    cat_features=cat_features
)

# Initialize scaler and naive bayes extimator
scaler = StandardScaler()
nb = GaussianNB()

# Create the pipeline
pipe = Pipeline([('scaler', scaler), ('laplace', laplace), ('nb', nb)])

For more examples, please have look at the notebook example of Laplace Mechanism's implementation.

As starting guidance, please refer to the source code for LaplaceMechanism in here.

Lingua principale
Python
Stelle
550
Fork
142
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di OpenMined/PyDP

Tutte le issue di OpenMined/PyDP

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.