scikit-learn-contrib/category_encoders

Error handling in inverse_transform is broken

已关闭

#190 创建于 2019年5月7日

 (0 条评论) (0 个反应) (0 位负责人)Python (397 个派生)batch import
bughelp wanted

仓库指标

星标
 (2,322 个星标)
PR 合并指标
 (平均合并 4天 10小时) (30 天内合并 2 个 PR)

描述

Inverse_transform should ideally handle absence of the columns dropped because of drop_invariant=True. But if it is not possible, inverse_transform() should at least return the correct error message instead of just crashing.

Example in form of a parameterized unit test:

    def test_inverse_wrong_feature_count_wit_drop_invariant(self):
        x = [['A', 'B', 'C'], ['D', 'E', 'C'], ['F', 'G', 'C']]  # the last column is constant 
        for encoder_name in {'BaseNEncoder', 'BinaryEncoder', 'OrdinalEncoder', 'OneHotEncoder'}:
            with self.subTest(encoder_name=encoder_name):
                enc = getattr(encoders, encoder_name)(drop_invariant=True)
                transformed = enc.fit_transform(x)

                # run inverse_transform() and check the raised exception text
                with self.assertRaises(ValueError) as cm:
                    enc.inverse_transform(transformed)
                self.assertTrue(str(cm.exception).startswith('Unexpected input dimension'))

贡献者指南