Columns and DataType Not Explicitly Set on line 62 of data_utils.py
Maintainer antworten meist innerhalb von 5 Tagen
Dieses Issue hat noch niemand übernommen.
Bewertung
Dieses Issue wurde noch nicht bewertet.
Beschreibung
Hello!
I found an AI-Specific Code smell in your project.
The smell is called: Columns and DataType Not Explicitly Set
You can find more information about it in this paper: https://dl.acm.org/doi/abs/10.1145/3522664.3528620.
According to the paper, the smell is described as follows:
| Problem | If the columns are not selected explicitly, it is not easy for developers to know what to expect in the downstream data schema. If the datatype is not set explicitly, it may silently continue the next step even though the input is unexpected, which may cause errors later. The same applies to other data importing scenarios. |
|---|---|
| Solution | It is recommended to set the columns and DataType explicitly in data processing. |
| Impact | Readability |
Example:
### Pandas Column Selection
import pandas as pd
df = pd.read_csv('data.csv')
+ df = df[['col1', 'col2', 'col3']]
### Pandas Set DataType
import pandas as pd
- df = pd.read_csv('data.csv')
+ df = pd.read_csv('data.csv', dtype={'col1': 'str', 'col2': 'int', 'col3': 'float'})
You can find the code related to this smell in this link: https://github.com/GoogleCloudPlatform/python-docs-samples/blob/161448f276cd6fbe697c83e2ec55c4828e9143fa/people-and-planet-ai/timeseries-classification/data_utils.py#L52-L72.
I also found instances of this smell in other files, such as:
File: https://github.com/GoogleCloudPlatform/python-docs-samples/blob/master/composer/2022_airflow_summit/data_analytics_process_expansion.py#L140-L150 Line: 145
.
I hope this information is helpful!
- Vorherrschende Sprache
- Jupyter Notebook
- Sterne
- 8.1k
- Forks
- 6.7k
- Ø Merge
- 4 T. 12 Std.
- Gemergte PRs (30 T.)
- 9
Entwicklungsumgebung
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus GoogleCloudPlatform/python-docs-samples
-
samples
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
GoogleCloudPlatform/python-docs-samples#14609 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
samples
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 92/100
GoogleCloudPlatform/python-docs-samples#14610 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
samples
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
GoogleCloudPlatform/python-docs-samples#14611 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
chore(generative_ai) Update model references for generative_ai samplesEvtl. wieder frei @XrossFox hat das vor 155 Tagen übernommen, und es ist kein Pull Request offen. Offensamples
GoogleCloudPlatform/python-docs-samples#14117 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 5 Tagen
-
Fix Pipeline dependency issues for geospatial-classification/serving_appEvtl. wieder frei @XrossFox hat das vor 167 Tagen übernommen, und es ist kein Pull Request offen. Offensamples
GoogleCloudPlatform/python-docs-samples#14091 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 5 Tagen