Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

TimeSeriesResult.to_df() turns API nulls into '' and emits a zero-length trailing bucket when filter_time_max lands on a period boundary

Abierto
#578 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
52/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
pandas, python
Área
api, data

Línea de trabajo

Start with result_conversions.create_dataframe and sort_breakdown in timeseries_result.py, then run the reproduction in the issue for boundary and non-boundary filter_time_max values. Trace how API nulls and the final bucket are represented before to_df(). Done means numeric nulls remain usable, zero-length periods are handled consistently, and the filter boundary behavior is documented or defined.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Environment: vortexasdk 1.0.32 (also reproduced on 1.0.29), Python 3.12, pandas 2.3.3, Windows.

A. to_df() replaces nulls with ''. result_conversions.create_dataframe runs pd.DataFrame(data).fillna(""). Any null from the API becomes an empty string and turns value into object dtype, so arithmetic such as df["value"] / 1000 raises TypeError. sort_breakdown in timeseries_result.py also pads missing breakdown entries with "value": "".

B. Zero-length trailing bucket. When filter_time_max falls exactly on a period boundary, for example midnight on the 1st with timeseries_frequency="month", the response has an extra final bucket for a zero-length period with count=0. With a rate unit (bpd) its value is None, so '' after to_df(). With a volume unit (b) it is 0, which reads like a real empty month. Months with genuinely no cargoes return value=0 as expected.

from datetime import datetime
from vortexasdk import CargoTimeSeries, Geographies, Products

def run(geo, prod, end):
    return CargoTimeSeries().search(
        filter_activity="loading_end", filter_origins=geo, filter_products=prod,
        filter_time_min=datetime(2025, 1, 1), filter_time_max=end,
        timeseries_frequency="month", timeseries_unit="bpd",
    ).to_df()

if __name__ == "__main__":
    geo = Geographies().search(term="Rotterdam").to_list()[0].id
    prod = Products().search(term="Diesel/Gasoil", exact_term_match=True).to_list()[0].id
    for end in (datetime(2025, 6, 1), datetime(2025, 6, 30, 23, 59)):
        df = run(geo, prod, end)
        print(end, df["value"].dtype)   # 2025-06-01 -> object, last row count=0 value=''
        print(df.tail(2))               # 2025-06-30 -> float64, last row is the real June

Expected:

  • Nulls preserved as NaN, keeping value numeric.
  • No bucket for a zero-length period, or documented behaviour (ideally a flag) for whether filter_time_max is inclusive.

Related: Products().reference(old_id) silently redirects a rotated product id to its replacement, but CargoTimeSeries with the old id returns all zeros with no error. Could id rotation be documented, or could the timeseries endpoint reject unknown product ids?

Lenguaje dominante
Python
Estrellas
25
Forks
12
Merge medio
20 h 18 min
PR fusionados (30 d)
1

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de VorTECHsa/python-sdk

Todos los issues de VorTECHsa/python-sdk

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.