Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

TimeSeriesResult.to_df() turns API nulls into '' and emits a zero-length trailing bucket when filter_time_max lands on a period boundary

未关闭
#578 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
52/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
pandas, python
领域
api, data

调研方向

Start with result_conversions.create_dataframe and sort_breakdown in timeseries_result.py, then run the reproduction in the issue for boundary and non-boundary filter_time_max values. Trace how API nulls and the final bucket are represented before to_df(). Done means numeric nulls remain usable, zero-length periods are handled consistently, and the filter boundary behavior is documented or defined.

由索引模型根据 Issue 内容生成。

描述

Environment: vortexasdk 1.0.32 (also reproduced on 1.0.29), Python 3.12, pandas 2.3.3, Windows.

A. to_df() replaces nulls with ''. result_conversions.create_dataframe runs pd.DataFrame(data).fillna(""). Any null from the API becomes an empty string and turns value into object dtype, so arithmetic such as df["value"] / 1000 raises TypeError. sort_breakdown in timeseries_result.py also pads missing breakdown entries with "value": "".

B. Zero-length trailing bucket. When filter_time_max falls exactly on a period boundary, for example midnight on the 1st with timeseries_frequency="month", the response has an extra final bucket for a zero-length period with count=0. With a rate unit (bpd) its value is None, so '' after to_df(). With a volume unit (b) it is 0, which reads like a real empty month. Months with genuinely no cargoes return value=0 as expected.

from datetime import datetime
from vortexasdk import CargoTimeSeries, Geographies, Products

def run(geo, prod, end):
    return CargoTimeSeries().search(
        filter_activity="loading_end", filter_origins=geo, filter_products=prod,
        filter_time_min=datetime(2025, 1, 1), filter_time_max=end,
        timeseries_frequency="month", timeseries_unit="bpd",
    ).to_df()

if __name__ == "__main__":
    geo = Geographies().search(term="Rotterdam").to_list()[0].id
    prod = Products().search(term="Diesel/Gasoil", exact_term_match=True).to_list()[0].id
    for end in (datetime(2025, 6, 1), datetime(2025, 6, 30, 23, 59)):
        df = run(geo, prod, end)
        print(end, df["value"].dtype)   # 2025-06-01 -> object, last row count=0 value=''
        print(df.tail(2))               # 2025-06-30 -> float64, last row is the real June

Expected:

  • Nulls preserved as NaN, keeping value numeric.
  • No bucket for a zero-length period, or documented behaviour (ideally a flag) for whether filter_time_max is inclusive.

Related: Products().reference(old_id) silently redirects a rotated product id to its replacement, but CargoTimeSeries with the old id returns all zeros with no error. Could id rotation be documented, or could the timeseries endpoint reject unknown product ids?

主要语言
Python
星标
25
派生
12
平均合并
20 小时 18 分钟
30 天内合并 PR
1

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

VorTECHsa/python-sdk 的其他 Issue

查看 VorTECHsa/python-sdk 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。