bug: get_dataset_run / get_dataset_runs / delete_dataset_run are unusable on Langfuse v4
I maintainer di solito rispondono entro 1 giorno
@hassiebp ci sta già lavorando.
Dal 27/9/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Describe the bug
The three dataset-run read helpers on the Langfuse client still call the pre-v4
dataset-run endpoints:
| Method | Defined at | Request |
|---|---|---|
Langfuse.get_dataset_run() |
langfuse/_client/client.py:2532 |
GET /api/public/datasets/{name}/runs/{run_name} |
Langfuse.get_dataset_runs() |
langfuse/_client/client.py:2557 |
GET /api/public/datasets/{name}/runs |
Langfuse.delete_dataset_run() |
langfuse/_client/client.py:2588 |
DELETE /api/public/datasets/{name}/runs/{run_name} |
On a Langfuse v4 deployment these paths are rejected, so all three raise
NotFoundError: 404 with the body
This endpoint is not available on deployments running in Langfuse v4 events_only mode.
Learn more about Langfuse v4 at: https://langfuse.com/docs/v4
The same SDK already ships the v4 read client for exactly this data —
client.api.experiments.list() and client.api.experiments.list_items()
(langfuse/api/experiments/raw_client.py:87 and :253) — but no helper uses it.
run_experiment() makes the gap directly visible: it writes its run through
POST /api/public/dataset-run-items, which is accepted on v4, returns a populated
dataset_run_id, and hands the caller a dataset_run_url. That run cannot then be
fetched back with the SDK's own read method.
Steps to reproduce
Against a v4 deployment (LANGFUSE_HOST pointing at it):
import datetime as dt
import time
import uuid
from langfuse import Langfuse
client = Langfuse()
dataset = f"repro-dataset-run-read-{uuid.uuid4().hex[:6]}"
client.create_dataset(name=dataset)
client.create_dataset_item(dataset_name=dataset, input={"q": "hello"}, id=f"{dataset}-item-1")
result = client.run_experiment(
name="repro-experiment",
run_name=f"{dataset}-run",
data=client.api.dataset_items.list(dataset_name=dataset).data,
task=lambda *, item, **kwargs: "answer",
)
print("run_experiment -> dataset_run_id:", result.dataset_run_id)
try:
client.get_dataset_run(dataset_name=dataset, run_name=f"{dataset}-run")
print("get_dataset_run -> OK")
except Exception as exc:
print("get_dataset_run ->", type(exc).__name__, "status:", getattr(exc, "status_code", None))
print(" body:", getattr(exc, "body", None))
# The events pipeline is asynchronous, so poll briefly for the run to become readable.
deadline = time.time() + 30
while True:
experiments = client.api.experiments.list(
from_start_time=dt.datetime.now(dt.timezone.utc) - dt.timedelta(hours=1),
name=f"{dataset}-run",
)
if experiments.data or time.time() > deadline:
break
time.sleep(2)
print("api.experiments.list ->", len(experiments.data), "run(s)")
for experiment in experiments.data:
print(" ", experiment.id, experiment.name, "items:", experiment.item_count)
client.flush()
Actual behavior
run_experiment -> dataset_run_id: d591fd225d7edec7
get_dataset_run -> NotFoundError status: 404
body: {'message': 'This endpoint is not available on deployments running in Langfuse v4 events_only mode. Learn more about Langfuse v4 at: https://langfuse.com/docs/v4'}
api.experiments.list -> 1 run(s)
d591fd225d7edec7 repro-dataset-run-read-6ca29b-run items: 1
The run exists and is readable — just not through the documented helper. Note that
dataset_run_id returned by run_experiment() and the id returned by
api.experiments.list() are the same value, so the data is not lost, only unreachable
from the helper that the SDK exposes for it.
Expected behavior
Either route these helpers through the v4 read path when the deployment is v4
(api.experiments.list / list_items, mapping by id/name and dataset_id), or
state in their docstrings that they are unavailable on v4 and point at
client.api.experiments. Today the docstrings say nothing about v4, so the only signal
a user gets is a 404.
If the first route is taken, the mapping is not a drop-in replacement. Measured against a
v4 deployment for a run written by run_experiment():
DatasetRunItem field |
v4 source |
|---|---|
dataset_run_id / dataset_run_name |
ExperimentItem.experimentId / .experimentName |
dataset_item_id |
ExperimentItem.experimentItemId |
trace_id |
ExperimentItem.traceId |
observation_id |
ExperimentItem.id — the experiment item id is the task span id |
created_at / updated_at |
ExperimentItem.startTime / .endTime |
id |
no equivalent — the legacy row id (a UUID) is not part of the experiment item |
and for DatasetRun: id/name/description/metadata/dataset_id come from
Experiment, but dataset_name has no counterpart (only datasetId is returned, so it
needs a datasets.list() lookup), metadata requires fields=metadata, and
startTime/endTime are clipped to the requested fromStartTime window rather than being
the run's createdAt/updatedAt. from_start_time is required and pagination is
cursor-based, whereas the current helpers are page-based.
Additional context
handle_fern_exceptionin these helpers only logs; the raisedNotFoundErrorkeeps the
generated dataclassstr()(it starts withheaders: {...}), so the useful server
message is only visible viaexc.body.- This is the SDK-side counterpart of
langfuse/langfuse#16835, and the two are
complementary rather than duplicates. #16835 reports that on a non-events_only
deployment the legacyGET /api/public/datasets/{name}/runsstill returns runs created
viaPOST /dataset-run-items, while the Experiments UI andGET /api/public/experiments
do not show them; the open PRlangfuse/langfuse#16859addresses that side. Here the
deployment is events_only: the legacy read path is rejected outright (404), while
GET /api/public/experimentsreturns the run correctly.#16859touches only
web/,worker/andpackages/— nolangfuse/files. - The rejection itself is intentional on the server side
(langfuse/langfuse#15313"reject dataset-run-item surface in events_only mode",
#15526"do not read dataset-runs in events_only mode"), which is why this looks like
an SDK-alignment issue rather than a server regression.
- Lingua principale
- Python
- Stelle
- 495
- Fork
- 358
- Merge medio
- 16h 54m
- PR unite (30g)
- 20
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di langfuse/langfuse-python
-
bug: GCS upload detection in MediaManager uses substring match on the full URLForse già presa @hassiebp l’ha presa oggi. Apertabug feat-multimodal-media sdk-python
langfuse/langfuse-python#1913 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
mask is not applied to create_dataset_item or create_score(comment=)Forse già presa @hassiebp l’ha presa 8 giorni fa. Apertabug compliance feat-data-masking sdk-python security
langfuse/langfuse-python#1896 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Scores bypass sample_rate since v4Forse già presa @hassiebp l’ha presa 13 giorni fa. Apertafeat-scores sdk-python unconfirmed-bug
langfuse/langfuse-python#1890 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
batch_evaluation fails on self-hosted v4 events_only deployments (uses unavailable v3 read endpoints)Forse già presa @hassiebp l’ha presa 26 giorni fa. Apertabug feat-evals sdk-python
langfuse/langfuse-python#1861 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
[HTTPXodus] Consider migrating from httpx to httpx2 (the actively maintained fork)Forse già presa @hassiebp l’ha presa 27 giorni fa. Apertaimprovement sdk-python
langfuse/langfuse-python#1856 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di langfuse/langfuse-python
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
kornia/kornia#5263 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
approved correction metadata
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
acl-org/acl-anthology#10133 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
BasedHardware/omi#20084 ·
I maintainer di solito rispondono entro 1 giorno
-
bug needs-acceptance wg/evaluation-quality
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
vllm-project/semantic-router#4424 ·
I maintainer di solito rispondono entro 1 giorno