Testing Center custom `string_comparison` evaluation errors on the JSONPath the server itself generates from `$.generatedData.outcome`
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 45/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Ambito
- cli, testing-qa
Direzione di ricerca
Non sono indicati file sorgente o test. Inizia riproducendo la valutazione personalizzata di string_comparison con il riferimento documentato $.generatedData.outcome e verifica come il server riscrive e analizza quel percorso. Il lavoro è completato quando il riferimento viene valutato correttamente, viene associato a un percorso analizzabile oppure viene rifiutato in fase di creazione con un errore chiaro.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
CLI Version:
@salesforce/cli/2.149.9
Architecture:
win32-arm64
Node Version:
node-v24.14.1
Plugin Version:
@oclif/plugin-autocomplete 3.2.56 (core)
@oclif/plugin-commands 4.1.63 (core)
@oclif/plugin-help 6.2.58 (core)
@oclif/plugin-not-found 3.2.93 (core)
@oclif/plugin-plugins 5.4.87 (core)
@oclif/plugin-search 1.2.54 (core)
@oclif/plugin-update 4.7.59 (core)
@oclif/plugin-version 2.2.57 (core)
@oclif/plugin-warn-if-update-available 3.1.73 (core)
@oclif/plugin-which 3.2.61 (core)
@salesforce/cli 2.149.9 (core)
agent 2.0.5 (user)
apex 4.1.0 (core)
api 2.0.9 (core)
auth 5.0.6 (core)
code-analyzer 5.12.0 (user)
data 5.1.5 (core)
deploy-retrieve 4.1.2 (core)
info 4.0.9 (core)
limits 4.0.3 (core)
marketplace 2.0.5 (core)
org 6.0.9 (core)
packaging 3.0.5 (core)
schema 4.0.5 (core)
settings 3.0.5 (core)
sobject 2.0.5 (core)
telemetry 4.0.5 (core)
templates 57.0.9 (core)
trust 4.0.9 (core)
user 5.0.1 (core)
OS and Version:
Windows_NT 10.0.26200
Shell:
powershell
Summary
A custom evaluation asserting on the agent's response text is unusable: the
server rewrites the documented $.generatedData.outcome reference into a
filter-expression JSONPath and then fails to parse its own rewrite.
Test case YAML (per the documented shape):
customEvaluations:
- label: "response contains estimat"
name: string_comparison
parameters:
- name: operator
value: contains
isReference: false
- name: actual
value: "$.generatedData.outcome"
isReference: true
- name: expected
value: "estimat"
isReference: false
Actual result
The evaluation returns status: ERROR on every case:
"errorMessage": "Error parsing JSONPath
'$.outputs[?(@.type == 'general.echo')].payload.planner_response.lastExecution.outcome'"
i.e. the server maps $.generatedData.outcome to a path containing a
filter expression [?(@.type == '...')], and the JSONPath implementation
evaluating it rejects filter syntax. The author's input is well-formed; the
failing path is generated internally.
Expected result
Either the generated path parses, or $.generatedData.outcome maps to a
filter-free path, or the reference is rejected at test-create time with a
clear message. An ERROR status at run time on the documented reference makes
deterministic response-text assertions impossible in Testing Center.
Why it matters
The LLM-judged expectedOutcome is lenient (see reproduction context in the
related issues: an agent whose instruction clause was removed still passed a
judge rubric that named the missing behavior). string_comparison is the
platform's only deterministic assertion over response text — precisely the
tool needed when the judge is too soft — and it errors as documented.
Environment
- sf CLI: <paste
sf version --verbose> - Org: Agentforce Developer Edition (
orgfarm-*.develop.my.salesforce.com)
- Lingua principale
- Nessun dato sulla lingua
- Stelle
- 571
- Fork
- 80
- Merge medio
- 2g 21h
- PR unite (30g)
- 3
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di forcedotcom/cli
-
investigating validated
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
forcedotcom/cli#3657 · 2 commenti ·
-
`sf agent preview` fails with `AgentApiNotFound` against staging (aws-stage1) orgs — `stage.api.salesforce.com` missing from endpoint fallbackForse già presa Una pull request collegata a questa issue è aperta o già unita. Apertaarea:afdx owned by another team
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
forcedotcom/cli#3645 · 2 commenti ·
-
bug investigating validated
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
forcedotcom/cli#3644 · 6 commenti ·
-
sf agent mcp asset replace --assets null throws raw TypeError instead of InvalidShapeForse già presa @konkonrong-lgtm l’ha presa 56 giorni fa. Apertaarea:afdx bug investigating owned by another team validated
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
forcedotcom/cli#3625 · 4 commenti ·
-
area:afdx bug owned by another team
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
forcedotcom/cli#3608 · 2 commenti ·
Tutte le issue di forcedotcom/cli
Issue simili
-
Help text refers to commands by alias or without the `pulumi` prefix, plus small wording errorsApertaBug pulumi/pulumi
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
andromarces/agent-loops#571 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
stripe/stripe-cli#2130 ·
I maintainer di solito rispondono entro 1 giorno
-
area/install reliability status/ready
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
CLI: TUI shows onboarding when the provider's API key is only in the environment (e.g. OPENROUTER_API_KEY)Forse già presa Una pull request collegata a questa issue è aperta o già unita. ApertaCLI
Difficoltà 2/5 1-3 ore Idoneità per principianti 67/100
cline/cline#14923 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno